Signaling of Downsampling Filter for Chroma Intra Prediction Mode from Luma

The method addresses the inefficiencies in existing video coding standards by allowing flexible downsampling filter selection and application in CfL intra prediction, enhancing compression efficiency and video quality.

JP2025517842APending Publication Date: 2025-06-12TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024518227
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-09
Filing Date
2022-11-15
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing video coding standards, such as AV1 and HEVC, lack efficient mechanisms for signaling and implementing downsampling filters for intra-prediction modes between components, which affects compression efficiency and prediction accuracy.

Method used

The proposed method involves a processor-based system that receives a current block from a coded video bitstream, determines the appropriate downsampling filter coefficients based on a syntax element, and applies these coefficients using different numbers of sampling positions for each filter, allowing for flexible downsampling in CfL intra prediction mode.

Benefits of technology

This approach enhances the flexibility and accuracy of chroma from luma (CfL) intra prediction by allowing the selection of optimal downsampling filters, leading to improved compression efficiency and video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517842000001_ABST
    Figure 2025517842000001_ABST
Patent Text Reader

Abstract

Receiving a current block from a coded video bitstream (VBS), obtaining from the VBS a syntax element indicating which of two or more downsampling filters (DSFs) is used to predict the current block in a chroma from luma (CfL) intra prediction mode, determining a plurality of filter coefficients according to a first DSF in response to the syntax element indicating that the first DSF is used for the current block, downsampling the current block based on the determined plurality of coefficients using a first number of sampling positions, determining a plurality of filter coefficients according to a second DSF in response to the syntax element indicating that the second DSF is used for the current block, and downsampling the current block based on the determined plurality of coefficients using a second number of sampling positions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 346,412, filed on May 27, 2022, and U.S. Patent Application No. 18 / 054,054, filed on November 9, 2022, in the United States Patent and Trademark Office, the disclosures of which are hereby incorporated by reference in their entireties.

[0002] Embodiments of the present disclosure relate to a set of advanced video coding techniques, and more particularly, to signaling a downsampling filter for an intra - prediction mode between components.

Background Art

[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. Developed by the Alliance for Open Media (AOMedia), a consortium established in 2015 as a successor to VP9, it includes semiconductor companies, video - on - demand providers, video content producers, software development companies, and web browser vendors. Many of the components of the AV1 project were sourced from previous research efforts by alliance members. Individual contributors started experimental technology platforms years ago: Xiph / Mozilla's Daala released code in 2010, Google's experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco's Thor was released on August 11, 2015. Based on the construction of the VP9 codebase, AV1 incorporates additional technologies, some of which were developed in these experimental formats. The first version of the AV1 reference codec, version 0.1.0, was released on April 7, 2016. The alliance released the AV1 bitstream specification on March 28, 2018, along with reference software - based encoders and decoders. On June 25, 2018, the validation version 1.0.0 of this specification was released. On January 8, 2019, the "AV1 Bitstream & Decoding Process Specification", which is the validation version 1.0.0 with Errata 1 of this specification, was released. The AV1 bitstream specification includes the reference video codec. "AV1 Bitstream & Decoding Process Specification" (version 1.0.0 with Errata 1), The Alliance for Open Media (January 8, 2019) is incorporated herein by reference in its entirety.

[0004] The High Efficiency Video Coding (HEVC) standard has been jointly developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) standardization bodies. To develop the HEVC standard, these two standardization bodies collaborate in a partnership known as the Joint Collaborative Team on Video Coding (JCT-VC). The first version of the HEVC standard was completed in January 2013, resulting in aligned texts published by both ITU-T and ISO / IEC. Subsequently, additional research was organized to extend the standard to support some additional application scenarios, including extended range use with improved accuracy and color format support, scalable video coding, and 3-D / stereo / multi-view video coding. In ISO / IEC, the HEVC standard became MPEG-H Part 2 (ISO / IEC 23008-2), and in ITU-T, it became ITU-T Recommendation H.265. The specification of the HEVC standard, "SERIES H:AUDIOVISUAL AND MULTIMEDIA SYSTEMS,Infrastructure of audiovisual services-Coding of moving video", ITU-T H.265, International Telecommunication Union (April 2015), is hereby incorporated by reference in its entirety into this specification.

[0005] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (version 1), 2014 (version 2), 2015 (version 3), and 2016 (version 4). Since then, they have been studying the potential need for standardization of future video coding technologies that can significantly outperform HEVC in compression capabilities. In October 2017, they announced the Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, 22 CfP responses for standard dynamic range (SDR), 12 CfP responses for high dynamic range (HDR), and 12 CfP responses in the 360 video category were submitted respectively. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th Joint Video Exploration Team-Joint Video Expert Team (JVET) meeting. The JVET carefully evaluated and officially started the standardization of next-generation video coding beyond HEVC, namely the so-called Versatile Video Coding (VVC). The specification of the VVC standard, "Versatile Video Coding (Draft 7)", JVET-P2001-vE, Joint Video Experts Team (October 2019), is hereby incorporated by reference in its entirety. Another specification of the VVC standard, "Versatile Video Coding (Draft 10)", JVET-S2001-vE, Joint Video Experts Team (July 2020), is hereby incorporated by reference in its entirety.

Summary of the Invention

Means for Solving the Problems

[0006] According to one aspect of the present disclosure, a method for performing CfL intra prediction from luma includes being executed by at least one processor, receiving a current block from a coded video bitstream, obtaining from the coded video bitstream a syntax element indicating which of two or more downsampling filters is used to predict the current block in CfL intra prediction mode, determining a plurality of filter coefficients according to a first downsampling filter in response to the syntax element indicating that the first downsampling filter is used for the current block, downsampling the current block based on the determined plurality of coefficients using a first number of sampling positions, determining a plurality of filter coefficients according to a second downsampling filter in response to the syntax element indicating that the second downsampling filter is used for the current block, downsampling the current block based on the determined plurality of coefficients using a second number of sampling positions, wherein the second number of sampling positions is different from the first number of sampling positions, and reconstructing the current block after downsampling the current block.

[0007] According to one aspect of the present disclosure, a device for performing CfL intra prediction from luma includes at least one memory configured to store program code, and at least one processor configured to access the program code and operate as instructed by the program code. The program code includes reception code configured to cause the at least one processor to receive a current block from a coded video bitstream, acquisition code configured to cause the at least one processor to obtain a syntax element from the coded video bitstream indicating which of two or more downsampling filters is used to predict the current block in CfL intra prediction mode, and downsampling code configured to cause the at least one processor to determine a plurality of filter coefficients according to a first downsampling filter in response to the syntax element indicating that the first downsampling filter is used for the current block, use a first number of sampling positions, and downsample the current block based on the determined plurality of coefficients, and to determine a plurality of filter coefficients according to a second downsampling filter in response to the syntax element indicating that the second downsampling filter is used for the current block, use a second number of sampling positions, and downsample the current block based on the determined plurality of coefficients, where the second number of sampling positions is different from the first number of sampling positions, and reconstruction code configured to cause the at least one processor to reconstruct the current block after downsampling the current block.

[0008] According to one aspect of the present disclosure, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors of a device to perform a chroma from luma (CfL) intra prediction, cause the one or more processors to receive a current block from a coded video bitstream, obtain a syntax element from the coded video bitstream indicating which of two or more downsampling filters is used to predict the current block in CfL intra prediction mode, cause a plurality of filter coefficients to be determined according to a first downsampling filter in response to the syntax element indicating that the first downsampling filter is used for the current block, downsample the current block based on the determined plurality of coefficients using a first number of sampling positions, cause a plurality of filter coefficients to be determined according to a second downsampling filter in response to the syntax element indicating that the second downsampling filter is used for the current block, downsample the current block based on the determined plurality of coefficients using a second number of sampling positions, wherein the second number of sampling positions is different from the first number of sampling positions, and cause the current block to be reconstructed after downsampling the current block, the one or more instructions including.

[0009] Further features, properties, and various effects of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Mode for Carrying Out the Invention

[0011] In the present disclosure, the term "block" can be interpreted as a prediction block, a coding block, or a coding unit (CU). The term "block" here can also be used to refer to a transform block.

[0012] In the present disclosure, the term "transform set" refers to a group of transform kernel (or candidate) options. A transform set can include one or more transform kernel (or candidate) options. According to embodiments of the present disclosure, when two or more transform options are available, an index can be signaled to indicate which transform option in the transform set is applied for the current block.

[0013] In the present disclosure, the term "prediction mode set" refers to a group of prediction mode options. A prediction mode set may include one or more prediction mode options. According to an embodiment of the present disclosure, when two or more prediction mode options are available, an index may be further signaled to indicate which one of the prediction mode options in the prediction mode set is currently applied to a block for performing a prediction.

[0014] In the present disclosure, the term "adjacent reconstruction sample set" refers to a group of reconstruction samples from previously decoded adjacent blocks or reconstruction samples in a previously decoded picture.

[0015] In the present disclosure, the term "neural network" refers to the general concept of a data processing structure having one or more layers as described herein with respect to "deep learning for video coding". According to an embodiment of the present disclosure, any neural network may be configured to implement the embodiment.

[0016] FIG. 1 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The system (100) may include at least two terminals (110, 120) interconnected via a network (150). In the case of unidirectional data transmission, the first terminal (110) may code video data at a local location for transmission to the other terminal (120) via the network (150). The second terminal (120) may receive the coded video data of the other terminal from the network (150), decode the coded data, and display the restored video data. Unidirectional data transmission may be common in media providing applications and the like.

[0017] FIG. 1 shows a second pair of terminals (130, 140) provided to support two-way transmission of coded video that may occur, for example, during a video conference. In the case of two-way transmission of data, each terminal (130, 140) may code video data captured at a local location for transmission to the other terminal via a network (150). Each terminal (130, 140) may also receive coded video data transmitted by the other terminal, may decode the coded data, and may display the restored video data on a local display device.

[0018] In FIG. 1, terminals (110-140) may be shown as servers, personal computers, and smartphones, and / or any other type of terminal. For example, terminals (110-140) may be laptop computers, tablet computers, media players, and / or dedicated video conferencing equipment. Network (150) represents any number of networks that transmit coded video data between terminals (110-140), including, for example, wired and / or wireless communication networks. Communication network (150) may exchange data over circuit-switched channels and / or packet-switched channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of this discussion, the architecture and topology of network (150) may not be important for the operation of the present disclosure, unless otherwise described below.

[0019] FIG. 2 shows the placement of video encoders and decoders in a streaming environment as an example of the use of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-related applications, such as, for example, storage of compressed video on digital media including video conferencing, digital television, CDs, DVDs, memory sticks, etc.

[0020] As shown in FIG. 2, the streaming system (200) may include an intake subsystem (213) that can include a video source (201) and an encoder (203). The video source (201) may be, for example, a digital camera and may be configured to create an uncompressed video sample stream (202). The uncompressed video sample stream (202) can provide a high data volume compared to an encoded video bitstream and can be processed by an encoder (203) coupled to the video source (201), which can be, for example, a camera. The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video bitstream (204) may include a lower data volume compared to the sample stream and can be stored on a streaming server (205) for future use. One or more streaming clients (206) can access the streaming server (205) to obtain a video bitstream (209) that can be a copy of the encoded video bitstream (204).

[0021] In an embodiment, the streaming server (205) may also function as a Media-Aware Network Element (MANE). For example, the streaming server (205) may be configured to prune the encoded video bitstream (204) to match one or more of the streaming clients (206) with potentially different bitstreams. In an embodiment, the MANE may be provided separately from the streaming server (205) in the streaming system (200).

[0022] The streaming client (206) can include a video decoder (210) and a display (212). The video decoder (210) can decode, for example, a video bitstream (209) that is an input copy of an encoded video bitstream (204) and generate an output video sample stream (211) that can be rendered on the display (212) or another rendering device (not shown). In some streaming systems, the video bitstreams (204, 209) can be encoded according to a particular video coding / compression standard. Examples of such standards include, but are not limited to, ITU-T Recommendation H.265. A video coding standard, informally known as Versatile Video Coding (VVC), is under development. Embodiments of the present disclosure can be used in the context of VVC.

[0023] FIG. 3 shows an exemplary functional block diagram of a video decoder (210) attached to a display (212) according to one embodiment of the present disclosure.

[0024] The video decoder (210) can include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (). In at least one embodiment, the video decoder (210) can include an integrated circuit, a series of integrated circuits, and / or other electronic circuits. The video decoder (210) may also be embodied, in part or in whole, in software operating on one or more CPUs having associated memory.

[0025] In this and other embodiments, the receiver (310) may receive one or more coded video sequences to be decoded by the decoder (210) one coded video sequence at a time, and the decoding of each coded video sequence is independent of other coded video sequences. The coded video sequence may be received from a channel (312) that may be a hardware / software link to a storage device storing the encoded video data. The receiver (310) may receive the encoded video data along with other data that may be transferred to respective usage entities (not shown), such as coded audio data and / or auxiliary data streams. The receiver (310) may separate the coded video sequence from the other data. To counter network jitter, a buffer memory (315) may be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter "parser"). When the receiver (310) is receiving data from a storage / transfer device with sufficient bandwidth and controllability or from a synchronous network, the buffer memory (315) may not be used or may be small. When used in a best-effort packet network such as the Internet, the buffer memory (315) may be required, may be relatively large, and may be of an adaptable size.

[0026] The video decoder (210) may include a parser (320) for reconstructing symbols (321) from an entropy-coded video sequence. The categories of these symbols include, for example, information used to manage the operation of the decoder (210) and potentially information for controlling a rendering device such as a display (212) that may be coupled to the decoder as shown in FIG. 2. The control information for the rendering device(s) may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (320) may parse / entropy-decode the received coded video sequence. The coding of the coded video sequence may be in accordance with a video coding technique or video coding standard and may follow principles well known to those skilled in the art, including variable-length coding, Huffman coding, arithmetic coding with or without context dependence, etc. The parser (320) may extract a set of sub-group parameters of at least one sub-group of pixels within the video decoder based on at least one parameter corresponding to a group from the coded video sequence. The sub-groups may include a Group of Pictures (GOP), picture, tile, slice, macroblock, coding unit (CU), block, transform unit (TU), prediction unit (PU), etc. The parser (320) may also extract information such as transform coefficients, quantization parameter values, motion vectors, etc. from the coded video sequence.

[0027] The parser (320) may perform an entropy decoding / syntax analysis operation on the video sequence received from the buffer memory (315) to create the symbols (321).

[0028] The reconstruction of symbol (321) can involve multiple different units depending on the type of the coded video picture or a part thereof (such as inter and intra pictures, inter and intra blocks), as well as other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (320). Such a flow of subgroup control information between the parser (320) and the following multiple units is not shown for clarity.

[0029] Beyond the function blocks already described, the decoder (210) can be conceptually subdivided into several functional units as described below. In an actual implementation operating under commercial constraints, many of these units interact closely with each other and can at least partially be integrated with each other. However, for the purpose of explaining the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0030] One unit can be the scaler / inverse transform unit (351). The scaler / inverse transform unit (351) can receive from the parser (320) quantization transform coefficients and control information including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc. as symbols (s) (321). The scaler / inverse transform unit (351) can output a block including sample values that can be input to the aggregator (355).

[0031] In some cases, the output samples of the scaler / inverse transform unit (351) may be related to intra-coded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from the previously reconstructed part of the current picture. Such prediction information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) uses the surrounding already reconstructed information fetched from the current (partially reconstructed) picture in the current picture memory (358) to generate a block of the same size and shape as the block being reconstructed. The aggregator (355) may, in some cases, add, sample by sample, the prediction information generated by the intra prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).

[0032] In other cases, the output samples of the scaler / inverse transform unit (351) may potentially be related to inter-coded, motion-compensated blocks. In such cases, the motion-compensation prediction unit (353) can access the reference picture memory (357) to fetch the samples used for prediction. After motion-compensating the samples fetched according to the symbols (321) related to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) to generate the output sample information (in this case, called residual samples or a residual signal). The address in the reference picture memory (357) where the motion-compensation prediction unit (353) fetches the prediction samples can be controlled by the motion vector. The motion vector can be in the form of, for example, symbols (321) having X, Y, and reference picture components and may be available to the motion-compensation prediction unit (353). Motion compensation can also include interpolation of the sample values fetched from the reference picture memory (357) when an exact motion vector with sub-samples is used, a motion vector prediction mechanism, etc.

[0033] The output samples of the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356). The video compression technique is controlled by parameters included in the coded video bitstream and can include in-loop filter techniques made available to the loop filter unit (356) as symbols (321) from the parser (320), but can also respond to meta information obtained during the decoding of the previous part (in decoding order) of the coded picture or coded video sequence, and can also respond to previously reconstructed and loop filtered sample values.

[0034] The output of the loop filter unit (356) can be an output sample stream that can be output to a rendering device such as a display (212) and can also be stored in the reference picture memory (357) for use in future inter-picture prediction.

[0035] Once fully reconstructed, a particular coded picture can be used as a reference picture for future prediction. When a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture can become part of the reference picture memory (357), and a new current picture memory can be reallocated before starting the reconstruction of the next coded picture.

[0036] The video decoder (210) may perform a decoding operation according to a predetermined video compression technique that may be documented in a standard such as ITU-T Rec.H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence is specified in the video compression technique document or standard, particularly in the profile document therein. Also, in order to conform to some video compression technique or standard, the complexity of the coded video sequence may also be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may in some cases be further restricted by the virtual reference decoder (HRD) specification and the metadata of the HRD buffer management signaled in the coded video sequence.

[0037] In one embodiment, the receiver (310) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data can be in the form of, for example, a temporal layer, a spatial layer, or an SNR enhancement layer, a redundant slice, a redundant picture, a forward error correction code, etc.

[0038] FIG. 4 shows an exemplary functional block diagram of a video encoder (203) associated with a video source (201) according to an embodiment of the present disclosure.

[0039] The video encoder (203) may include, for example, an encoder which is a source coder (430), a coding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy coder (445), a controller (450), and a channel (460).

[0040] The encoder (203) may receive video samples from a video source (201) (which is not part of the encoder) that can capture video images (s) to be coded by the encoder (203).

[0041] The video source (201) can provide the source video sequence to be coded by the encoder (203) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media supply system, the video source (201) can be a storage device that stores previously prepared video. In a video conferencing system, the video source (201) can be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0042] According to one embodiment, the encoder (203) can code and compress pictures of the source video sequence into the coded video sequence (443) in real time or under any other arbitrary time constraints required by the application. Enforcing an appropriate coding speed is one function of the controller (450). The controller (450) may also control other functional units as described later and may be functionally coupled to these units. The coupling is not shown for clarity. The parameters set by the controller (450) can include rate control related parameters (such as picture skip, quantizer, lambda value of rate distortion optimization techniques), picture size, Group of Pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller (450) since they can relate to the video encoder (203) optimized for a specific system design.

[0043] Some video encoders operate in what those skilled in the art would readily recognize as a "coding loop." As an overly simplified explanation, the coding loop can consist of an encoding part of a source coder (430) (responsible for creating symbols based on the input picture to be coded and the (one or more) reference pictures), and a (local) decoder (433) incorporated in an encoder (203) that reconstructs symbols to create sample data that a (remote) decoder would also create when the compression between the symbols and the coded video bitstream is reversible in a particular video compression technique. The reconstructed sample stream can be input into a reference picture memory (434). Since bit-exact results are obtained regardless of the position of the decoder (local or remote) by decoding the symbol stream, the contents of the reference picture memory are also bit-exact between the local encoder and the remote encoder. In other words, the prediction part of the encoder "sees" the same sample values as the reference picture samples that the decoder will "see" when using prediction during decoding. This basic principle of the synchronization of reference pictures (and the resulting drift if synchronization cannot be maintained, for example due to channel errors) is known to those skilled in the art.

[0044] The operation of the "local" decoder (433) can be considered the same as the operation of the "remote" decoder (210) already described in detail in relation to FIG. 3. However, since the symbols are available and the encoding / decoding of the symbols into the coded video sequence by the entropy coder (445) and the parser (320) can be reversible, the entropy decoding part of the decoder (210) including the channel (312), the receiver (310), the buffer memory (315), and the parser (320) may not be fully implemented in the local decoder (433).

[0045] What can be said at this point is that any decoder technology, except for the parse / entropy decoding that exists within the decoder, may need to exist in substantially the same functional form within the corresponding encoder. For this reason, the disclosed subject matter focuses on decoder operations. The description of encoder technologies can be omitted since they can be the reverse of the decoder technologies that have been comprehensively described. More detailed descriptions are necessary only in specific areas and are presented below.

[0046] As part of the operation, the source coder (430) can perform motion compensation predictive coding that predictively codes an input frame by referring to one or more previously coded frames from a video sequence designated as a "reference frame". In this way, the coding engine (432) codes the difference between a pixel block of the input frame and a pixel block (s) of a reference frame (s) that can be selected as a predictive reference (s) to the input frame.

[0047] The local video decoder (433) can decode the coded video data of a frame that can be designated as a reference frame based on the symbols created by the source coder (430). The operation of the coding engine (432) can advantageously be an irreversible process. When the coded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence can typically be a reproduction of the source video sequence with some error. The local video decoder (433) can reproduce the decoding process that can be performed by the video decoder for the reference frame and store the reconstructed reference frame in the reference picture memory (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame having the same content as the reconstructed reference frame obtained by the remote video decoder (without transmission errors).

[0048] Predictor (435) can perform predictive search for the coding engine (432). That is, for a new frame to be coded, predictor (435) can search the reference picture memory (434) by obtaining sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors and block shapes that can serve as appropriate predictive references for the new picture. Predictor (435) can operate on one sample block per pixel block to find an appropriate predictive reference. In some cases, the input picture can have predictive references drawn from a plurality of reference pictures stored in the reference picture memory (434) as determined by the search results obtained by predictor (435).

[0049] Controller (450) can manage the coding operations of video coder (430), including, for example, setting parameters and subgroup parameters used for encoding video data.

[0050] The outputs of all the aforementioned functional units can be entropy-coded by entropy coder (445). The entropy coder converts the symbols generated by various functional units into an encoded video sequence by reversibly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0051] Transmitter (440) can buffer the encoded video sequence(s) created by entropy coder (445) for transmission via communication channel (460), which can be a hardware / software link to a storage device that will store the encoded video data. Transmitter (440) can merge the encoded video data from video coder (430) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0052] The controller (450) may manage the operation of the encoder (203). During coding, the controller (450) may assign a specific coded picture type to each coded picture, which may affect the coding technique applicable to each picture. For example, a picture may often be assigned as an intra picture (I picture), a predicted picture (P picture), or a bi-directionally predicted picture (B picture).

[0053] An intra picture (I picture) may be a picture that can be coded and decoded without using any other frame in the sequence as a source of prediction. Some video codecs allow for various types of intra pictures, such as, for example, an independent decoder refresh (IDR) picture. Those skilled in the art are aware of these variations of I pictures and their respective uses and characteristics.

[0054] A predicted picture (P picture) may be one that can be coded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0055] A bi-directionally predicted picture (B picture) may be one that can be coded and decoded using intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predicted pictures may use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0056] Source pictures are generally spatially subdivided into a plurality of sample blocks (e.g., blocks of samples of 4×4, 8×8, 4×8, or 16×16 samples each) and can be coded block by block. The blocks can be coded predictively by referring to other (already coded) blocks as determined by the coding assignment applied to each block of the picture. For example, blocks of an I picture may be coded non-predictively or may be coded predictively by referring to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be coded non-predictively via spatial prediction or via temporal prediction that refers to one previously coded reference picture. Blocks of a B picture can be coded non-predictively via spatial prediction or via temporal prediction that refers to one or two previously coded reference pictures.

[0057] The video coder (203) can perform coding operations according to a predetermined video coding technology or standard such as ITU-T Rec.H.265. In its operation, the video coder (203) can perform various compression operations including predictive coding operations that utilize temporal and spatial redundancies in the input video sequence. Thus, the coded video data can conform to the syntax specified by the video coding technology or standard being used.

[0058] In one embodiment, the transmitter (440) can transmit additional data along with the coded video. The video coder (430) can include such data as part of the coded video sequence. The additional data can include temporal layer / spatial layer / SNR enhancement layer, other forms of redundant data such as redundant pictures and slices, supplementary enhancement information (SEI) messages, visual user utility information (VUI) parameter set fragments, and the like.

[0059] [Directional Intra Prediction in AV1] VP9 supports eight directional modes corresponding to angles from 45 degrees to 207 degrees. To exploit more diverse spatial redundancy in the direction texture, in AV1, the directional intra mode is extended to angles set with finer granularity. The original eight angles are slightly modified to be nominal angles, and these eight nominal angles are named V_PRED(542), H_PRED(543), D45_PRED(544), D135_PRED(545), D113_PRED(5446), D157_PRED(547), D203_PRED(548), and D67_PRED(549), which are shown in Figure 5 for the current block (541). For each nominal angle, there are seven finer angles, so AV1 has a total of 56 direction angles. The predicted angle is represented by adding an angle delta that is -3 to 3 times the step size of 3 degrees to the nominal intra angle. In AV1, eight nominal modes are first signaled together with five non-angle smoothing modes. Then, if the current mode is an angle mode, an index is further signaled to indicate the angle delta with respect to the corresponding nominal angle. To implement the directional prediction mode in AV1 by a general method, all 56 directional intra prediction modes of AV1 are implemented using a unified direction predictor that projects each pixel to a reference sub-pixel position and interpolates the reference pixels by a 2-tap bilinear filter.

[0060] [Non-directional Smoothing Intra Predictor in AV1] AV1 has five non - directional smoothing intra - prediction modes: DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. In DC prediction, the average value of the left and upper adjacent samples is used as the predictor for the block to be predicted. In the case of the PAETH predictor, first the upper, left, and upper - left reference samples are fetched, and then the value closest to (upper + left - upper - left) is set as the predictor for the pixel to be predicted. FIG. 6 shows the positions of the upper sample (554), left sample (556), and upper - left sample (558) for the pixel (552) in the current block (550). In the SMOOTH mode, SMOOTH_V mode, and SMOOTH_H mode, the current block (550) is predicted using vertical or horizontal bilinear interpolation, or the average value in both directions.

[0061] [Chroma predicted from luma] In addition to the above - mentioned modes, chroma from luma (CfL) is a chroma - only intra - predictor that models chroma pixels as a linear function of the corresponding reconstructed luma pixels. CfL prediction can be expressed as shown in Equation 1 below.

[0062] [Equation]

[0063] In the formula, L AC represents the AC contribution of the luma component, α represents the parameter of the linear model, and DC represents the DC contribution of the chroma component.

[0064] FIG. 7 provides a graph of the linear function described by equation (1). As seen in FIG. 7 and equation (1), the reconstructed luma pixels are subsampled to chroma resolution and then the average value is subtracted to form the AC contribution. Instead of requiring the decoder to calculate a scaling parameter as in some prior arts to approximate the chroma AC component from the AC contribution, AV1 CfL can determine the parameter α based on the original chroma pixels and can signal them within the bitstream. This reduces the complexity of the decoder and obtains a more accurate prediction. Regarding the DC contribution of the chroma component, it is sufficient for most chroma contents and can be calculated using the intra DC mode which has a mature high-speed implementation form.

[0065] In CfL mode, when some samples in a collocated luma block are outside the picture boundary, these samples can be padded and used to calculate the average of the luma samples. As shown in FIG. 8, the samples in area 802 of the current block 800 are outside the picture as indicated by the picture boundary 804, and these samples can be padded by copying the values of the nearest available samples within the current block 800, for example, the samples included in area 806 of the current block 800.

[0066] In CfL mode, the luma subsampling step is combined with the average subtraction step as shown in FIG. 7. In this way, not only is the equation simplified, but also the subsampling division and the corresponding rounding error are removed. An example of the equation corresponding to the combination of both steps is given by equation (2), and this can be simplified to form equation (3).

[0067]

Equation

[0068]

Equation

[0069] Note that integer division is used in both equations. MxN represents a matrix of pixels in the luma plane.

[0070] Based on the supported chroma subsampling, Sx × Sy ∈ {1, 2, 4}, and since both M and N are powers of 2, M × N is also a power of 2.

[0071] For example, in the context of 4:2:0 chroma subsampling, instead of applying a box filter, the proposed approach only requires summing four reconstructed luma pixels that match the chroma pixel. That is, a 4-tap {1 / 4, 1 / 4, 1 / 4, 1 / 4} filter is used to downsample the collocated luma samples to match the chroma resolution. Then, when CfL scales that luma pixel to improve the prediction accuracy, the approach of some embodiments described herein can only scale by 2.

[0072] [Chroma Downsampling Format] In embodiments, there may be different YUV formats, i.e., chroma formats, depending on different chroma downsampling phases, an example of which is shown in FIG. 9. Different chroma formats define different downsampling grids or phases for different color components. In the case of the 4:2:0 format, as shown in FIG. 9, there can be two different downsampling formats, sometimes referred to as 4:2:0 MPEG1 or 4:2:0 MPEG2.

[0073] [Downsampling Filter] In the case of the luma downsampling filter in AV1, the following equation (4) is applied to derive the reconstructed luma sample.

[0074] [Equation]

[0075] In the above formula (4) and the following formula, rec’ L represents the pixel value of the downsampled and reconstructed luma pixel at location (i,j), which can be the location of the corresponding chroma pixel. Also, rec L (x,y) represents the restored luma sample value located at position (x,y).

[0076] Assume a chroma downsampling format 1000 shown in FIG. 10 that the downsampling filter in AV1 can correspond to the 4:2:0 MPEG1 downsampling format of FIG. 9. Specifically, FIG. 10 shows a chroma sample 1002 and four corresponding luma samples 1004 indexed from 0 to 3. In an embodiment, the chroma sample can be a chroma pixel, and the luma sample can be a luma pixel. In an embodiment, the luma sample can be used to determine the value of the downsampled luma sample corresponding to the chroma sample, or the value of the downsampled luma pixel corresponding to the chroma pixel according to the above formula (4).

[0077] In some implementations of the CfL mode, only one downsampling filter is supported, but for some content or different chroma downsampling formats, the current unique downsampling filter may not be the optimal filter.

[0078] Therefore, an embodiment can support two or more downsampling filters for luma reconstructed samples when an inter-component prediction mode such as the CfL prediction mode is selected.

[0079] The embodiments may also be applied to modes other than the CfL prediction mode. For example, it may be another prediction mode that uses one color component to predict another color component that requires downsampling for one or more color components. Thus, the embodiments may also be applied, for example, by replacing luma with one specific color component (e.g., R) and replacing chroma with another specific color component (e.g., G or B).

[0080] According to an embodiment, the plurality of downsampling filters may include the downsampling filter described above with respect to FIG. 10 and Equation (4), and may also include a 6-tap filter. FIG. 11 shows a chroma downsampling format 1100 corresponding to a 6-tap filter according to an embodiment. As seen in FIG. 11, the chroma downsampling format 1100 includes a chroma sample 1102 and six corresponding luma samples 1104 indexed from 0 to 5. In an embodiment, the luma samples 1104 may be reconstructed luma samples, which may be used to derive the downsampled luma pixels corresponding to the chroma sample 1102 that may be located at the location (i, j) based on the following Equation (5).

[0081] [Number]

[0082] In an embodiment, it can be assumed that the 6-tap downsampling filter corresponds to a 4:2:0 MPEG2 downsampling format when the chroma downsampling format is considered.

[0083] In an embodiment, additional filters, for example four filters, can be used. For example, downsampling filters corresponding to the following Equations (6) to (10) may be used.

[0084] [Number]

[0085]

Number

[0086]

Number

[0087]

Number

[0088]

Number

[0089] In an embodiment, a high-level syntax element such as a flag or an index may be signaled to indicate which downsampling filter is used in a downsampling process that involves a coding / decoding process that requires downsampling luma to a lower resolution and aligning it with chroma, such as an intra prediction mode between components like the CfL prediction mode or other inter-component prediction methods. In an embodiment, the high-level syntax element may be signaled in at least one of a sequence header or a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a video parameter set (VPS), a slice header (SH), a picture header (PH), a frame header, a tile header, a coding tree unit (CTU) header, a superblock header, and a block having a specific predefined block size (e.g., 32×32, 64×64).

[0090] In an embodiment, the filter coefficients can be customized and the coefficients can be explicitly signaled in the bitstream at a high level syntax. The high level syntax can include at least one of a sequence header or SPS, PPS, APS, VPS, SH, PH, a frame header, a tile header, a CTU header, a superblock header, and a block having a specific predefined block size (e.g., 32×32, 64×64).

[0091] In an embodiment, when the filter coefficients are explicitly signaled, a first flag indicating whether the filter has an odd or even number of taps can be signaled.

[0092] In an embodiment, when the filter coefficients are explicitly signaled, another number N is signaled to derive the number of filter taps. For example, when the filter has an odd number of filter taps, the number of filter taps can be derived as one of 2*N + 1 or 2*N - 1. As another example, when the filter has an even number of filter taps, the number of filter taps can be derived based on one of 2*N or 2*(N - 1) or 2*(N + 1).

[0093] In an embodiment, N filter coefficients can be further signaled and the filter is derived based on assuming that the filter is symmetric and determining whether the filter has an odd or even number of taps and the N signaled filter coefficients.

[0094] FIG. 12 is a flowchart of a process 1200 for performing chroma from luma (CfL) intra prediction according to an embodiment.

[0095] As shown in FIG. 12, in operation 1202, process 1200 includes receiving a current block from a coded video bitstream.

[0096] As further shown in FIG. 12, in operation 1204, process 1200 includes obtaining a syntax element indicating which of two or more downsampling filters is used to predict the current block in the CfL intra prediction mode from the coded video bitstream. In an embodiment, the downsampling filter may be any of the filters described above. In an embodiment, the syntax element may be a high-level syntax element as described above. In an embodiment, the syntax element may be signaled in at least one of a sequence header, a sequence parameter set, a picture parameter set, an adaptive parameter set, a video parameter set, a slice header, a picture header, a frame header, a tile header, a coding tree unit header, a superblock header, or a block having a predetermined block size.

[0097] As further shown in FIG. 12, in operation 1206, process 1200 includes determining which of two or more downsampling filters is used for the current block based on the syntax element.

[0098] As further shown in FIG. 12, based on a syntax element indicating a first downsampling filter among two or more downsampling filters, in operation 1208, process 1200 determines a plurality of filter coefficients according to the first downsampling filter, and in operation 1210, downsamples the current block based on the determined plurality of coefficients using a first number of sampling positions. After the downsampling is performed, process 1200 can proceed to operation 1216.

[0099] As further shown in FIG. 12, based on a syntax element indicating a second downsampling filter among two or more downsampling filters, in operation 1212, process 1200 determines a plurality of filter coefficients according to the second downsampling filter, and in operation 1214, uses a second number of sampling positions to downsample a current block based on the determined plurality of coefficients. In an embodiment, the second number of sampling positions may be different from the first number of sampling positions. After downsampling is performed, process 1200 can proceed to operation 1216.

[0100] As further shown in FIG. 12, in operation 1216, process 1200 includes reconstructing a current block after downsampling the current block.

[0101] In an embodiment, the syntax element may include a first flag indicating whether the number of taps corresponding to a plurality of filter coefficients is even or odd.

[0102] In an embodiment, the syntax element may indicate a predetermined number used to derive the number of taps corresponding to a plurality of filter coefficients.

[0103] In an embodiment, based on the number of taps being odd, the number of taps may be determined to be equal to one of 2*N + 1 or 2*N + 1, where N indicates a predetermined number.

[0104] In an embodiment, based on the number of taps being even, the number of taps may be determined to be equal to one of 2*N, 2*(N + 1), or 2*(N + 1), where N indicates a predetermined number.

[0105] In an embodiment, the value of at least one filter coefficient among a plurality of filter coefficients may be explicitly signaled, and the remaining filter coefficients among the plurality of filter coefficients may be determined based on the value of the at least one filter coefficient.

[0106] In some embodiments, the syntax element may include a first flag indicating whether the number of taps corresponding to a plurality of filter coefficients is even or odd, and the remaining filter coefficients may be further determined based on the first flag.

[0107] In embodiments, it may be assumed that the plurality of filter coefficients are symmetric.

[0108] FIG. 12 shows exemplary blocks of process 1200, but in some implementations, process 1200 may include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in FIG. 12. Additionally or alternatively, two or more of the blocks of process 1200 may be executed in parallel.

[0109] Furthermore, the proposed method may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium to execute one or more of the proposed methods.

[0110] The techniques of the embodiments of the present disclosure described above can be implemented as computer software physically stored on one or more computer-readable media using computer-readable instructions. For example, FIG. 13 shows a computer system (900) suitable for implementing embodiments of the disclosed subject matter.

[0111] The computer software can be coded using any suitable machine code or computer language that can be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or via interpretation, microcode execution, etc.

[0112] The command can be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.

[0113] The components shown in FIG. 13 for the computer system (900) are essentially exemplary and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure. The component configuration should not be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiment of the computer system (900).

[0114] The computer system (900) may include a specific human interface input device. Such a human interface input device may respond to input by one or more human users via, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), audio input (e.g., voice, clapping), visual input (e.g., gesture), olfactory input (not shown). The human interface device can also be used to capture specific media that is not necessarily directly related to conscious input by humans, such as audio (e.g., voice, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (including 2D video, 3D video such as stereoscopic video).

[0115] The input human interface device may include one or more (only one of each shown) of a keyboard (901), a mouse (902), a trackpad (903), a touch screen (910), a data glove, a joystick (905), a microphone (906), a scanner (907), and a camera (908).

[0116] The computer system (900) may also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, by tactile output, sound, light, and smell / taste. Such human interface output devices can include tactile output devices (e.g., tactile feedback by a touch screen (910), data glove, or joystick (905), although there may also be tactile feedback devices that do not function as input devices). For example, such devices can include audio output devices (such as speakers (909), headphones (not shown)), visual output devices (screens (910) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without a touch screen input function and each with or without a tactile feedback function, some of which can output three-dimensional images, two-dimensional visual images, or more than four dimensions by means such as virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0117] The computer system (900) may also include optical media having a CD / DVD or similar medium (921) such as a CD / DVD ROM / RW (920), a thumb drive (922), a removable hard drive or solid state drive (923), legacy magnetic media such as tapes and floppy disks (not shown), special ROM / ASIC / PLD-based devices such as security dongles (not shown), and other human-accessible storage devices and the media associated therewith.

[0118] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.

[0119] The computer system (900) can also include an interface to one or more communication networks. The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include local area networks such as Ethernet and wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., and television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, and vehicle and industrial including CANBus. A particular network generally requires an external network interface adapter attached to a particular general-purpose data port or peripheral bus (949) (such as a USB port of the computer system (900)). Other networks are generally integrated into the core of the computer system 900 by attachment to a system bus as described later (such as an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (900) can communicate with other entities. Such communication can be only unidirectional reception (such as broadcast TV), only unidirectional transmission (such as CANbus to a particular CANbus device), or bidirectional with other computer systems using, for example, local or wide area digital networks. Such communication can include communication to a cloud computing environment (955). Specific protocols and protocol stacks can be used for each of those networks and network interfaces as described above.

[0120] The aforementioned human interface device, human-accessible storage device, and network interface (954) can be attached to the core (940) of the computer system (900).

[0121] The core (940) can include one or more central processing units (CPUs) (941), a graphics processing unit (GPU) (942), a dedicated programmable processing device in the form of a field programmable gate array (FPGA) (943), a hardware accelerator (944) for specific tasks, etc. These devices can be connected via a system bus (948) together with a read-only memory (ROM) (945), a random access memory (946), an internal mass storage such as an internal non-user-accessible hard drive, SSD (947). In some computer systems, the system bus (948) can be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (948) of the core or via a peripheral bus (949). Architectures for peripheral buses include PCI, USB, etc. The graphics adapter (950) may be included in the core (940).

[0122] The CPU (941), GPU (942), FPGA (943), and accelerator (944) can execute specific instructions that can together constitute the aforementioned computer code. The computer code can be stored in the ROM (945) or the RAM (946). Migration data can also be stored in the RAM (946), while persistent data can be stored, for example, in the internal mass storage (947). Fast storage and retrieval for any of the memory devices can be enabled using a cache memory that can be closely associated with one or more CPUs (941), GPUs (942), mass storage (947), ROM (945), RAM (946), etc.

[0123] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be specially designed and constructed for the purposes of this disclosure or can be of the kinds available and well known to those skilled in the computer software arts.

[0124] Rather than being limiting, by way of example, the architecture corresponding to the computer system (900), specifically the core (940), can provide functionality as a result of software executed within one or more tangible computer-readable media. Such computer-readable media can be associated with user-accessible mass storage as described above, as well as specific storage of the core (940) that is of a non-transitory nature, such as internal core mass storage (947) or ROM (945). The software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (940). The computer-readable media can include one or more memory devices or chips according to specific needs. The software can cause the core (940), and specifically the processor (including CPU, GPU, FPGA, accelerator, etc.) therein, to define data structures stored in the RAM (946) and modify such data structures according to processes defined by the software, to execute specific processes described herein, or specific portions of specific processes. Additionally or alternatively, the computer system can provide functionality as a result of logic wired or otherwise embodied in circuitry (e.g., accelerator (944)) that can operate instead of or in conjunction with software to execute specific processes described herein or specific portions of specific processes. References to software can, as necessary, include logic, and vice versa. References to computer-readable media can, as necessary, include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0125] Embodiments of the present disclosure may be used individually or combined in any order. Further, each of the embodiments (and their methods) may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-transitory computer-readable medium.

[0126] The foregoing disclosure provides examples and explanations, but is not intended to be exhaustive or to limit the embodiments to the exact forms disclosed. Modifications and variations are possible in light of the above disclosure or may be acquired from practice of the implementation forms.

[0127] The term element as used in this application is intended to be construed broadly as hardware, firmware, or a combination of hardware and software.

[0128] Even if a combination of features is recited in the claims and / or disclosed in this specification, such a combination is not intended to limit the disclosure of possible implementation forms. In fact, many of these features can be combined in ways that are not explicitly recited in the claims and / or not explicitly disclosed in the specification. Each of the dependent claims listed below can depend directly on only one claim, but combinations of each dependent claim with any other claim in the set of claims are included in the disclosure of possible implementation forms.

[0129] Elements, operations, or instructions used in this application are not to be construed as being critical or essential, unless so specified. Also, the articles "a" and "an" used in this application are intended to include one or more things and can be used in the sense of "one or more". Further, the term "set" used in this application is intended to include one or more things (e.g., related things, unrelated things, combinations of related and unrelated things, etc.) and can be used in the sense of "one or more". Where only one thing is intended, the term "one" or a similar expression is used. Also, the terms "has", "have", "having", etc. used in this application are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise specified.

[0130] Although this disclosure describes several non-limiting and exemplary embodiments, there are changes, substitutions, and various alternative equivalents within the scope of this disclosure. Thus, those skilled in the art will understand that, although not explicitly illustrated or described herein, numerous systems and methods that embody the principles of this disclosure and are thus within the spirit and scope of this disclosure can be devised.

Description of Reference Numerals

[0131] 100 Communication system 110 Terminal 120 Terminal 130 Terminal 140 Terminal 150 Network 200 Streaming system 201 Video source 202 Uncompressed video sample stream 203 Encoder 204 Encoded video bitstream 205 Streaming server 206 Streaming client 209 Video bit stream 210 Video decoder 211 Video sample stream 212 Display 213 Capture subsystem 310 Receiver 312 Channel 315 Buffer memory 320 Entropy decoder / parser 321 Symbol 351 Scaler / inverse transform unit 352 Intra prediction unit 353 Motion compensation prediction unit 355 Aggregator 356 Loop filter unit 357 Reference picture memory 358 Current picture memory 430 Source coder 432 Coding engine 433 Decoder 434 Reference picture memory 435 Predictor 440 Transmitter 443 Video sequence 445 Entropy coder 450 Controller 460 Channel 550 Current block 552 Pixel 554 Upsample 556 Left sample 558 Upper left sample 800 Current block 802 Area 804 Picture boundary 806 Area 900 Computer system 901 Keyboard 902 Mouse 903 Track pad 905 Joystick 906 Microphone 907 Scanner 908 Camera 909 Speaker 910 Touch Screen 921 Medium 922 Thumb Drive 923 Solid State Drive 940 Core 941 Central Processing Unit (CPU) 942 Graphics Processing Unit (GPU) 943 Field Programmable Gate Array (FPGA) 944 Hardware Accelerator 945 Read Only Memory (ROM) 946 Random Access Memory 947 Internal Mass Storage 948 System Bus 949 Peripheral Bus 950 Graphics Adapter 954 Network Interface 955 Cloud Computing Environment 1000 Chroma Downsampling Format 1002 Chroma Sample 1004 Luma Sample 1100 Chroma Downsampling Format 1102 Chroma Sample 1104 Luma Sample 1102 Chroma Sample 1200 Process

Claims

**Claim 1** A method for performing a chroma from luma (CfL) intra prediction, the method being executed by at least one processor, receiving a current block from a coded video bitstream, obtaining, from the coded video bitstream, a syntax element indicating which of two or more downsampling filters is used to predict the current block in CfL intra prediction mode, in response to the syntax element indicating that a first downsampling filter is used for the current block, determining a plurality of filter coefficients according to the first downsampling filter, downsampling the current block based on the determined plurality of coefficients using a first number of sampling positions, in response to the syntax element indicating that a second downsampling filter is used for the current block, determining the plurality of filter coefficients according to the second downsampling filter, downsampling the current block based on the determined plurality of coefficients using a second number of sampling positions, wherein the second number of sampling positions is different from the first number of sampling positions, and reconstructing the current block after downsampling the current block. A method comprising: **Claim 2** The method according to claim 1, wherein the syntax element includes a first flag indicating whether the number of taps corresponding to the plurality of filter coefficients is even or odd. **Claim 3** The method according to claim 2, wherein the syntax element indicates a predetermined number used to derive the number of the taps corresponding to the plurality of filter coefficients. **Claim 4** Based on the number of the taps being odd, the number of the taps is determined to be equal to one of 2*N + 1 or 2*N + 1, where N indicates the predetermined number. The method according to claim 3. **Claim 5** Based on the number of the taps being even, the number of the taps is determined to be equal to one of 2*N, 2*(N + 1), or 2*(N + 1), where N indicates the predetermined number. The method according to claim 3. **Claim 6** The value of at least one of the plurality of filter coefficients is explicitly signaled, and the remaining filter coefficients of the plurality of filter coefficients are determined based on the value of the at least one filter coefficient. The method according to claim 1. **Claim 7** The syntax element includes a first flag indicating whether the number of taps corresponding to the plurality of filter coefficients is even or odd, and the remaining filter coefficients are further determined based on the first flag. The method according to claim 7. **Claim 8** The plurality of filter coefficients are assumed to be symmetric. The method according to claim 7. **Claim 9** A device for performing chroma from luma (CfL) intra prediction, the device comprising: at least one memory configured to store program code; at least one processor configured to access the program code and operate as instructed by the program code, the program code comprising: reception code configured to cause the at least one processor to receive a current block from a coded video bitstream; acquisition code configured to cause the at least one processor to acquire from the coded video bitstream a syntax element indicating which of two or more downsampling filters is used to predict the current block in CfL intra prediction mode; the at least one processor, in response to the syntax element indicating that a first downsampling filter is used for the current block, causing the plurality of filter coefficients to be determined according to the first downsampling filter; using a first number of sampling positions to downsample the current block based on the determined plurality of coefficients; in response to the syntax element indicating that a second downsampling filter is used for the current block, causing the plurality of filter coefficients to be determined according to the second downsampling filter; Downsampling code configured to downsample the current block based on the determined plurality of coefficients using a second number of sampling positions, the second number of sampling positions being different from the first number of sampling positions, At least one processor and reconstruction code configured to cause the at least one processor to reconstruct the current block after downsampling the current block. A device comprising: **Claim 10** The method of claim 9, wherein the syntax element includes a first flag indicating whether the number of taps corresponding to the plurality of filter coefficients is even or odd. **Claim 11** The method of claim 10, wherein the syntax element indicates a predetermined number used to derive the number of the taps corresponding to the plurality of filter coefficients. **Claim 12** Based on the number of the taps being odd, the number of the taps is determined to be equal to one of 2*N + 1 or 2*N + 1, where N indicates the predetermined number. The method of claim 11. **Claim 13** Based on the number of the taps being even, the number of the taps is determined to be equal to one of 2*N, 2*(N + 1), or 2*(N + 1), where N indicates the predetermined number. The method of claim 11. **Claim 14** The value of at least one filter coefficient among the plurality of filter coefficients is explicitly signaled, The remaining filter coefficients among the plurality of filter coefficients are determined based on the value of the at least one filter coefficient. The method of claim 9. **Claim 15** The method of claim 14, wherein the syntax element includes a first flag indicating whether the number of taps corresponding to the plurality of filter coefficients is even or odd, and the remaining filter coefficients are further determined based on the first flag. **Claim 16** The method of claim 14, wherein the plurality of filter coefficients are assumed to be symmetric. **Claim 17** A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a device for performing chroma from luma (CfL) intra prediction, cause the one or more processors to Receive a current block from a coded video bitstream, Obtain a syntax element indicating which of two or more downsampling filters is used to predict the current block in the CfL intra prediction mode from the coded video bitstream, In response to the syntax element indicating that the first downsampling filter is used for the current block, Cause a plurality of filter coefficients to be determined according to the first downsampling filter, Use a first number of sampling positions to downsample the current block based on the determined plurality of coefficients, In response to the syntax element indicating that the second downsampling filter is used for the current block, Cause the plurality of filter coefficients to be determined according to the second downsampling filter, Use a second number of sampling positions, different from the first number of sampling positions, to downsample the current block based on the determined plurality of coefficients, A non-transitory computer-readable medium including one or more instructions to reconstruct the current block after downsampling the current block. **Claim 18** The temporary computer-readable medium according to claim 17, wherein the syntax element includes a first flag indicating whether the number of taps corresponding to the plurality of filter coefficients is even or odd. **Claim 19** The temporary computer-readable medium according to claim 18, wherein the syntax element indicates a predetermined number used to derive the number of the taps corresponding to the plurality of filter coefficients. **Claim 20** The value of at least one of the plurality of filter coefficients is explicitly signaled, The remaining filter coefficients of the plurality of filter coefficients are determined based on the value of the at least one filter coefficient, the temporary computer-readable medium according to claim 17.