Improved signaling method for scaling parameters in chroma-from-luma intra prediction mode

The method addresses inefficiencies in signaling scaling parameters for cross-component intra prediction by deriving a predicted scaling factor from neighboring samples, improving decoder efficiency and prediction accuracy in video coding.

JP7733125B2Active Publication Date: 2025-09-02TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023556781
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-11-08
Filing Date
2022-11-09
Publication Date
2025-09-02
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

Existing video coding technologies, such as AV1 and HEVC, face challenges in efficiently signaling scaling parameters for cross-component intra prediction modes, leading to suboptimal decoder complexity and prediction accuracy.

Method used

A method for cross-component intra prediction that involves deriving a predicted scaling factor based on neighboring samples of a chroma block and a co-located luma block, using this factor to reconstruct the chroma block, and signaling the scaling parameters more efficiently through a reordering process.

Benefits of technology

Improves decoder efficiency and prediction accuracy by reducing the complexity of signaling scaling parameters, enhancing the overall video coding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007733125000003
    Figure 0007733125000003
  • Figure 0007733125000004
    Figure 0007733125000004
  • Figure 0007733125000005
    Figure 0007733125000005
Patent Text Reader

Abstract

A method and apparatus for performing cross-component intra prediction includes: receiving a current chroma block from a coded bitstream; determining from the coded bitstream a scaling factor to be used for the current chroma block in a chroma for luma (CfL) intra prediction mode; deriving a predicted scaling factor based on a first nearby sample of the current chroma block and a second nearby sample of a luma block co-located with the current chroma block; using the predicted scaling factor as the scaling factor to be used for the current chroma block in the CfL intra prediction mode; and reconstructing the current chroma block after scaling the current chroma block based on the predicted scaling factor.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority from U.S. Provisional Application No. 63 / 349,472, filed June 6, 2022, and U.S. Application No. 17 / 982,967, filed November 8, 2022, in the United States Patent and Trademark Office, the entire disclosures of which are incorporated herein by reference.

[0002]

[0002] Technical field Embodiments of the present disclosure relate to a family of advanced video coding techniques, and more particularly to signaling scaling parameters for cross-component intra prediction modes. [Background technology]

[0003] AOMedia Video 1 (AV1) is an open video coding format designed for video transmission over the Internet. It was developed as a successor to VP9 by the Alliance for Open Media (AOMedia), a consortium founded in 2015 that includes semiconductor companies, video-on-demand providers, video content producers, software developers, and web browser vendors. Many of the components of the AV1 project were sourced from previous research efforts by Alliance members. Individual contributors have initiated experimental technology platforms for several years: Xiph's / Mozilla's Daala published its code in 2010, Google's experimental VP9 evolution project VP10 was announced on September 12, 2014, and Cisco's Thor on August 11, 2015. Building on the VP9 codebase, AV1 incorporates additional technologies, some of which were developed in these experimental formats. The first version of the AV1 reference codec, version 0.1.0, was published on April 7, 2016. The Alliance announced the release of the AV1 Bitstream Specification on March 28, 2018, along with references to software-based encoders and decoders. Validated version 1.0.0 of the specification was released on June 25, 2018. The "AV1 Bitstream & Decoding Process Specification" was released on January 8, 2019, as validated version 1.0.0 with Errata 1 of the specification. The AV1 Bitstream Specification includes a reference video codec. The "AV1 Video Bitstream & Decoding Process Specification" (Version 1.0.0 with Errata 1), Alliance for Open Media (January 8, 2019), is incorporated herein by reference in its entirety.

[0004] The High Efficiency Video Coding (HEVC) standard is being jointly developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) standards organizations. To advance the HEVC standard, these two organizations collaborate in a partnership known as the Joint Collaborative Team on Video Coding (JCT-VC). The first version of the HEVC standard was finalized in January 2013, resulting in an aligned text published by both the ITU-T and ISO / IEC. Subsequently, additional work was organized to extend the standard to support several additional application scenarios, including extended range usage with increased precision and color format support, scalable video coding, and 3D / stereo / multiview video coding. Within ISO / IEC, the HEVC standard became MPEG-H Part 2 (ISO / IEC 23008-2), and within ITU-T, it became ITU-T Recommendation H.265. The specification for the HEVC standard, “SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services - Coding of moving video,” ITU-T H.265, International Telecommunication Union (April 2015), is hereby incorporated by reference in its entirety.

[0005]

[0005] ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) published the H.265 / HEVC (High Efficiency Video Coding) standard in 2013 (Version 1), 2014 (Version 2), 2015 (Version 3), and 2016 (Version 4). Since then, the authorities have been studying the potential need for standardization of future video coding technologies that may significantly outperform HEVC in compression capabilities. In October 2017, the authorities issued a Joint Call for Proposals on Video Compression with Capability beyond HEVC (CfP). By February 15, 2018, 22 CfP responses had been submitted for standard dynamic range (SDR), 12 for high dynamic range (HDR), and 12 for the 360 ​​video category. In April 2018, all received CfP responses were evaluated at the 122nd MPEG / 10th Joint Video Exploration Team-Joint Video Expert Team (JVET) meeting. After careful evaluation, the JVET officially initiated standardization of the next-generation video coding beyond HEVC, known as Versatile Video Coding (VVC). The VVC standard specification, “Versatile Video Coding (Draft 7),” JVET-P2001-vE, Joint Video Experts Team (October 2019), is incorporated herein by reference in its entirety. Another specification for the VVC standard, “Versatile Video Coding (Draft 10)”, JVET-S2001-vE, Joint Video Experts Team (July 2020), is incorporated herein by reference in its entirety. Summary of the Invention

[0006] According to one aspect of the disclosure, a method for performing cross-component intra prediction is performed by at least one processor, the method including: receiving a current chroma block from a coded bitstream; receiving a chroma-for-luma (chroma f r o m determining from the coded bitstream a scaling factor to be used for a current chroma block in a CfL (Cal for Luma, CfL) intra prediction mode; deriving a predicted scaling factor based on a first neighboring sample of the current chroma block and a second neighboring sample of a co-located luma block with the current chroma block; using the predicted scaling factor as the scaling factor to be used for the current chroma block in the CfL intra prediction mode; and reconstructing the current chroma block after scaling the current chroma block based on the predicted scaling factor.

[0007] According to one aspect of the disclosure, a device for performing cross-component intra prediction includes at least one memory configured to store program code, and at least one processor configured to access the program code and perform operations directed by the program code, the program code including: receiving code configured to cause the at least one processor to receive a current chroma block from a coded bitstream; r o mthe coding scheme includes: a determining code configured to cause at least one processor to determine, from the coded bitstream, a scaling factor to be used for a current chroma block in a CfL (CfL-luma) intra-prediction mode; a derivation code configured to cause at least one processor to derive a predicted scaling factor based on first neighboring samples of the current chroma block and second neighboring samples of a co-located luma block; a using code configured to cause at least one processor to use the predicted scaling factor as the scaling factor to be used for the current chroma block in a CfL intra-prediction mode; and a reconstructing code configured to cause at least one processor to reconstruct the current chroma block after scaling the current chroma block based on the predicted scaling factor.

[0008] According to one aspect of the disclosure, a non-transitory computer-readable medium is provided that stores instructions, which, when executed by one or more processors of a device for performing cross-component intra prediction, cause the one or more processors to: receive a current chroma block from a coded bitstream; r o m determining from the coded bitstream a scaling factor to be used for the current chroma block in a CfL (CfL luma) intra prediction mode; deriving a predicted scaling factor based on a first neighboring sample of the current chroma block and a second neighboring sample of a co-located luma block to the current chroma block; using the predicted scaling factor as the scaling factor to be used for the current chroma block in a CfL intra prediction mode; and reconstructing the current chroma block after scaling the current chroma block based on the predicted scaling factor. [Brief explanation of the drawings]

[0009]

[0009] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Figure 1]

[0010] FIG. 1 is a schematic diagram of a simplified block diagram of a communication system according to an embodiment. [Figure 2]

[0011] FIG. 2 is a schematic illustration of a simplified block diagram of a communication system according to an embodiment. [Figure 3]

[0012] FIG. 3 is a schematic illustration of a simplified block diagram of a decoder according to an embodiment. [Figure 4]

[0013] FIG. 4 is a schematic diagram of a simplified block diagram of an encoder according to an embodiment. [Figure 5]

[0014] FIG. 5 is a diagram illustrating eight nominal angles for AV1 according to an embodiment. [Figure 6]

[0015] FIG. 6 is a diagram illustrating a current block and samples according to an embodiment. [Figure 7]

[0016] FIG. 7 is an explanatory diagram of a linear function corresponding to a chroma-for-luma mode according to an embodiment. [Figure 8]

[0017] FIG. 8 is a diagram illustrating a current block and neighboring samples according to an embodiment. [Figure 9]

[0018] FIG. 9 is a flowchart of a method for signaling scaling parameters for cross-component intra prediction according to an embodiment. [Figure 10]

[0019] FIG. 10 is a diagram of a computer system suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0010]

[0020] In this disclosure, the term "block" may be interpreted as a prediction block, a coding block, or a coding unit (CU). The term "block" here may also be used to refer to a transform block.

[0011]

[0021] In this disclosure, the term "transform set" refers to a group of transform kernel (or candidate) options. A transform set may include one or more transform kernel (or candidate) options. According to embodiments of the present disclosure, when more than one transform option is available, an index may be signaled to indicate which of the transform options in the transform set is applied to the current block.

[0012]

[0022] In this disclosure, the term "prediction mode set" refers to a group of prediction mode options. A prediction mode set may include one or more prediction mode options. According to embodiments of the present disclosure, when more than one prediction mode option is available, an index may be further signaled to indicate which of the prediction mode options in the prediction mode set is applied to the current block to perform prediction.

[0013]

[0023] In this disclosure, the term "neighboring reconstructed sample set" refers to a group of reconstructed samples from nearby previously decoded blocks or reconstructed samples within a previously decoded picture.

[0014]

[0024] In this disclosure, the term "neural network" refers to the general concept of a structure having one or more layers that processes data as described herein in connection with "Deep Learning for Video Coding." According to embodiments of the present disclosure, any neural network may be configured to implement the embodiments.

[0015]

[0025] 3 shows a simplified block diagram of a communication system (100) according to an embodiment of the present disclosure. The communication system (100) may include at least two terminals (110, 120) interconnected via a network (150). Regarding unidirectional data transmission, for example, a first terminal (110) may locally code video data for transmission to another terminal (120) via the network (150). The second terminal (120) may receive the other terminal's coded video data from the network (150), decode the coded data, and display the recovered video data. Unidirectional data transmission may be common in media serving applications, for example.

[0016]

[0026] 1 illustrates a second pair of terminals (130, 140) provided to support bidirectional transmission of coded video, such as might occur during a video conference. For the bidirectional transmission of data, each terminal (130, 140) is capable of coding locally captured video data for transmission to the other terminal over the network (150). Each terminal (130, 140) is also capable of receiving coded video data transmitted by the other terminal, decoding the coded data, and displaying the recovered video data on a local display device.

[0017]

[0027] In FIG. 1 , the terminals (110-140) may be depicted as servers, personal computers, smartphones, and / or any other type of terminal. For example, the terminals (110-140) may be laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. The network (150) represents any number of networks that carry coded video data between the terminals (110-140), including, for example, wired and / or wireless communication networks. The communication network (150) may exchange data over circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this description, the architecture and topology of the network (150) may not be important to the operation of the present disclosure, unless otherwise described herein.

[0018]

[0028] Figure 4 illustrates the placement of a video encoder and a video decoder in a streaming environment as an example application of the disclosed subject matter, which may be equally applicable to other video-enabled applications including, for example, video conferencing, digital TV broadcasting, and storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.).

[0019]

[0029] As shown in FIG. 2, the streaming system (200) may include a capture subsystem (213) that may include a video source (201) and an encoder (203). The video source (201) may be, for example, a digital camera and may be configured to create an uncompressed video sample stream (202). The uncompressed video sample stream (202) may provide a higher amount of data compared to an encoded video bitstream and may be processed by an encoder (203) coupled to the video source (201), which may be, for example, a camera. The encoder (203) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream (204) may include a smaller amount of data compared to the sample stream and may be stored on a streaming server (205) for future use. One or more streaming clients (206) can access the streaming server (205) to retrieve a video bitstream (209), which may be a copy of the encoded video bitstream (204).

[0020]

[0030] In embodiments, the streaming server (205) may also function as a Media-Aware Network Element (MANE). For example, the streaming server (205) may be configured to prune the encoded video bitstream (204) to tailor potentially divergent bitstreams to one or more of the streaming clients (206). In embodiments, a MANE may be provided separately from the streaming server (205) in the streaming system (200).

[0021]

[0031] The streaming client (206) may include a video decoder (210) and a display (212). The video decoder (210) may, for example, decode a video bitstream (209), which may be an incoming copy of the encoded video bitstream (204), and create an outgoing video sample stream (211) that can be rendered on the display (212) or another rendering device (not shown). In some streaming systems, the video bitstreams (204, 209) may be encoded according to a particular video coding / compression standard. Examples of such standards include, but are not limited to, ITU-R Recommendation H.265. A video coding standard informally known as Versatile Video Coding (VVC) is under development. Embodiments of the present disclosure may be used in the context of VVC.

[0022]

[0032] FIG. 3 illustrates an exemplary functional block diagram of a video decoder (210) attached to a display (212) according to an embodiment of the present disclosure.

[0023]

[0033] The video decoder (210) may include a channel (312), a receiver (310), a buffer memory (315), an entropy decoder / parser (320), a scaler / inverse transform unit (351), an intra prediction unit (352), a motion compensation prediction unit (353), an aggregator (355), a loop filter unit (356), a reference picture memory (357), and a current picture memory (358). In at least one embodiment, the video decoder (210) may include an integrated circuit, a series of integrated circuits, and / or other electronic circuitry. The video decoder (210) may also be partially or entirely implemented in software running on one or more CPUs with associated memory.

[0024]

[0034] In this and other embodiments, the receiver (310) can receive one or more coded video sequences to be decoded by the decoder (210), one coded video sequence at a time, where the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequences can be received from a channel (312), which can be a hardware / software link to a storage device that stores the coded video data. The receiver (310) can receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which can be transported using entities (not shown). The receiver (310) can separate the coded video sequences from other data. To address network jitter, a buffer memory (315) can be coupled between the receiver (310) and the entropy decoder / parser (320) (hereinafter referred to as the "parser"). If the receiver (310) is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory (315) may not be used or may be small. For use on best-effort packet networks such as the Internet, a buffer memory (315) may be required, which may be relatively large and may be adaptively sized.

[0025]

[0035] The video decoder (210) may include a parser (320) for reconstructing symbols (321) from the entropy-coded video sequence. These symbol categories include, for example, information used to manage the operation of the decoder (210) and potentially information for controlling a rendering device, such as a display (212), which may be coupled to the decoder as shown in FIG. 2. The control information for the rendering device(s) may be in the form of, for example, a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (320) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context effects, etc. The parser (320) can retrieve a set of subgroup parameters for at least one of the subgroups of pixels at a video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroups can include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (320) can also retrieve information from the coded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0026]

[0036] The parser (320) is capable of performing an entropy decoding / parsing operation on the video sequence received from the buffer memory (315) to create symbols (321).

[0027]

[0037] The symbol reconstruction (321) may involve several different units, depending on the type of coded video picture or part thereof (inter and intra picture, inter and intra block, etc.), and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed by the parser (320) from the coded video sequence. The flow of such subgroup control information between the parser (320) and subsequent units is not depicted for the sake of clarity.

[0028]

[0038] Beyond the functional blocks already described, the decoder (210) may be conceptually subdivided into a number of functional units, as described below. In a practical implementation operating within commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.

[0029]

[0039] One unit may be a scalar / inverse transform unit (351), which may receive quantized transform coefficients as well as control information (including which transform to use, block size, quantization factor, quantization scaling matrix, etc.) as symbols (321) from the parser (320). The scalar / inverse transform unit (351) may output blocks containing sample values ​​that may be input to an aggregator (355).

[0030]

[0040] In some cases, the output samples of the scaler / inverse transform (351) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) generates blocks of the same size and shape as the block being reconstructed using surrounding already reconstructed information fetched from the current (partially reconstructed) picture from the current picture memory (358). The aggregator (355) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).

[0031]

[0041] In other cases, the output samples of the scalar / inverse transform unit (351) may relate to an inter-coded, potentially motion-compensated, block. In such cases, the motion-compensated prediction unit (353) may access the reference picture memory (357) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (321) associated with the block, these samples may be added by the aggregator (355) to the output of the scalar / inverse transform unit (351) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (357) from which the motion-compensated prediction unit (353) fetches the prediction samples may be controlled by a motion vector. The motion vector may be available to the motion-compensated prediction unit (353) in the form of a symbol (321), which may have, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values ​​as fetched from reference picture memory (357), motion vector prediction mechanisms, etc., when sub-sample precise motion vectors are in use.

[0032]

[0042] The output samples of the aggregator (355) can be subjected to various loop filtering techniques in the loop filter unit (356). Video compression techniques can include in-loop filtering techniques controlled by parameters contained in the coded video bitstream and made available to the loop filter unit (356) as symbols (321) from the parser (320), but can also be responsive to previously reconstructed loop-filtered sample values ​​as well as to meta-information obtained during decoding of previous parts (in decoding order) of the coded picture or coded video sequence.

[0033]

[0043] The output of the loop filter unit (356) may be a sample stream that can be output to a rendering device such as a display (212) and may also be stored in a reference picture memory (357) for use in future inter-picture prediction.

[0034]

[0044] Once a given coded picture is fully reconstructed, it may be used as a reference picture for future prediction. Once a coded picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (320)), the current reference picture may become part of the reference picture memory (357), and a new current picture memory may be reallocated before beginning reconstruction of a subsequent coded picture.

[0035]

[0045] The video decoder (210) may perform decoding operations according to a given video compression technique, which may be documented in a standard such as ITU-T Rec. H.265. The coded video sequence may conform to the syntax specified by the video compression technique or standard used, in the sense that it conforms to the syntax of the video compression technique or standard as specified in the video compression technique's document or standard, specifically in a profile document therein. To comply with some video compression techniques or standards, the complexity of the coded video sequence may be within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may be further limited in some cases by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.

[0036]

[0046] In an embodiment, the receiver (310) may receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or SNR enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0037]

[0047] FIG. 4 illustrates an exemplary functional block diagram of a video encoder (203) associated with a video source (201) according to an embodiment of the present disclosure.

[0038]

[0048] The video encoder (203) may include, for example, an encoder, such as a source coder (430), a coding engine (432), a (local) decoder (433), a reference picture memory (434), a predictor (435), a transmitter (440), an entropy coder (445), a controller (450), and a channel (460).

[0039]

[0049] The encoder (203) is capable of receiving video samples from a video source (201) (not part of the encoder) that is capable of capturing video images to be coded by the encoder (203).

[0040]

[0050] The video source (201) may provide the source video sequence to be coded by the encoder (203) in the form of a digital video sample stream that may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, ...), any color space (e.g., BT.601 YCrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (201) may be a storage device capable of storing pre-prepared video. In a video conferencing system, the video source (201) may be a camera that captures local image information as a video sequence. The video data may be provided as multiple individual pictures or images that convey motion when viewed in sequence. The picture itself may be organized as a spatial array of pixels, where each pixel may contain one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples, and the following discussion will focus on samples.

[0041]

[0051] According to one embodiment, the encoder (203) is capable of coding and compressing pictures of a source video sequence into a coded video sequence (443) in real time or under any other time constraints required by the application. Imposing an appropriate coding rate is one function of the controller (450). The controller (450) may also control and be operatively coupled to other functional units, as described below. The coupling is not depicted for clarity. Parameters set by the controller (450) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art will readily identify other functions of the controller (450) as they may be relevant to a video encoder (203) optimized for a particular system design.

[0042]

[0052] Some video encoders operate in what those skilled in the art readily recognize as a "coding loop." As a simplified explanation, the coding loop can consist of a source coder (430) (e.g., responsible for generating symbols based on an input picture to be coded and reference pictures) and a (local) decoder (533) embedded in the encoder (203), which reconstructs the symbols to create sample data that a (remote) decoder will also create if the compression between the symbols and the coded video bitstream is lossless for a given video compression technology. The reconstructed sample stream can be input to a reference picture memory (434). Because decoding the symbol stream yields bit-exact results independent of the location (local or remote) of the decoder, the reference picture memory contents are also bit-exact between the local and remote encoders. In other words, the predictor of the encoder "sees" as reference picture samples exactly the same sample values ​​that the decoder would "see" if it were to use prediction during decoding. This basic principle of reference picture synchronization (i.e., failure to maintain synchronization due to channel errors results in drift) is known to those skilled in the art.

[0043]

[0053] The operation of the "local" decoder (433) may be the same as that of the "remote" decoder (210), already described in detail above in connection with Figure 3. However, because symbols are available and the encoding / decoding of the symbols into a coded video sequence by the entropy coder (445) and parser (320) may be lossless, the entropy decoding portion of the decoder (210), including the channel (312), receiver (310), buffer memory (315), and parser (320), may not be fully implemented in the local decoder (433).

[0044]

[0054] An insight that can be made at this point is that, with the exception of analysis / entropy decoding that is present in the decoder, any decoder technology may need to exist in substantially identical functional form in the corresponding encoder. For this reason, the disclosed subject matter focuses on the operation of the decoder. Descriptions of the encoder technology can be omitted, as it may be the inverse of the decoder technology that has been comprehensively described. Only in certain areas is more detailed description required and is provided below.

[0045]

[0055] As part of its operation, the source coder (430) may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as “reference frames.” In this manner, the coding engine (432) codes the differences between pixel blocks of the input frame and pixel blocks of reference frames that can be selected as predictive references for the input frame.

[0046]

[0056] The local video decoder (433) can decode the coded video data of a frame that can be designated as a reference frame based on the symbols generated by the source coder (430). The operation of the coding engine (432) can advantageously be a non-lossless process. When the coded video data can be decoded by a video decoder (not shown in FIG. 4), the reconstructed video sequence can typically be a replica of the source video sequence with some errors. The local video decoder (433) repeats the decoding process that can be performed by the video decoder with respect to the reference frame, which can cause the reconstructed reference frame to be stored in the reference picture memory (434). In this way, the encoder (203) can locally store a copy of the reconstructed reference frame that has common content as a reconstructed reference picture to be obtained by the far-end video decoder (assuming there are no transmission errors).

[0047]

[0057] The predictor (435) can perform a prediction search for the coding engine (432). That is, for a new frame to be coded, the predictor (435) can search the reference picture memory (434) for sample data (such as candidate reference pixel blocks) or predetermined metadata (such as reference picture motion vectors, block shapes, etc.) that may serve as suitable prediction references for the new picture. The predictor (435) can operate on a sample-block-pixel-block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (435), an input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (434).

[0048]

[0058] The controller (450) may manage the coding operations of the source coder (430), including, for example, setting parameters and subgroup parameters used to encode the video data.

[0049]

[0059] All outputs of the aforementioned functional units may be subjected to entropy coding in an entropy coder (445), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as Huffman coding, variable length coding, arithmetic coding, etc.

[0050]

[0060] The transmitter (440) can buffer the coded video sequence, as produced by the entropy coder (445), and prepare it for transmission over a communication channel (460), which may be a hardware or software link to a storage device that stores the coded video data. The transmitter (440) can merge the coded video data from the video coder (430) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0051]

[0061] The controller (450) can manage the operation of the encoder (203). During coding, the controller (450) can assign a particular coded picture type to each coded picture, which can affect the coding technique that can be applied to each picture. For example, a picture can sometimes be designated as an intra-picture (I-picture), a predicted picture (P-picture), or a bidirectionally predicted picture (B-picture).

[0052]

[0062] An intra picture (I-picture) can be one that can be coded and decoded without using any other frame in a sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (“IDR”) pictures. Those skilled in the art are aware of these variations of I-pictures and their respective uses and characteristics.

[0053]

[0063] A predictive picture (P picture) can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0054]

[0064] Bi-directionally predicted pictures (B-pictures) can be coded and decoded using intra- or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a block.

[0055]

[0065] A source picture is typically spatially subdivided into multiple sample blocks (e.g., 4x4, 8x8, 4x8, or 16x16 sample blocks, respectively) and can be coded block by block. Blocks can be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the respective picture. For example, blocks of an I-picture can be nonpredictively coded, or they can be predictively coded with reference to previously coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P-picture can be nonpredictively coded with spatial or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture can be nonpredictively coded with spatial or temporal prediction with reference to one or two previously coded reference pictures.

[0056]

[0066] The video encoder (203) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In this operation, the video coder (203) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. The coded video data may therefore conform to a syntax specified by the video coding technique or standard used.

[0057]

[0067] In some embodiments, the transmitter (440) can transmit additional data along with the coded video. The video coder (430) can include such data as part of the coded video sequence. The additional data can include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, Supplemental Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0058]

[0068] [Directional intra prediction in AV1] VP9 supports eight directional modes, corresponding to angles from 45 to 207 degrees. To exploit more spatial redundancy in the directional structure, AV1 expands the directional intra-modes to a finer set of angles. The original eight angles are slightly modified to create nominal angles. These eight nominal angles are named V_PRED (542), H_PRED (543), D45_PRED (544), D135_PRED (545), D113_PRED (546), D157_PRED (547), D203_PRED (548), and D67_PRED (549), and are shown in Figure 5 relative to the current block (541). For each nominal angle, there are seven finer angles, resulting in a total of 56 directional angles in AV1. The prediction angle is represented by the nominal intra-angle plus the angle delta, where the angle delta is -3 to 3 times the 3-degree step size. In AV1, eight nominal modes are first signaled along with five non-angular smooth modes. Then, if the current mode is an angular mode, an index is further signaled to indicate the angle delta relative to the corresponding nominal angle. To implement directional prediction modes in AV1 in a general way, all 56 directional intra-prediction modes in AV1 are implemented using a unified directional predictor, which projects each pixel to a reference sub-pixel location and interpolates the reference pixel with a 2-tap bilinear filter.

[0059]

[0069] [Non-directional smooth intra predictor in AV1] In AV1, there are five non-directional smooth intra prediction modes: DC, PAETH, SMOOTH, SMOOTH_V, and SMOOTH_H. For DC prediction, the average of nearby samples to the left and above is used as the predictor for the predicted block. For the PAETH predictor, the top, left, and top-left reference samples are first fetched, and then the value closest to (top + left - top-left) is set as the predictor for the predicted pixel. Figure 6 shows the locations of the top sample (554), left sample (556), and top-left sample (558) for pixel (552) in the current block (550). For SMOOTH, SMOOTH_V, and SMOOTH_H modes, the current block (550) is predicted using quadratic interpolation in the vertical or horizontal direction, or an average in both directions.

[0060]

[0070] [Chroma predicted from Luma] In addition to the above modes, Chroma from Luma (CfL) is a chroma-only intra predictor, which models chroma pixels as a linear function of the corresponding reconstructed luma pixels. CfL prediction can be expressed as shown in Equation (1) below:

[0061]

number

[0062]

[0071] Figure 7 provides a graphical illustration of the linear function described by Equation (1). As seen in Figure 7 and Equation (1), the reconstructed luma pixels are subsampled to the chroma resolution, and then the mean is subtracted to form the AC contributions. Instead of requiring the decoder to calculate scaling parameters to approximate the chroma AC components from the AC contributions, as in some background art, AV1 CfL can determine the parameter α based on the original chroma pixels and signal them in the bitstream. This reduces decoder complexity and results in more accurate predictions. The DC contributions of the chroma components may be calculated using intra-DC mode, which is sufficient for most chroma content and has mature, fast implementations.

[0063]

[0072] When the chroma-from-luma (CfL) mode is selected, the joint sign of the scaling factors for the U and V components may be signaled first. The sign of a scaling factor may be negative, zero, or positive. Furthermore, the (0,0) combination may be disallowed in the CfL mode because it results in a "DC" prediction. Therefore, there are a total of eight (3*3-1=8) possible combinations of signs for the two scaling factors. As a result, the joint sign may be signaled using an octal symbol. Only one context may be employed to signal the joint sign.

[0064]

[0073] For signaling the magnitude of the scaling parameter, 16-ary symbols may be used to represent values ​​ranging from 0 to 2 in steps of 1 / 8. Note that 16-ary symbols can fully utilize the capabilities of a multi-symbol entropy encoder. The context for signaling the scaling parameter may depend on the value of the joint code.

[0065]

[0074] There may be a strong correlation between neighboring samples and samples in the current block, but this correlation is not exploited in the signaling of the joint sign and magnitude of the scaling parameters in CfL mode.

[0066]

[0075] Therefore, embodiments provide an improved method for signaling scaling parameters in CfL mode. The embodiments described herein may also be applied to other modes, such as modes similar to CfL mode, but in which luma is replaced by one specific color component (e.g., R) and chroma is replaced by another specific color component (e.g., G or B).

[0067]

[0076] In an embodiment, a linear model using a scaling parameter alpha_predict may be derived based on nearby samples of the current chroma block and nearby samples of the co-located luma block. This scaling parameter alpha_predict may be used to signal the actual scaling factor in CfL mode for the current block. The scaling parameter alpha_predict may also be referred to as the predicted scaling parameter for the current chroma block.

[0068]

[0077] In an embodiment, neighboring samples above and / or to the left of the current chroma block and neighboring samples above and / or to the left of the co-located luma block may be included in the process of deriving the scaling parameter alpha_predict. An example is shown in Figure 8, where the neighboring samples involved are shown as dotted squares and the current block samples are shown as white squares. As another example, if the YUV format of the current frame / video sequence is not YUV444, the neighboring samples of the co-located luma block may be downsampled before deriving the predicted scaling coefficients in the linear model.

[0069]

[0078] In some embodiments, a least mean squares method may be employed to derive the predicted scaling parameters of the linear model, for example, the same method used in the cross-component linear model (CCLM) mode defined in the VVC standard.

[0070]

[0079] In an embodiment, the predicted scaling parameters may be derived separately for the U and V components. For the U component of the current chroma block, nearby samples of the U component of the current chroma block and nearby samples of the co-located luma block may be used to derive predicted scaling parameters. For the V component of the current chroma block, nearby samples of the V component of the current chroma block and nearby samples of the co-located luma block may be used to derive predicted scaling parameters.

[0071]

[0080] In an embodiment, the predicted scaling parameters may be used to reorder the available sign values ​​of the scaling coefficients in the CfL mode. The index of the re-ordered sign value may be signaled in the bitstream. An example of the reordering of the available sign values ​​is shown in Table 1. In this example, the sign value can be either negative (-1), positive (+1), or zero (0). If scale_predict is greater than 0, the available sign values ​​may be reordered as positive (+1), or zero (0) negative (-1). Thus, if the sign value in the CfL mode for the current chroma block is negative, 2 may be signaled in the bitstream because the index for negative (-1) in Table 1 is 2 when scale_predict is greater than 0.

[0072] Table 1: Examples of permutation code values

[0073] [Table 1]

[0081] In an embodiment, the predicted scaling parameters may be used to reorder the magnitudes (absolute values) of the scaling coefficients in CfL mode. The indices of the reordered magnitude values ​​(absolute values) of the scaling coefficients are signaled in the bitstream. In an embodiment, the absolute magnitudes of the scaling parameters may be reordered based on the difference between the absolute values ​​of the predicted scaling parameters and the absolute values ​​of the available scaling factors. For example, if the absolute value of the predicted scaling parameter is 3 / 8, the available scaling factors may be reordered as follows: (3 / 8,2 / 8,4 / 8,1 / 8,5 / 8,6 / 8,7 / 8,1,9 / 8,10 / 8,11 / 8,12 / 8,13 / 8,14 / 8,15 / 8,16 / 8)

[0082] In an embodiment, to signal the value of a scaling factor in CfL mode, a first flag, which may be referred to as 0_flag, may be signaled in the bitstream to indicate whether the value of the scaling factor is equal to 0. If the scaling factor is not equal to 0, the predicted scaling factor may be used to permute all available scaling factors, including positive and negative scaling factors, except for 0. An index within the set of permuted scaling factors may be signaled.

[0074]

[0083] In an embodiment, the predicted scaling factor may be used to sort all available scaling factors based on the difference between the predicted scaling factor and the available scaling factors. For example, if the predicted scaling factor is 5 / 8, all available scaling factors may be sorted as follows: (5 / 8,4 / 8,6 / 8,3 / 8,7 / 8,2 / 8,8 / 8,1 / 8, 9 / 8,-1 / 8,10 / 8,-2 / 8,11 / 8,-3 / 8,12 / 8,-4 / 8, 13 / 8,-5 / 8,14 / 8,-6 / 8,15 / 8,-7 / 8,16 / 8,-8 / 8, -9 / 8,-10 / 8,-11 / 8,-12 / 8,-13 / 8,-14 / 8,-15 / 8,-16 / 8) The first 16 scaling factors may be put into the first set, and the last 16 scaling factors may be put into the second set. For example, if the scaling factor of the CfL mode for the current block is 3 / 8, the set index is 0 and the index within the set is 3.

[0075]

[0084] In an embodiment, the reordered scaling factors may be divided into multiple sets based on the indices of the reordered scaling factors, and the set indices and indices within the sets may be signaled in the bitstream.

[0076]

[0085] In an embodiment, to signal the values ​​of the scaling factors in CfL mode, the predicted scaling factors may be used to sort all available scaling factors (including 0), and the index of the selected scaling factor among the sorted set of scaling factors may be signaled.

[0077]

[0086] In an embodiment, to signal the values ​​of the scaling factors in CfL mode, the predicted scaling factors may be used to select a subset of scaling factors from the full set of available scaling factors, and the selection of the scaling factors in the subset of scaling factors may be signaled.

[0078]

[0087] In an embodiment, a first set of scaling factors having a first precision may be defined, the predicted scaling factors may be used to select a subset of the first set of scaling factors, then a second set of scaling factors having a second precision may be further selected, and then the selected subset of the first set of scaling factors together with the second set of scaling factors forms a third set of scaling factors, which may include scaling factors with variable precision. For example, the first precision may include, but is not limited to, ¼, ⅛, and 1 / 16 precision. For example, the second precision may include, but is not limited to, ⅛, 1 / 16, 1 / 32, 1 / 64, and 1 / 128 precision. In an embodiment, the second precision may be higher than the first precision.

[0079]

[0088] FIG. 9 is a flowchart of a process 1000 for performing cross-component intra prediction according to an embodiment.

[0080]

[0089] As shown in FIG. 9, in operation 1002, the process 1000 includes receiving a current chroma block from a coded bitstream.

[0081]

[0090] As further shown in FIG. 9, in operation 1004, the process 1000 calculates CfL(chroma f r o m This includes determining from the coded bitstream the scaling factor to be used for the current chroma block in intra prediction modes (i.e., luma).

[0082]

[0091] As further shown in FIG. 9, in operation 1006, the process 1000 includes deriving a predicted scaling factor based on a first nearby sample of the current chroma block and a second nearby sample of a luma block co-located with the current chroma block.

[0083]

[0092] As further shown in FIG. 9, in operation 1008, the process 1000 includes using the predicted scaling factor as the scaling factor to be used for the current chroma block in a CfL intra prediction mode.

[0084]

[0093] As further shown in FIG. 9, in operation 1010, the process 1000 includes reconstructing the current chroma block after scaling the current chroma block based on the predicted scaling factor.

[0085]

[0094] In an embodiment, the predicted scaling factor may be derived based on at least one chroma sample located above or to the left of the current chroma block and at least one luma sample located above or to the left of the current luma block.

[0086]

[0095] In an embodiment, the predicted scaling factor may be derived using a least mean squares operation.

[0087]

[0096] In an embodiment, the predicted scaling factors may be derived separately for the first and second chroma components.

[0088]

[0097] In an embodiment, the predicted scaling factor may be included in a plurality of scaling factors; a scaling parameter is used to order at least one of the signs available for the plurality of scaling factors and the magnitudes of the absolute values ​​corresponding to the scaling factors.

[0089]

[0098] In an embodiment, a scaling parameter may be used to order the multiple scaling factors excluding those toward 0 based on a flag indicating that the scaling factor is not zero; a predicted scaling factor may be selected from among the multiple scaling factors.

[0090]

[0099] In an embodiment, the predicted scaling factor is included in a plurality of scaling factors; a scaling parameter may be used to order the plurality of scaling factors; and an index may be used to select the predicted scaling factor from among the plurality of scaling factors.

[0091]

[0100] In embodiments, a scaling parameter may be used to select a subset of scaling factors from among the plurality of scaling factors; a predicted scaling factor may be selected from among the subset of scaling factors based on syntax elements signaled in the coded bitstream.

[0092]

[0101] 9 illustrates example blocks of process 1000, in some implementations process 1000 may include additional, fewer, different, or differently arranged blocks relative to the blocks illustrated in FIG 9. Additionally or alternatively, two or more blocks of process 1000 may be performed in parallel.

[0093]

[0102] Additionally, the proposed methods may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.

[0094]

[0103] The techniques of the disclosed embodiments described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 10 illustrates a computer system (900) suitable for implementing embodiments of the disclosed subject matter.

[0095]

[0104] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code that contains instructions that may be executed directly by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or that may go through interpretation, microcode execution, etc.

[0096]

[0105] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0097]

[0106] 10 for computer system 900 are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components shown in the exemplary embodiment of computer system 900.

[0098]

[0107] The computer system (900) may include certain human interface input devices that may respond to input by one or more human users, for example, via tactile input (e.g., keystrokes, swipes, data glove movements), auditory input (e.g., voice, claps), visual input (e.g., gestures), or olfactory input (not shown). Human interface devices may also be used to capture certain media that do not necessarily involve direct human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still-image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic pictures).

[0099]

[0108] Input human interface devices may include one or more of the following (only one of each is depicted): a keyboard (901), a mouse (902), a trackpad (903), a touch screen (910), a data glove, a joystick (905), a microphone (906), a scanner (907), and a camera (908).

[0100]

[0109] The computer system (900) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (910), data gloves, or joystick (905), although there may also be haptic feedback devices that do not function as input devices). For example, such devices may be auditory output devices (e.g., speakers (909), headphones (not shown)), visual output devices (e.g., screens (910) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capability, each with or without haptic feedback capability, some of which may be capable of outputting two-dimensional visual output, three- or more-dimensional output by means such as stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0101]

[0110] The computer system (900) may also include human-accessible storage devices and their associated media, such as optical media including CDs / DVDs, ROM / RW (920) using media such as CDs / DVDs (2021), thumb drives (922), removable hard drives or solid state drives (923), legacy magnetic media (not shown) such as tapes and floppy disks (not shown), and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0102]

[0111] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not encompass transmission media, carrier waves, or other transitional signals.

[0103]

[0112] The computer system 900 may also include interfaces to one or more communication networks. Networks may be, for example, wireless, wired, or optical. Networks may further be local, wide area, metropolitan, vehicular, industrial, real-time, delay-tolerant, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks (including GSM, 3G, 4G, 5G, LTE, etc.), TV wired or wireless wide area digital networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicular and industrial networks including CANBus, etc. Certain networks generally require an external network interface adapter (e.g., a USB port on the computer system 900) attached to a particular general-purpose data port or peripheral bus (949); others are generally integrated into the core of the computer system 900 by attaching to the system bus, as described below (e.g., an Ethernet interface is integrated into a PC computer system, and a cellular network interface is integrated into a smartphone computer system). Using any of these networks, the computer system 900 can communicate with other entities. Such communications can be one-way receive-only (e.g., broadcast TV), one-way transmit-only (e.g., from a CANbus to a particular CANbus device), or bidirectional, such as to other computer systems using local or wide-area digital networks. Such communications can include communications to a cloud computing environment (955). Predetermined protocols and protocol stacks can be used for each of these networks and network interfaces, as described above.

[0104]

[0113] The aforementioned human interface devices, human-accessible storage devices, and network interfaces (954) may be attached to the core (940) of the computer system (900).

[0105]

[0114] A core (940) may include one or more central processing units (CPUs) (941), graphics processing units (GPUs) (942), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (943), task-specific hardware accelerators (944), etc. These devices may be connected via a system bus (948) to read-only memory (ROM) (945), random access memory (946), and internal mass storage such as an internal non-user-accessible hard drive, SSD, etc. (947). In some computer systems, the system bus (948) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (948) or via a peripheral bus (949). Architectures for peripheral buses include PCI, USB, etc. A graphics adapter (950) may be included in the core (940).

[0106]

[0115] The CPU (941), GPU (942), FPGA (943), and accelerator (944) can execute specific instructions that, in combination, may constitute the aforementioned computer code. This computer code may be stored in ROM (945) or RAM (946). Temporary data may be stored in RAM (946), while persistent data may be stored, for example, in internal mass storage (947). A cache memory, which may be closely associated with one or more of the CPU (941), GPU (942), mass storage (947), ROM (945), RAM (946), etc., may be used to enable fast storage and retrieval of data from any memory device.

[0107]

[0116] The computer-readable medium can contain computer code for performing various computer-implemented operations, and the medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.

[0108]

[0117] By way of example and not limitation, the architecture corresponding to the computer system (900), and in particular the core (940), may provide functionality as a result of processor(s) (including CPUs, GPUs, FPGAs, accelerators, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may include user-accessible mass storage, as discussed above, as well as media associated with the core's (940) specific storage of a non-transitory nature, such as the core's internal mass storage (947) or ROM (945). Software implementing various embodiments of the present disclosure may be stored within such devices and executed by the core (940). The computer-readable media may include one or more memory devices or chips, depending on the particular needs. Software can cause the cores (940) and specifically the processors therein (including CPUs, GPUs, FPGAs, etc.) to execute the particular processes or portions of the particular processes described herein (including defining data structures stored in RAM (946) and modifying such data structures according to the software-defined processes). Additionally or alternatively, the computer system can provide functionality as a result of hardwired logic or logic otherwise embedded in circuitry (e.g., accelerators (944)), which can operate in place of or in conjunction with software to execute the particular processes or portions of the particular processes described herein. References to software can encompass logic, where appropriate, and vice versa. References to computer-readable media can encompass circuitry (e.g., integrated circuits (ICs)) that stores software for execution, circuitry that embodies logic for execution, or both, where appropriate. The present disclosure encompasses any appropriate combination of hardware and software.

[0109]

[0118] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the embodiments (and methods thereof) may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium.

[0110]

[0119] The above disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementation to the precise form disclosed. Modifications and variations are possible in light of the above disclosure or may be learned from practice of the implementation.

[0111]

[0120] As used herein, the term component is intended to be broadly interpreted as hardware, firmware, or a combination of hardware and software.

[0112]

[0121] Although combinations of features may be recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. Indeed, many of these features may be combined in ways not specifically recited in the claims and / or specifically disclosed in the specification. Although each dependent claim listed below may depend directly on only one claim, the disclosure of possible implementations includes each dependent claim in combination with all other claims in the claim set.

[0113]

[0122] No element, act, or instruction used herein should be construed as critical or essential unless expressly described as such. Also, as used herein, the article words "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Furthermore, the term "set" is intended to include one or more items (e.g., related items, unrelated items, or a combination of related and unrelated items) and may be used interchangeably with "one or more." Where only one item is intended, "one" or similar language is used. Also, as used herein, the terms "has," "have," "having," etc. are intended to be open-ended terms. Furthermore, the phrase "based on" is intended to mean "based at least in part on," unless expressly stated otherwise.

[0114]

[0123] While this disclosure describes several non-limiting exemplary embodiments, there are alterations, substitutions, and various substitute equivalents that fall within the scope of this disclosure. It will thus be recognized that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore are within the spirit and scope thereof.

[0115]

[0124] Additional notes (Appendix 1) 1. A method executed by at least one processor for performing cross-component intra prediction, the method comprising: receiving a current chroma block from the coded bitstream; CfL(chroma f r o m determining from the coded bitstream a scaling factor to be used for the current chroma block in an intra prediction mode (luma); deriving a predicted scaling factor based on first neighboring samples of the current chroma block and second neighboring samples of a co-located luma block; using the predicted scaling factor as the scaling factor to be used for the current chroma block in the CfL intra prediction mode; and reconstructing the current chroma block after scaling the current chroma block based on the predicted scaling factor; A method comprising:

[0116] (Appendix 2) 2. The method of claim 1, wherein the predicted scaling factor is derived based on at least one chroma sample located above or to the left of the current chroma block and at least one luma sample located above or to the left of the current luma block.

[0117] (Appendix 3) 2. The method of claim 1, wherein the predicted scaling factor is derived using a least mean squares operation.

[0118] (Appendix 4) 2. The method of claim 1, wherein the predicted scaling factors are derived separately for a first chroma component and a second chroma component. (Appendix 5) 2. The method of claim 1, wherein the predicted scaling factor is included in a plurality of scaling factors; A method wherein a scaling parameter is used to order at least one of the available sign values ​​for the plurality of scaling factors and the magnitudes of the absolute values ​​corresponding to the scaling factors.

[0119] (Appendix 6) 2. The method of claim 1, wherein a scaling parameter is used to order the plurality of scaling factors excluding those toward 0 based on a flag indicating that the scaling factor is not zero; The method, wherein the predicted scaling factor is selected from among the plurality of scaling factors.

[0120] (Appendix 7) 2. The method of claim 1, wherein the predicted scaling factor is included in a plurality of scaling factors; a scaling parameter is used to order the plurality of scaling factors; and A method wherein an index is used to select the predicted scaling factor from among the plurality of scaling factors.

[0121] (Appendix 8) 2. The method of claim 1, wherein a scaling parameter is used to select a subset of scaling factors from a plurality of scaling factors; The method of claim 1, wherein the predicted scaling factor is selected from among the subset of scaling factors based on syntax elements signaled in the coded bitstream.

[0122] (Appendix 9) 1. A device for performing cross-component intra prediction, the device comprising: at least one memory configured to store program code; and at least one processor configured to access said program code and to perform operations directed by said program code; the program code comprising: receiving code configured to cause the at least one processor to receive a current chroma block from a coded bitstream; CfL(chroma f r o mdetermining, from the coded bitstream, a scaling factor to be used for the current chroma block in a (luma) intra prediction mode; derivation code configured to cause the at least one processor to derive a predicted scaling factor based on first neighboring samples of the current chroma block and second neighboring samples of a co-located luma block; using code configured to cause the at least one processor to use the predicted scaling factor as a scaling factor to be used for the current chroma block in the CfL intra prediction mode; and reconstruction code configured to cause the at least one processor to reconstruct the current chroma block after scaling the current chroma block based on the predicted scaling factor; Including, the device.

[0123] (Appendix 10) 10. The device of claim 9, wherein the predicted scaling factor is derived based on at least one chroma sample located above or to the left of the current chroma block and at least one luma sample located above or to the left of the current luma block.

[0124] (Appendix 11) 10. The device of claim 9, wherein the predicted scaling factor is derived using a least mean squares operation.

[0125] (Appendix 12) 10. The device of claim 9, wherein the predicted scaling factors are derived separately for a first chroma component and a second chroma component.

[0126] (Appendix 13) 10. The device of claim 9, wherein the predicted scaling factor is included in a plurality of scaling factors; A device, wherein a scaling parameter is used to order at least one of the available code values ​​for the plurality of scaling factors and the magnitude of the absolute value corresponding to the scaling factor.

[0127] (Appendix 14) 10. The device of claim 9, wherein a scaling parameter is used to order the plurality of scaling factors excluding those toward 0 based on a flag indicating that the scaling factor is not zero; The predicted scaling factor is selected from among the plurality of scaling factors.

[0128] (Appendix 15) 10. The device of claim 9, wherein the predicted scaling factor is included in a plurality of scaling factors; a scaling parameter is used to order the plurality of scaling factors; and The device, wherein an index is used to select the predicted scaling factor from among the plurality of scaling factors.

[0129] (Appendix 16) 10. The device of claim 9, wherein a scaling parameter is used to select a subset of scaling factors from the plurality of scaling factors; The device, wherein the predicted scaling factor is selected from among the subset of scaling factors based on syntax elements signaled in the coded bitstream.

[0130] (Appendix 17) 1. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a device for performing cross-component intra prediction, cause the one or more processors to: receiving a current chroma block from the coded bitstream; CfL(chroma f r o m determining from the coded bitstream a scaling factor to be used for the current chroma block in an intra prediction mode (luma); deriving a predicted scaling factor based on first neighboring samples of the current chroma block and second neighboring samples of a co-located luma block; using the predicted scaling factor as the scaling factor to be used for the current chroma block in the CfL intra prediction mode; and reconstructing the current chroma block after scaling the current chroma block based on the predicted scaling factor; A medium that allows this to be done.

[0131] (Appendix 18) 18. The non-transitory computer-readable medium of Claim 17, wherein the predicted scaling factor is derived based on at least one chroma sample located above or to the left of the current chroma block and at least one luma sample located above or to the left of the current luma block.

[0132] (Appendix 19) 18. The non-transitory computer-readable medium of claim 17, wherein the predicted scaling factor is derived using a least mean squares operation.

[0133] (Appendix 20) 18. The non-transitory computer-readable medium of claim 17, wherein the predicted scaling factors are derived separately for a first chroma component and a second chroma component.

Claims

1. 1. A method executed by at least one processor for performing cross-component intra prediction, the method comprising: receiving a current chroma block from a coded bitstream; a determining step of determining a scaling factor to be used for the current chroma block in a chroma from luma (CfL) intra prediction mode, deriving a predicted scaling factor based on first neighboring samples of the current chroma block and second neighboring samples of a co-located luma block; and a determining step including using the predicted scaling factor to determine, from among a plurality of available scaling factors, a scaling factor to be used for the current chroma block in the CfL intra prediction mode; and deriving chroma samples of the current chroma block by using the scaling factor to be used determined in the determining step in the CfL intra prediction mode; A method comprising:

2. 2. The method of claim 1, wherein the predicted scaling factor is derived based on at least one chroma sample located above or to the left of the current chroma block and at least one luma sample located above or to the left of the current luma block.

3. 10. The method of claim 1, wherein the predicted scaling factor is derived using a least mean squares operation.

4. The method of claim 1 , wherein the predicted scaling factors are derived separately for a first chroma component and a second chroma component.

5. 2. The method of claim 1, wherein the predicted scaling factor is used to order at least one of the available code values ​​for the plurality of scaling factors and the magnitudes of the absolute values ​​corresponding to the available plurality of scaling factors.

6. 2. The method of claim 1, wherein the predicted scaling factor is used to order a plurality of scaling factors, excluding those toward 0, based on a flag indicating that the scaling factor to be used is not 0; A method wherein the scaling factor used is selected from among the plurality of scaling factors.

7. 10. The method of claim 1, wherein the predicted scaling factor is used to order the plurality of scaling factors; and A method wherein an index is used to select the scaling factor to be used from among the plurality of scaling factors.

8. 10. The method of claim 1, wherein the predicted scaling factors are used to select a subset of scaling factors from the plurality of scaling factors; The method of claim 1, wherein the scaling factor to be used is selected from among the subset of scaling factors based on syntax elements signaled in the coded bitstream.

9. 1. A device for performing cross-component intra prediction, comprising: at least one memory configured to store program code; and at least one processor configured to access said program code and to perform operations directed by said program code; 9. A device comprising: a processor configured to execute a method according to any one of claims 1 to 8;

10. A computer program product causing a processor of a device performing cross-component intra prediction to perform the method of any one of claims 1 to 8.