Cross-component prediction

By introducing syntax elements into video encoding technology to control the weighted average of multiple chromaticity prediction blocks, combining inter-frame and cross-component prediction modes, the problem of inefficient cross-component prediction in the prior art is solved, and more efficient and accurate video encoding is achieved.

CN120153649APending Publication Date: 2025-06-13TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004670.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-24
Filing Date
2024-06-04
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing video encoding technology has problems with inefficiency in cross-component prediction, especially when dealing with weighted averages of multiple chromaticity prediction blocks, the determination of weights is not flexible enough, which affects the encoding effect.

Method used

By introducing a first syntax element into the code stream, it indicates whether to predict the chromaticity block by weighted average of multiple chromaticity prediction blocks, and determine the prediction block of the chromaticity block according to the inter prediction mode and the cross-component prediction mode, and the prediction accuracy of the chromaticity block is improved by weighted average.

Benefits of technology

It improves the efficiency and accuracy of video encoding, and through flexible weight settings, it adapts to the needs of different encoding modes, and improves the compression effect of the code stream.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120153649A_ABST
    Figure CN120153649A_ABST
Patent Text Reader

Abstract

In one method, a first chroma prediction block of a chroma block of a current block is determined based on an inter prediction mode. A second chroma prediction block of the chroma block of the current block is determined based on a cross-component prediction mode in which the second chroma prediction block is obtained based on reconstructed luma samples of the luma block of the current block. A prediction block of the chroma block is determined as a weighted average of a first chroma prediction block of the plurality of chroma prediction blocks and a second chroma prediction block of the plurality of chroma prediction blocks. A first syntax element is encoded into the code stream, the first syntax element indicating that a chroma block of a current block is predicted by a weighted average of a plurality of chroma prediction blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by Reference

[0001] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 528,589, filed on Jul. 24, 2023, entitled "On Improvement of Cross-Component Prediction", the entire content of which is incorporated herein by reference. Technical Field

[0002] The present disclosure describes aspects generally related to video coding. Background Art

[0003] The background description provided herein is intended to present the background of the present disclosure as a whole. The extent to which the work of the presently named inventors, described in the background art section and in various aspects of this specification, was carried out does not indicate that it was prior art at the time of filing of the present disclosure, and has never been expressly or implicitly admitted to be prior art of the present disclosure.

[0004] Image / video compression can help transmit image / video data between different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial redundancy and temporal redundancy. In one example, a video codec can use a technique called intra prediction, which can compress an image based on spatial redundancy. For example, intra prediction can use reference data from the currently reconstructed picture to perform sample prediction. In another example, a video codec can use a technique called inter prediction, which can compress an image based on temporal redundancy. For example, inter prediction can utilize motion compensation to predict samples in the current picture from a previously reconstructed picture. Motion compensation can be represented by a motion vector (MV). Summary of the Invention

[0005] Aspects of the present disclosure provide a bitstream, methods and apparatuses for video encoding / decoding. In some examples, the apparatus for video encoding / decoding includes a processing circuit.

[0006] According to one aspect of the present disclosure, a method for processing visual media data is provided. In this method, a bitstream of visual media data is processed according to formatting rules. In one example, the bitstream includes a first syntax element associated with a current block in a current picture, the current block including a chrominance block and a luminance block, and the first syntax element indicates whether the chrominance block is predicted by a weighted average of a plurality of chrominance prediction blocks. The formatting rules specify that when the first syntax element indicates that the chrominance block is predicted by a weighted average of a plurality of chrominance prediction blocks, a first chrominance prediction block of the plurality of chrominance prediction blocks is determined based on an inter-frame prediction mode. A second chrominance prediction block of the plurality of chrominance prediction blocks is determined based on a cross-component prediction mode, in which the second chrominance prediction block is obtained based on reconstructed luminance samples of a luminance block filtered according to filter coefficients of a filter. The filter coefficients of the filter are based on (i) chrominance prediction of the chrominance block and luminance prediction of the luminance block or (ii) merge candidates in a merge list encoded in the cross-component prediction mode. The formatting rules specify that a prediction block of the chrominance block is determined as a weighted average of the first chrominance prediction block and the second chrominance prediction block.

[0007] In one example, when both the upper adjacent block and the left adjacent block of the current block are encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.75. When at least one of the upper adjacent block and the left adjacent block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.5. When neither the upper adjacent block nor the left adjacent block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.25.

[0008] In one example, the formatting rules specify that a merge list is constructed based on a plurality of cross-component prediction encoded blocks, the plurality of cross-component prediction encoded blocks including at least one of (i) adjacent spatially adjacent blocks, (ii) non-adjacent spatially adjacent blocks, (iii) history-based adjacent blocks, (iv) temporally collocated blocks, and (v) temporally shifted blocks in a reference picture of the current picture.

[0009] According to another aspect of the present disclosure, a method for video coding is provided. In this method, a first chrominance prediction block of a chrominance block of a current block is determined based on an inter-frame prediction mode. A second chrominance prediction block of the chrominance block of the current block is determined based on a cross-component prediction mode, in which the second chrominance prediction block is obtained based on reconstructed luminance samples of a luminance block of the current block. A prediction block of the chrominance block is determined as a weighted average of the first chrominance prediction block of the plurality of chrominance prediction blocks and the second chrominance prediction block of the plurality of chrominance prediction blocks. A first syntax element is encoded into the bitstream, the first syntax element indicating that the chrominance block of the current block is predicted by a weighted average of a plurality of chrominance prediction blocks.

[0010] In one example, the cross-component prediction mode includes a first cross-component prediction mode in which a second chrominance prediction block is obtained based on reconstructed luminance samples of a luminance block filtered according to filter coefficients of a first filter, and the filter coefficients of the first filter are obtained based on a chrominance prediction of a chrominance block and a luminance prediction of the luminance block. The cross-component prediction mode includes a second cross-component prediction mode in which a second chrominance prediction block is obtained based on reconstructed luminance samples of a luminance block filtered by filter coefficients of a second filter. The filter coefficients of the second filter are obtained based on merge candidates in a merge list.

[0011] In one example, when encoding both the upper adjacent block and the left adjacent block of a current block in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.75. When encoding at least one of the upper adjacent block and the left adjacent block of the current block in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.5. When neither the upper adjacent block nor the left adjacent block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.25.

[0012] According to another aspect of the present disclosure, there is provided an apparatus for video decoding. The apparatus includes a processing circuit. The processing circuit is configured to receive a bitstream that includes a first syntax element associated with a current block in a current picture. The current block includes a chrominance block and a luminance block. The first syntax element indicates whether to predict the chrominance block by a weighted average of a plurality of chrominance prediction blocks. When the first syntax element indicates predicting the chrominance block by a weighted average of a plurality of chrominance prediction blocks, the processing circuit is configured to (i) determine a first chrominance prediction block of the plurality of chrominance prediction blocks based on an inter prediction mode, and (ii) determine a second chrominance prediction block of the plurality of chrominance prediction blocks based on a cross-component prediction mode in which the second chrominance prediction block is obtained based on reconstructed luminance samples of the luminance block. The processing circuit is configured to determine a prediction block of the chrominance block as a weighted average of the first chrominance prediction block and the second chrominance prediction block.

[0013] In one example, the cross-component prediction mode includes a first cross-component prediction mode in which a second chrominance prediction block is obtained based on reconstructed luminance samples of a luminance block filtered according to filter coefficients of a first filter. The filter coefficients of the first filter are obtained based on a chrominance prediction of a chrominance block and a luminance prediction of the luminance block. The cross-component prediction mode includes a second cross-component prediction mode in which a second chrominance prediction block is obtained based on reconstructed luminance samples of a luminance block filtered by filter coefficients of a second filter, and the filter coefficients of the second filter are obtained based on merge candidates in a merge list.

[0014] In one example, when applying a cross-component prediction mode to a current block, a first syntax element is written in the bitstream and signaled.

[0015] In one example, the processing circuitry is configured to determine a predicted block of a luminance block by copying an inter-predicted luminance block when the first syntax element indicates that a chrominance block is predicted by a weighted average of multiple chrominance prediction blocks and the cross-component prediction mode is applied to the current block. In one example, the processing circuitry is configured to receive a flag in the bitstream. The flag indicates whether a weighted average of multiple luminance prediction blocks is used to predict the luminance block.

[0016] In one example, when encoding both the upper neighboring block and the left neighboring block of a current block in the cross-component prediction mode, the weight of a second chrominance prediction block in the weighted average is 0.75. When encoding one of the upper neighboring block and the left neighboring block of the current block in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.5. When neither the upper neighboring block nor the left neighboring block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.25.

[0017] In one example, the processing circuitry is configured to construct a merge list based on multiple cross-component prediction encoded blocks, the multiple cross-component prediction encoded blocks including at least one of (i) adjacent spatially neighboring blocks, (ii) non-adjacent spatially neighboring blocks, (iii) history-based neighboring blocks, (iv) temporally collocated blocks, and (v) temporally shifted blocks in a reference picture of a current picture. The merge list is constructed based on a predefined scan order of adjacent spatially neighboring blocks, history-based neighboring blocks, non-adjacent spatially neighboring blocks, and temporally collocated blocks from the reference picture.

[0018] In one example, when a second syntax element in the bitstream indicates that the merge list is used for the cross-component prediction mode, the processing circuitry is configured to determine a merge candidate from the merge list indicated by an index in the bitstream.

[0019] In one example, when the second syntax element indicates that a second cross-component prediction mode is not applied, the processing circuitry is configured to (i) obtain filter coefficients of a first filter based on chrominance prediction of a chrominance block and luminance prediction of a luminance block, or (ii) determine the filter coefficients of the first filter according to information signaled in the bitstream.

[0020] In one example, the processing circuitry is configured to divide the merge list into multiple subgroups. The processing circuitry is configured to determine a merge candidate from the merge list indicated by an index in the bitstream, the index including a first part and a second part, the first part indicating which one of the multiple subgroups to select, and the second part indicating which one of the merge candidates to select from the selected subgroup.

[0021] Aspects of the present disclosure also provide an apparatus for video encoding. The apparatus for video encoding includes processing circuitry configured to implement any of the methods for video encoding described herein.

[0022] Aspects of the present disclosure also provide a method for video decoding. The method includes any of the methods implemented by an apparatus for video decoding.

[0023] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to execute any of the methods for video decoding / encoding described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0025] Figure 1 is a schematic diagram of an example of a block diagram of a communication system (100).

[0026] Figure 2 is a schematic diagram of an example of a block diagram of a decoder.

[0027] Figure 3 is a schematic diagram of an example of a block diagram of an encoder.

[0028] Figure 4 is a schematic diagram of weighted derivation based on neighboring blocks.

[0029] Figure 5 is a schematic diagram of cross-component prediction without mixing with chrominance prediction values.

[0030] Figure 6 is a schematic diagram of cross-component prediction without mixing with chrominance prediction values.

[0031] Figure 7 shows a flowchart outlining a decoding method according to some aspects of the present disclosure.

[0032] Figure 8 shows a flowchart outlining an encoding method according to some aspects of the present disclosure.

[0033] Figure 9 is a schematic diagram of a computer system according to one aspect. DETAILED DESCRIPTION

[0034] Figure 1A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application for the disclosed subject matter, video encoders, and video decoders in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, Digital Video Discs (DVDs), memory sticks, etc.

[0035] The video processing system (100) includes an acquisition subsystem (113), and the acquisition subsystem may include a video source (101) such as a digital camera that creates, for example, an uncompressed video picture stream (102). In one example, the video picture stream (102) includes samples taken by the digital camera. The video picture stream (102) is depicted as a thick line to emphasize the high data volume of the video picture stream compared to the encoded video data (104) (or encoded video bitstream). The video picture stream (102) may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (104) (or encoded video bitstream (104)), depicted as a thin line to emphasize the lower data volume compared to the video picture stream (102), may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 the client subsystem (106) and the client subsystem (108) in, may access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes an incoming copy (107) of the encoded video data and produces an output video picture stream (111) that can be presented on a display (112) (such as a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), video data (107), and video data (109) (such as video bitstreams) may be encoded according to certain video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In an example, a video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0036] Note that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may further include a video encoder (not shown).

[0037] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be provided in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 the video decoder (110) in the example of

[0038] The receiver (231) may receive one or more encoded video sequences to be decoded by the video decoder (210), and the one or more encoded video sequences are included in a bitstream, for example. In one aspect, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequence may be received from a channel (201), and the channel may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that can be forwarded to their respective using entities (not labeled). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be provided outside the video decoder (210) (not labeled). In still other applications, a buffer memory (not labeled) may be provided outside the video decoder (210) to prevent network jitter, for example, and another buffer memory (215) may be configured inside the video decoder (210) to handle playback timing, for example. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be needed, or the buffer memory may be made smaller. For use on a service packet network such as the Internet, the buffer memory (215) may also be needed, and the buffer memory may be relatively large, advantageously may have an adaptive size, and may be implemented at least partially in an operating system or a similar element (not labeled) outside the video decoder (210).

[0039] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210), and potential information for controlling a display device (212) (e.g., a display screen), such as a display device that is not part of the electronic device (230) but can be coupled to the electronic device (230), as shown in Figure 2 shown. The control information for the display device may be in the form of a parameter set segment (not labeled) of Supplementary Enhancement Information (SEI) or Video Usability Information (VUI). The parser (220) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be based on a video coding technology or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (220) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. The subgroups may include Group of Picture (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, MVs, and so on.

[0040] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0041] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. Which units are involved and the way they are involved may be controlled by the parser (220) through subgroup control information parsed from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (220) and the multiple units below are not described.

[0042] In addition to the functional blocks already mentioned, the video decoder (210) can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.

[0043] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform mode, block size, quantization factor, quantization scaling matrix, etc. to use. The scaler / inverse transform unit (251) can output blocks including sample values, and the sample values can be input into the aggregator (255).

[0044] In some cases, the output samples of the scaler / inverse transform unit (251) can belong to intra-coded blocks. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture but can use predictive information from a previously reconstructed portion of the current picture. Such predictive information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses the surrounding reconstructed information extracted from the current picture buffer (258) to generate a block with the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.

[0045] In other cases, the output samples of the scaler / inverse transform unit (251) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (253) can access the reference picture memory (257) to extract samples for prediction. After motion-compensating the extracted samples according to the symbols (221) belonging to the block, these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case, called residual samples or residual signals), thereby generating output sample information. The motion compensation prediction unit (253) obtaining the prediction samples from an address within the reference picture memory (257) can be controlled by a motion vector, and the motion vector is in the form of the symbol (221) for use by the motion compensation prediction unit (253), and the symbol (221) can have, for example, an X component, a Y component, and a reference picture component. Motion compensation can also include interpolation of sample values extracted from the reference picture memory (257), a motion vector prediction mechanism, etc. when using sub-sample accurate motion vectors.

[0046] The output samples of the aggregator (255) may undergo various loop filtering techniques in the loop filter unit (256). Video compression techniques may include in-loop filter techniques that are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and that may be available as symbols (221) from the parser (220) for the loop filter unit (256). Video compression may also respond to meta-information obtained during the decoding of previous (in decoding order) portions of the encoded picture or encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0047] The output of the loop filter unit (256) may be a sample stream that may be output to the display device (212) and stored in the reference picture memory (257) for subsequent inter-picture prediction.

[0048] Once fully reconstructed, some encoded pictures may be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture has been fully reconstructed and the encoded picture (e.g., via the parser (220)) has been identified as a reference picture, the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting the reconstruction of subsequent encoded pictures.

[0049] The video decoder (210) may perform decoding operations according to a predetermined video compression technique or a standard such as the ITU-T H.265 recommendation. The encoded video sequence may conform to the syntax specified by the video compression technique or standard in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured in, e.g., mega samples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0050] In one aspect, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (210) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.

[0051] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is provided in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used to replace Figure 1 the video encoder (103) in the example of

[0052] The video encoder (303) may receive video samples from a video source (301) (which is not Figure 3 part of the electronic device (320) in the example), and the video source may capture video images to be encoded by the video encoder (303). In another embodiment, the video source (301) is part of the electronic device (320).

[0053] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits,...), any color space (e.g., BT.601 Y CrCB, RGB,...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following focuses on describing the samples.

[0054] According to one aspect, a video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other required time constraints. Implementing an appropriate encoding speed is a function of a controller (350). In some aspects, the controller (350) controls and is functionally coupled to other functional units as described below. For simplicity, couplings are not labeled in the figures. Parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, GOP layout, maximum motion vector search range, etc. The controller (350) may be configured to have other suitable functions that relate to optimizing the video encoder (303) for a certain system design.

[0055] In some aspects, the video encoder (303) is configured to operate in an encoding loop. As an overly simplified description, in an example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data. The reconstructed sample stream (sample data) is input into a reference picture memory (334). Since the decoding of the symbol stream produces bit-exact results independent of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples “seen” by the prediction part of the encoder are exactly the same as the sample values that the decoder will “see” when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs, for example, when synchronization cannot be maintained due to channel errors) is also used in some related technologies.

[0056] The operation of the “local” decoder (333) may be the same as that of the “remote” decoder of the video decoder (210) described in detail above in connection with Figure 2 However, briefly referring additionally to Figure 2 , when symbols are available and the entropy encoding / decoding of the symbols into the encoded video sequence by the entropy encoder (345) and the parser (220) can be lossless, the entropy decoding part of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).

[0057] In one aspect, any decoder technology other than parsing / entropy decoding that exists in a decoder exists in a corresponding encoder in an identical or substantially identical functional form. Thus, the disclosed subject matter focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is reciprocal to the decoder technology described comprehensively. More detailed descriptions of certain aspects will be provided below.

[0058] During operation, in some examples, a source encoder (330) may perform motion-compensated predictive coding that predictively encodes an input picture by referring to one or more previously encoded pictures designated as "reference pictures" in a video sequence. In this way, an encoding engine (332) encodes the difference between a pixel block of the input picture and a pixel block of a reference picture, which may be selected as a prediction reference for the input picture.

[0059] A local video decoder (333) may decode the encoded video data of a picture that may be designated as a reference picture based on symbols created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data is decoded at a video decoder ( Figure 3 not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on a reference picture and may store the reconstructed reference picture in a reference picture cache (334). In this way, the video encoder (303) may locally store a copy of the reconstructed reference picture that has the same content (in the absence of transmission errors) as the reconstructed reference picture that will be obtained by a remote video decoder.

[0060] A predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search a reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as an appropriate prediction reference for the new picture. The predictor (335) may operate on a per-pixel block basis of sample blocks to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (335), an input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (334).

[0061] A controller (350) may manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0062] The outputs of all the above functional units can be entropy - encoded in an entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable - length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.

[0063] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) to prepare for transmission over a communication channel (360), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can merge the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0064] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, pictures can typically be assigned to any of the following picture types:

[0065] An intra - picture (I - picture), which can be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra - pictures, including, for example, Independent Decoder Refresh (IDR) pictures.

[0066] A predictive picture (P - picture), which can be a picture that can be encoded and decoded using intra - prediction or inter - prediction, where the intra - prediction or inter - prediction uses motion vectors and reference indices to predict the sample values of each block.

[0067] A bi - predictive picture (B - picture), which can be a picture that can be encoded and decoded using intra - prediction or inter - prediction, where the intra - prediction or inter - prediction uses two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0068] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, a block of an I picture may be non-predictively coded, or a block of the I picture may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. A block of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. A block of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0069] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T Rec. H.265. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0070] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0071] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In an embodiment, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0072] In some aspects, bidirectional prediction techniques can be used for inter - picture prediction. According to bidirectional prediction techniques, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in decoding order (but may be past and future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be predicted by a combination of the first reference block and the second reference block.

[0073] In addition, merge mode techniques can be used for inter - picture prediction to improve coding efficiency.

[0074] According to some aspects of the present disclosure, predictions such as inter - picture prediction and intra - picture prediction are performed on a per - block basis. For example, according to the High Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are segmented into Coding Tree Units (CTUs) for compression. CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three Coding Tree Blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be recursively split into one or more coding units (CUs) in a quadtree. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In an example, each CU is analyzed to determine the prediction type for the CU, such as an inter - prediction type or an intra - prediction type. Additionally, depending on temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In one aspect, prediction operations in encoding (encoding / decoding) are performed on a per - prediction - block basis. Taking the luminance prediction block as the prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0075] Note that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more integrated circuits. In another aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using one or more processors that execute software instructions.

[0076] Aspects of the present disclosure include techniques for improving cross-component prediction.

[0077] Video coding has been widely applied in many applications, such as broadcasting, video recording, and video streaming. Many emerging video coding standards, such as H.264, H.265 / HEVC, H.266 / VVC, and AV1, have been published and widely adopted in these video applications. In one aspect, a hybrid video codec includes multiple coding modules, such as intra prediction, inter prediction, transform coding, quantization, entropy coding, and post-loop filtering. In the present disclosure, improvements to cross-component prediction are provided to enhance the signaling and prediction of cross-component prediction.

[0078] In one aspect of the present disclosure, a first method is provided for combining multiple prediction signals, such as combining an inter prediction signal with an intra prediction signal. An example of the first method is described in Equation 1. P combi =(1 - ω)×P modeA +ω×P modeB Equation (1) where ω is a weight value between 0 and 1, P modeA is obtained by a prediction process using the method according to Mode A, and P modeB is obtained by a prediction process using the method according to Mode B. Then the two prediction samples are combined using the weighted average ω. Depending on the coding modes of adjacent blocks (such as Figure 4 the coding modes of the upper adjacent block (402) and the left adjacent block (404) of the current block (400) as shown), the weighted average ω can be signaled or calculated. Examples of weighted derivation using the upper adjacent block and the left adjacent block are as follows: (1) If the upper adjacent block is available and encoded by Mode B, set isModeBTop to 1, otherwise set isModeBTop to 0; (2) If the left adjacent block is available and encoded by Mode B, set isModeBLeft to 1, otherwise set isModeBLeft to 0; (3) If (isModeBLeft + isModeBTop) equals 2, then set ω to 0.75; (4) Otherwise, if (isModeBLeft + isModeBTop) equals 1, then set ω to 0.5; and (5) Otherwise, set ω to 0.25.

[0079] In one aspect of the present disclosure, a second method is provided, such as an inter-block cross-component prediction (CCP) method, to predict chrominance samples from reconstructed luminance samples when encoding a current block by using a motion vector (MV) to point to a predicted block in a reference picture in an inter-frame prediction mode, or by using a block vector (BV) to point to a predicted block in a reconstructed picture in an alternative prediction mode. The BV can be signaled explicitly, inherited from an adjacent encoded block, or obtained implicitly by comparing the distortion cost between a template of the current encoded block and a reference template within a predefined reconstructed region.

[0080] Figure 5 and Figure 6 Two examples of the CCP method are shown from the decoder side. As Figure 5 and Figure 6 shown, a cross-component filter is obtained by using a predicted block of a luminance block and a predicted block of a chrominance block. The obtained filter is applied to the reconstructed luminance block to predict the chrominance block. Figure 5 Shows an inter-component prediction for predicting a chrominance prediction block without mixing with a chrominance prediction value. Figure 6 Shows an inter-component prediction for predicting a chrominance prediction block with mixing with a chrominance prediction value. In Figure 6 it is possible to mix the predicted chrominance block with the chrominance prediction value using an MV or a BV to produce a final chrominance prediction block.

[0081] Figure 5 Examples of cross-component prediction (500) of chrominance blocks are provided. As Figure 5As shown, the luminance prediction value (502) and chrominance prediction value (504) of the block can be applied to obtain the filtering coefficients (506) of the filter, such as obtaining the cross-component filtering coefficients of the cross-component filter. In the example, the luminance prediction value (502) and chrominance prediction value (504) are obtained based on any suitable prediction mode, such as the inter prediction mode, intra prediction mode, Intra Block Copy (IBC) mode, Cross-Component Linear Model (CCLM), Multi-Model Linear Model (MMLM), Convolutional Cross-Component Intra Prediction Model (CCCM), and Gradient Linear Model (GLM). The cross-component filtering coefficients can be obtained based on any one of the cross-component modes. For example, the cross-component filtering coefficients can be obtained based on one of CCLM, MMLM, CCCM, and GLM. The obtained cross-component filtering coefficients (506) can be further applied to the reconstructed luminance block (507) to generate the predicted chrominance block (509). The reconstructed luminance block (507) can be determined as the sum of the luminance prediction value (502) and the luminance residual (510). The chrominance reconstruction samples (516) can be determined as the sum of the predicted chrominance block (509) and the chrominance residual (512). The chrominance residual (512) can be the difference between the chrominance block and the chrominance prediction value (504). In addition, the luminance reconstruction samples (514) can be determined as the sum of the luminance prediction value (502) and the luminance residual (510).

[0082] Figure 6 An example of cross-component prediction (600) of a chrominance block is shown, where the predicted chrominance block (609) is blended with the chrominance prediction value (604). As Figure 6 shown, the chrominance reconstruction samples (616) are determined as a weighted combination of the chrominance prediction value (604) and the predicted chrominance block (609) based on the weighting factor ω.

[0083] In one aspect of the present disclosure, a third method for cross-component prediction is provided. In the third method, a merge list is constructed by using neighboring blocks from a reference picture, such as (1) neighboring spatially adjacent blocks, (2) non-neighboring spatially adjacent blocks, (3) history-based spatially adjacent blocks, (4) temporally collocated blocks, and / or (5) temporally shifted blocks (where the shifted motion vector is obtained from a neighboring block), and is used to obtain a prediction mode (e.g., a cross-component prediction mode). Cross-component prediction may include intra cross-component prediction and inter-block cross-component prediction. For intra cross-component prediction, a merge list construction may be obtained. For inter-block cross-component prediction shown in the second method, a merge list construction for the second method may be constructed, such as based on neighboring blocks.

[0084] In an example of the third method, filter coefficients of a filter, such as filter coefficients (506) or (606), may be derived based on merge candidates in the merge list. For example, when a merge candidate is encoded in a cross-component prediction mode, such as one of CCLM, MMLM, CCCM, and GLM, the filter coefficients obtained for the merge candidate encoded in the cross-component prediction mode may be applied in the cross-component prediction of the third method. Thus, it may not be necessary Figure 5 for the filter coefficients (506) derived in Figure 6 and the filter coefficients (606) derived in

[0085] In the present disclosure, combined inter and intra prediction may refer to the combined inter and intra prediction in the first method to predict a predicted sample by using a weighted average of predicted samples from two different prediction modes. Inter-block cross-component prediction (CCP) may refer to the inter-block cross-component prediction in the second method, where a cross-component prediction mode is obtained by using a predicted block, and then when encoding the current block in an inter prediction or in a prediction mode where a block vector (BV) points to a predicted block in a reconstructed picture, the cross-component prediction mode is applied to a luminance reconstructed block to predict a chrominance prediction block. The block vector may be signaled explicitly, inherited from a neighboring encoded block, or obtained implicitly by comparing the template of the current encoded block with a reference template within a predefined reconstructed region. The merge list may refer to a merge list construction for cross-component prediction, and the cross-component prediction may be, but is not limited to, the second method (e.g., inter-block cross-component prediction) or other existing intra cross-component predictions.

[0086] In the present disclosure, a weighted average mode (or method) is provided, in which a final chrominance prediction block can be obtained by using a weighted average of a plurality of chrominance prediction blocks. The plurality of chrominance prediction blocks can be obtained from one or more of the methods discussed above. Encoding information, such as a first syntax element or a first flag, can be signaled to indicate whether the weighted average mode is used. When the first flag is true (or a first value such as 1), it means that the weighted average of a plurality of chrominance prediction blocks is applied. For example, the weighted average of (i) a chrominance prediction block of a chrominance component of a current block based on inter-frame prediction and (ii) a chrominance prediction block of a chrominance component based on one of a second method or a third method is applied. Otherwise, when the first flag is false (or a second value such as 0), the weighted average method (or mode) is not applied to the current block.

[0087] In one aspect, the first flag (or the first syntax) is signaled based on one or more conditions. For example, when a prediction method (such as a second method or a third method) is applied to the current block, the first flag can be signaled.

[0088] In one aspect, the first flag is signaled at an encoding structure level (such as an encoding block level, a transform block level, etc.).

[0089] In one aspect, when the first flag is true (or a first value such as 1), the second method or the third method is applied to the current block. Thus, in the weighted average method indicated by the first flag, the inter-frame prediction block can be Mode A, and the prediction block based on the second method or the third method can be Mode B. When the second method or the third method is selected and signaled, a combined prediction block of the current block can be obtained based on the inter-frame prediction block and the prediction block from the second method or the third method.

[0090] In one aspect, when the first flag is true and the second method or the third method is applied to the current block, a luminance prediction block of a luminance component of the current block is formed by directly copying the luminance inter-frame prediction block.

[0091] In one aspect, the first flag is signaled only for the chrominance component. For the luminance component, another syntax element or another flag is used to indicate whether the first method is used.

[0092] In one aspect, the derivation of the weight (e.g., ω) depends on neighboring blocks and the prediction modes of the neighboring blocks. For example, the weight can be obtained based on whether the neighboring upper block and / or the neighboring left block are encoded by the second method.

[0093] In an example, when encoding both the adjacent upper block and the adjacent left block in the second method, the weight of the second method is the first weight (e.g., 0.75). In an example, when encoding at least one of the adjacent upper block and the adjacent left block in the second method, the weight of the second method is the second weight (e.g., 0.5). In an example, when neither the adjacent upper block nor the adjacent left block is encoded in the second method, the weight of the second method is the third weight (e.g., 0.25).

[0094] In one aspect, the weight derivation depends on whether the adjacent upper block and the adjacent left block are encoded in the third method.

[0095] In an example, when encoding both the adjacent upper block and the adjacent left block in the third method, the weight of the third method is 0.75. In an example, when encoding at least one of the adjacent upper block and the adjacent left block in the third method, the weight of the third method is 0.5. In an example, when neither the adjacent upper block nor the adjacent left block is encoded in the third method, the weight of the third method is 0.25.

[0096] In one aspect, the weight derivation depends on whether the adjacent upper block and the adjacent left block are encoded in the second method or the third method.

[0097] In an example, when encoding both the adjacent upper block and the adjacent left block in the second method or the third method, the weight of the second method or the third method is 0.75. In an example, when encoding at least one of the adjacent upper block and the adjacent left block in the second method or the third method, the weight of the second method or the third method is 0.5. In an example, when neither the adjacent upper block nor the adjacent left block is encoded in the second method or the third method, the weight of the second method or the third method is 0.25.

[0098] In the present disclosure, a second list (e.g., a merge list) may be constructed for cross-component prediction modes. The cross-component prediction modes may include but are not limited to cross-component intra prediction and inter-block cross-component prediction. The merge list may be constructed from cross-component prediction encoded blocks. The cross-component prediction encoded blocks may include (1) adjacent spatially adjacent blocks, (2) non-adjacent spatially adjacent blocks, (3) history-based adjacent blocks, (4) temporally collocated blocks, and / or (5) temporally shifted blocks (where the shifted motion vectors are obtained from adjacent blocks) in a reference picture. In an example, the cross-component prediction encoded blocks are encoded according to any suitable cross-component intra prediction or inter-block cross-component prediction in the second method. In an example, the cross-component prediction encoded blocks are encoded in a cross-component prediction mode, such as one of CCLM, MMLM, CCCM, and GLM. In an example, the filtering coefficients obtained for the cross-component prediction encoded blocks may be applied to the cross-component prediction of the current block.

[0099] In one aspect, signaling is used to convey coding information (such as a second syntax element or a second flag) to indicate whether to use a second list (or a merge list). For example, the second flag can indicate whether to use the second list to obtain the filtering coefficients for cross-component prediction of the current block. If the second flag is true (or a first value such as 1), the second list is used. Then, another syntax element, such as an index, can be signaled to indicate which candidate in the second list is selected. Thus, the current block can be predicted based on cross-component prediction that includes filtering coefficients copied (or obtained) from the selected candidate in the merge list. Otherwise, when the second flag is false (or a second value such as 0), the obtained and / or signaled cross-component prediction can be used for the current block. In the example, the obtained cross-component prediction is Figure 5 and Figure 6 the cross-component prediction shown, where the filtering coefficients of the cross-component prediction are obtained based on chrominance prediction values and luminance prediction values.

[0100] In one aspect, the second list is constructed based on candidates in a predefined scan order: adjacent spatially adjacent coded blocks from reference pictures, history-based adjacent coded blocks, non-adjacent spatially adjacent coded blocks, and / or temporally collocated coded blocks.

[0101] For example, adjacent spatially adjacent coded blocks are scanned (or identified) first. If adjacent spatially adjacent coded blocks are available, they are filled in the merge list. Subsequently, if the history information is not empty, history-based adjacent coded blocks are inserted. Additionally, the availability of non-adjacent spatially adjacent coded blocks is checked, and thus the availability of temporally collocated coded blocks is checked from the reference frame.

[0102] In one aspect, the candidates in the second list are reordered. The reordering can be performed according to a criterion or other guidelines. For example, the candidates can be reordered in ascending order based on the cost of template matching of the corresponding candidates. In the example, the prediction parameters (such as filtering coefficients) of each candidate are applied to the template of the candidate to calculate the template matching cost.

[0103] In one aspect, after constructing the merge list, the second list (or the merge list) is divided into multiple groups, such as N groups, where N is a non-negative number. The above index can be divided into two syntax elements (or two syntax parts): one index part can indicate the group ID of the N groups, which ranges from (0, N - 1), and the other index part can indicate the index (or position) of the selected candidate in the group identified by the group ID.

[0104] Figure 7A flowchart is shown that outlines a method (700) according to an aspect of the present disclosure. The method (700) can be used in a video decoder. In various aspects, the method (700) is executed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some aspects, the method (700) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the method (700). The method starts at (S701) and proceeds to (S710).

[0105] At (S710), a bitstream is received, the bitstream including a first syntax element associated with a current block in a current picture. The current block includes a chrominance block and a luminance block. The first syntax element indicates whether the chrominance block is predicted by a weighted average of a plurality of chrominance prediction blocks.

[0106] At (S720), when the first syntax element indicates that the chrominance block is predicted by a weighted average of a plurality of chrominance prediction blocks, a first chrominance prediction block of the plurality of chrominance prediction blocks is determined based on an inter prediction mode. A second chrominance prediction block of the plurality of chrominance prediction blocks is determined based on a cross-component prediction mode, in which the second chrominance prediction block is obtained based on reconstructed luminance samples of the luminance block.

[0107] At (S730), the prediction block of the chrominance block is determined as a weighted average of the first chrominance prediction block and the second chrominance prediction block.

[0108] In an example, the cross-component prediction mode includes a first cross-component prediction mode, in which the second chrominance prediction block is obtained based on reconstructed luminance samples of the luminance block filtered according to the filter coefficients of a first filter. The filter coefficients of the first filter are obtained based on the chrominance prediction of the chrominance block and the luminance prediction of the luminance block. The cross-component prediction mode includes a second cross-component prediction mode, in which the second chrominance prediction block is obtained based on reconstructed luminance samples of the luminance block filtered by the filter coefficients of a second filter. The filter coefficients of the second filter are obtained based on merge candidates in a merge list.

[0109] In an example, when the cross-component prediction mode is applied to the current block, the first syntax element is signaled in the bitstream.

[0110] In an example, when the first syntax element indicates that the chrominance block is predicted by a weighted average of a plurality of chrominance prediction blocks and the cross-component prediction mode is applied to the current block, the prediction block of the luminance block is determined by copying an inter prediction block of the luminance. In an example, a flag in the bitstream is received. The flag indicates whether the luminance block is predicted by a weighted average of a plurality of luminance prediction blocks.

[0111] In the example, when encoding both the upper neighboring block and the left neighboring block of the current block in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.75. When encoding one of the upper neighboring block and the left neighboring block of the current block in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.5. When neither the upper neighboring block nor the left neighboring block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.25.

[0112] In the example, the merge list is constructed based on a plurality of cross-component prediction encoded blocks, the plurality of cross-component prediction encoded blocks including at least one of (i) adjacent spatially neighboring blocks, (ii) non-adjacent spatially neighboring blocks, (iii) history-based neighboring blocks, (iv) temporally collocated blocks, and (v) temporally shifted blocks in a reference picture of the current picture. The merge list is constructed based on a predefined scan order of adjacent spatially neighboring blocks, history-based neighboring blocks, non-adjacent spatially neighboring blocks, and temporally collocated blocks from the reference picture.

[0113] In the example, when a second syntax element in the bitstream indicates that the merge list is used for the cross-component prediction mode, a merge candidate is determined from the merge list according to an index in the bitstream.

[0114] In the example, when the second syntax element indicates that the second cross-component prediction mode is not applied, the filtering coefficient of the first filter is obtained based on the chrominance prediction of the chrominance block and the luminance prediction of the luminance block, or the filtering coefficient of the first filter is determined according to information signaled in the bitstream.

[0115] In the example, the merge list is divided into a plurality of subgroups. A merge candidate is determined from the merge list according to an index in the bitstream. The index includes a first part and a second part, the first part indicating which one of the plurality of subgroups to select, and the second part indicating which merge candidate to select from the selected subgroup.

[0116] Then, the method proceeds to (S799) and ends.

[0117] The method (700) can be adjusted appropriately. At least one step in the method (700) can be modified and / or omitted. At least one additional step can be added. Any suitable execution order can be used.

[0118] Figure 8A flowchart is shown that outlines a method (800) according to one aspect of the present disclosure. The method (800) can be used in a video encoder. In various aspects, the method (800) is executed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some aspects, the method (800) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the method (800). The method starts at (S801) and proceeds to (S810).

[0119] At (S810), a first chrominance prediction block of the chrominance block of the current block is determined based on an inter-frame prediction mode. A second chrominance prediction block of the chrominance block of the current block is determined based on a cross-component prediction mode, in which the second chrominance prediction block is obtained based on the reconstructed luminance samples of the luminance block of the current block.

[0120] At (S820), the prediction block of the chrominance block is determined as a weighted average of the first chrominance prediction block of the plurality of chrominance prediction blocks and the second chrominance prediction block of the plurality of chrominance prediction blocks.

[0121] At (S830), a first syntax element is encoded into the bitstream, and the first syntax element indicates that the chrominance block of the current block is predicted by a weighted average of the plurality of chrominance prediction blocks.

[0122] In an example, the cross-component prediction mode includes a first cross-component prediction mode, in which the second chrominance prediction block is obtained based on the reconstructed luminance samples of the luminance block filtered according to the filtering coefficients of a first filter. The filtering coefficients of the first filter are obtained based on the chrominance prediction of the chrominance block and the luminance prediction of the luminance block. The cross-component prediction mode includes a second cross-component prediction mode, in which the second chrominance prediction block is obtained based on the reconstructed luminance samples of the luminance block filtered by the filtering coefficients of a second filter. The filtering coefficients of the second filter are obtained based on the merge candidates in the merge list.

[0123] In an example, when both the upper adjacent block and the left adjacent block of the current block are encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.75. When at least one of the upper adjacent block and the left adjacent block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.5. When neither the upper adjacent block nor the left adjacent block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.25.

[0124] Then, the processing proceeds to (S899) and ends.

[0125] The method (800) can be adjusted appropriately. At least one step in the method (800) can be modified and / or omitted. At least one additional step can be added. Any suitable execution order can be used.

[0126] In one aspect, a method of processing visual media data includes: processing a bitstream of visual media data according to formatting rules. For example, the bitstream can be a bitstream decoded / encoded by any of the decoding and / or encoding methods described herein. The formatting rules can specify one or more constraints of the bitstream and / or one or more processes performed by a decoder and / or an encoder.

[0127] In an example, the bitstream includes a first syntax element associated with a current block in a current picture. The current block includes a chrominance block and a luminance block. The first syntax element indicates whether the chrominance block is predicted by a weighted average of a plurality of chrominance prediction blocks. The formatting rules specify that when the first syntax element indicates that the chrominance block is predicted by a weighted average of a plurality of chrominance prediction blocks, a first chrominance prediction block of the plurality of chrominance prediction blocks is determined based on an inter-frame prediction mode. A second chrominance prediction block of the plurality of chrominance prediction blocks is determined based on a cross-component prediction mode, in which the second chrominance prediction block is obtained based on reconstructed luminance samples of a luminance block filtered according to filter coefficients of a filter. The filter coefficients of the filter are based on (i) the chrominance prediction of the chrominance block and the luminance prediction of the luminance block, or (ii) a merge candidate in a merge list encoded in the cross-component prediction mode. The formatting rules specify that the prediction block of the chrominance block is determined as a weighted average of the first chrominance prediction block and the second chrominance prediction block.

[0128] The techniques described above can be implemented as computer software that uses computer-readable instructions and is physically stored on one or more computer-readable media. For example, Figure 9 FIG. shows a computer system (900) suitable for implementing certain embodiments of the disclosed subject matter.

[0129] The computer software can be encoded using any suitable machine code or computer language, and any suitable machine code or computer language can be assembled, compiled, linked, or similar mechanisms to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode, etc.

[0130] The instructions can be executed on various types of computers or components thereof, such as personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0131] Figure 9 The components of the computer system (900) shown are exemplary and are not intended to place any limitation on the use or functionality scope of computer software for implementing aspects of the present disclosure. The configuration of the components should not be construed as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiments of the computer system (900).

[0132] The computer system (900) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to one or more human users through inputs such as, for example: tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, clapping), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to human conscious inputs, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, captured images from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.

[0133] The input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard (901), mouse (902), touchpad (903), touch screen (910), data glove (not shown), joystick (905), microphone (906), scanner (907), camera (908).

[0134] The computer system (900) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile outputs, sounds, lights, and smells / tastes. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback of the touch screen (910), data glove (not shown), or joystick (905), but may also be tactile feedback devices that are not input devices), audio output devices (e.g., speakers (909), headphones (not shown)), visual output devices (e.g., screens (910) including CRT screens, LCD screens, plasma screens, OLED screens, each screen having or not having touch screen input functionality, each screen having or not having tactile feedback functionality, some of which are capable of outputting two-dimensional visual outputs or more than three-dimensional outputs through devices such as stereoscopic image outputs, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted)).

[0135] The computer system (900) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (920) with media (921) such as CD / DVD, thumb drives (922), removable hard disk drives or solid state drives (923), traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0136] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0137] The computer system (900) may also include an interface to one or more communication networks (955). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local area network, a wide area network, a metropolitan area network, vehicle and industrial networks, real-time networks, delay-tolerant networks, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, and so on. Some networks typically require an external network interface adapter (e.g., a USB port of the computer system (900)) connected to certain common data ports or peripheral buses (949); as described below, other network interfaces are typically integrated into the kernel of the computer system (900) by connecting to the system bus (e.g., an Ethernet interface connected to a PC computer system or a cellular network interface connected to a smartphone computer system). The computer system (900) may communicate with other entities using any of these networks. Such communication may be one-way reception only (e.g., broadcast television), one-way transmission only (e.g., CANbus connected to certain CANbus devices), or two-way, e.g., connecting to other computer systems using a local area network or a wide area network digital network. As described above, certain protocols and protocol stacks may be used on each of those networks and network interfaces.

[0138] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel (940) of the computer system (900).

[0139] The kernel (940) may include one or more central processing units (CPUs) (941), a graphics processing unit (GPU) (942), a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) (943), a hardware accelerator (944) for certain tasks, a graphics adapter (950), etc. These devices, as well as a read-only memory (ROM) (945), a random access memory (946), an internal mass storage (947) such as an internal hard disk drive, SSD, etc. that is not accessible to users, may be connected via a system bus (948). In some computer systems, the system bus (948) may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be directly connected to the system bus (948) of the kernel or connected to the system bus (948) of the kernel via a peripheral bus (949). In one example, a screen (910) may be connected to the graphics adapter (950). The architecture of the peripheral bus includes PCI, USB, etc.

[0140] The CPU (941), GPU (942), FPGA (943), and accelerator (944) may execute certain instructions, which may be combined to form the above-mentioned computer code. The computer code may be stored in the ROM (945) or the RAM (946). Transitional data may also be stored in the RAM (946), while permanent data may be stored, for example, in the internal mass storage (947). Fast storage and retrieval of any storage device may be performed by using a cache, which may be closely associated with one or more CPUs (941), GPUs (942), mass storage (947), ROM (945), RAM (946), etc.

[0141] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code that are specially designed and constructed for the purposes of the present disclosure, or the medium and the computer code may be of the type well-known and available to those skilled in the field of computer software.

[0142] As a non-limiting example, software included in one or more tangible computer-readable media (including CPUs, GPUs, FPGAs, accelerators, etc.) can be executed by one or more processors such that a computer system having an architecture (900), particularly a core (940), can provide functionality. Such computer-readable media can be media associated with the user-accessible mass storage as described above, as well as certain non-transitory memories of the core (940), such as the on-core mass memory (947) or ROM (945). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (940). Depending on specific needs, the computer-readable media can include one or more storage devices or chips. The software can cause the core (940), particularly the processors therein (including CPUs, GPUs, FPGAs, etc.), to execute specific processes or specific portions of specific processes described herein, including defining data structures stored in the RAM (946) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality by hardwired logic or logic otherwise embodied in circuitry (e.g., accelerator (944)) that can replace the software or operate in conjunction with the software to execute specific processes or specific portions of specific processes described herein. In appropriate cases, portions referring to software can include logic, and vice versa. In appropriate cases, portions referring to computer-readable media can include circuitry (e.g., integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or including both. The present disclosure encompasses any suitable combination of hardware and software.

[0143] As used in this disclosure, "at least one" or "one of" is intended to include any one or combination of the recited elements. For example, when referring to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C, it is intended to include only A, only B, only C, or any combination among A, B, and C. When referring to one of A or B and one of A and B, it is intended to include A or B, or (A and B). Where applicable, such as when the elements are not mutually exclusive, using "one of" does not exclude any combination of the recited elements.

[0144] Although the present disclosure has described examples of multiple aspects, there are modifications, permutations, and various alternative equivalents that fall within the scope of the present disclosure. Accordingly, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure.

Claims

1. A method for processing visual media data, the method comprising: Processing the code stream of the visual media data according to the format rules, wherein: The code stream includes a first syntax element associated with a current block in a current picture, the current block includes a chrominance block and a luminance block, the first syntax element indicating whether the chrominance block is predicted by a weighted average of a plurality of chrominance prediction blocks; and The format rules specify: When the first syntax element indicates that the chroma block is predicted by a weighted average of the plurality of chroma prediction blocks, (i) a first chroma prediction block of the plurality of chroma prediction blocks is determined based on an inter-prediction mode, and (ii) a second chroma prediction block of the plurality of chroma prediction blocks is determined based on an inter-component prediction mode, in which the second chroma prediction block is obtained based on reconstructed luma samples of the luma block filtered according to a filter coefficient of a filter, wherein the filter coefficient of the filter is obtained based on (i) a chroma prediction of the chroma block and a luma prediction of the luma block, or (ii) a merge candidate in a merge list encoded in the inter-component prediction mode; and A prediction block of the chroma block is determined as a weighted average of the first chroma prediction block and the second chroma prediction block.

2. The method according to claim 1, wherein: When both the upper neighboring block and the left neighboring block of the current block are encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.75; When at least one of the upper neighboring block and the left neighboring block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.5; as well as When both the upper neighboring block and the left neighboring block of the current block are not encoded in the cross-component prediction mode, the weight of the second chroma prediction block in the weighted average is 0.

25.

3. The method according to claim 1 or 2, wherein: The format rules specify: The merge list is constructed based on multiple cross-component prediction coding blocks, wherein the multiple cross-component prediction coding blocks include at least one of (i) adjacent spatial neighboring blocks, (ii) non-adjacent spatial neighboring blocks, (iii) history-based neighboring blocks, (iv) temporally collocated blocks, and (v) temporally shifted blocks in a reference picture of the current picture.

4. A video encoding method, comprising: (i) determining a first chroma prediction block of a chroma block of a current block based on an inter-prediction mode, and (ii) determining a second chroma prediction block of the chroma block of the current block based on an inter-component prediction mode, wherein the second chroma prediction block is obtained based on a reconstructed luma sample of a luma block of the current block; Determine a prediction block of the chroma block as a weighted average of the first chroma prediction block of a plurality of chroma prediction blocks and the second chroma prediction block of the plurality of chroma prediction blocks; as well as A first syntax element is encoded into a code stream, wherein the first syntax element indicates that the chroma block of the current block is predicted by the weighted average of the multiple chroma prediction blocks.

5. The method according to claim 4, wherein: The cross-component prediction mode includes one of the following: a first cross-component prediction mode, in which the second chroma prediction block is obtained based on the reconstructed luma samples of the luma block filtered according to filter coefficients of a first filter, wherein the filter coefficients of the first filter are obtained based on a chroma prediction of the chroma block and a luma prediction of the luma block; as well as A second cross-component prediction mode, in which the second chroma prediction block is obtained based on the reconstructed luma samples of the luma block filtered by filter coefficients of a second filter, and the filter coefficients of the second filter are obtained based on merge candidates in a merge list.

6. The method according to claim 4 or 5, wherein: When both the upper neighboring block and the left neighboring block of the current block are encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.75; When at least one of the upper neighboring block and the left neighboring block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.5; as well as When both the upper neighboring block and the left neighboring block of the current block are not encoded in the cross-component prediction mode, the weight of the second chroma prediction block in the weighted average is 0.

25.

7. A device for video decoding, comprising: The processing circuit is configured to: Receive a code stream, the code stream comprising a first syntax element associated with a current block in a current picture, the current block comprising a chrominance block and a luminance block, the first syntax element indicating whether the chrominance block is predicted by a weighted average of a plurality of chrominance prediction blocks; When the first syntax element indicates that the chroma block is predicted by the weighted average of the plurality of chroma prediction blocks, (i) determining a first chroma prediction block of the plurality of chroma prediction blocks based on an inter prediction mode, and (ii) determining a second chroma prediction block of the plurality of chroma prediction blocks based on an inter-component prediction mode, in which the second chroma prediction block is obtained based on a reconstructed luma sample of the luma block; as well as A prediction block of the chroma block is determined as a weighted average of the first chroma prediction block and the second chroma prediction block.

8. The device according to claim 7, wherein: The cross-component prediction mode includes one of the following: a first cross-component prediction mode, in which the second chroma prediction block is obtained based on the reconstructed luma samples of the luma block filtered according to filter coefficients of a first filter, wherein the filter coefficients of the first filter are obtained based on a chroma prediction of the chroma block and a luma prediction of the luma block; as well as A second cross-component prediction mode, in which the second chroma prediction block is obtained based on the reconstructed luma samples of the luma block filtered by filter coefficients of a second filter, and the filter coefficients of the second filter are obtained based on merge candidates in a merge list.

9. The device according to claim 7 or 8, wherein: When the cross-component prediction mode is applied to the current block, the first syntax element is signaled in the codestream.

10. The device according to any one of claims 7 to 9, wherein: The processing circuit is configured to: when the first syntax element indicates that the chroma block is predicted by the weighted average of the plurality of chroma prediction blocks and the cross-component prediction mode is applied to the current block, determining a prediction block for the luma block by copying a luma inter prediction block, or A flag in the code stream is received, the flag indicating whether to use a weighted average of multiple luma prediction blocks to predict the luma block.

11. The device according to any one of claims 7 to 10, wherein: When both the upper neighboring block and the left neighboring block of the current block are encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.75; When one of the upper neighboring block and the left neighboring block of the current block is encoded in the cross-component prediction mode, the weight of the second chrominance prediction block in the weighted average is 0.5; as well as When both the upper neighboring block and the left neighboring block of the current block are not encoded in the cross-component prediction mode, the weight of the second chroma prediction block in the weighted average is 0.

25.

12. The device according to any one of claims 8 to 11, wherein: The processing circuit is configured to: The merge list is constructed based on multiple cross-component prediction coding blocks, wherein the multiple cross-component prediction coding blocks include: at least one of (i) adjacent spatial neighboring blocks, (ii) non-adjacent spatial neighboring blocks, (iii) history-based adjacent blocks, (iv) temporally collocated blocks, and (v) temporally shifted blocks in a reference picture of the current picture, and the merge list is constructed based on a predefined scanning order of the adjacent spatial neighboring blocks, the history-based adjacent blocks, the non-adjacent spatial neighboring blocks, and the temporally collocated blocks from the reference picture.

13. The device according to any one of claims 8 to 12, wherein: The processing circuit is configured to: When a second syntax element in the codestream indicates that the merge list is used for the cross-component prediction mode, the merge candidate is determined from the merge list indicated by an index in the codestream.

14. The device according to any one of claims 8 to 13, wherein: The processing circuit is configured to: When the second syntax element indicates that the second cross-component prediction mode is not applied, obtaining the filter coefficients of the first filter based on the chroma prediction of the chroma block and the luma prediction of the luma block, or The filter coefficient of the first filter is determined according to information signaled in the code stream.

15. The device according to any one of claims 8 to 14, wherein: The processing circuit is configured to: dividing the merged list into a plurality of subgroups; and The merge candidate is determined from the merge list indicated by an index in the code stream, wherein the index includes a first part and a second part, the first part indicating which subgroup of the multiple subgroups is selected, and the second part indicating which merge candidate of the merge candidates is selected from the selected subgroup.