Method and device for video processing and medium
By employing decoder derivation and a multi-hypothesis CCP mode in video encoding and decoding, the problem of insufficient encoding and decoding efficiency in existing technologies is solved, achieving more efficient encoding and decoding performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2026-03-27
AI Technical Summary
The efficiency of existing video encoding and decoding technologies needs to be further improved.
A decoder-derived method is used to determine whether to use a real-time computed model or a derived model. The number of neighboring blocks is encoded or decoded using a prediction mode or method. A multi-hypothesis CCP mode is applied to convert video units and bitstreams.
It improves the performance of video encoding and decoding, and increases encoding and decoding efficiency.
Smart Images

Figure CN121753330A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to the application of inter-frame CCP / CCCM modes. Background Technology
[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multi-Functional Video Codec (VVC) standard. However, the encoding and decoding efficiency of video encoding and decoding technologies is generally expected to be further improved. Summary of the Invention
[0003] Embodiments of this disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for a conversion between video units and a bitstream of video, determining, based on a decoder-derived method, whether to use at least one of the following: a real-time computed model or a derived model; and performing the conversion based on at least one of the following: a real-time computed model or a derived model. Compared to conventional solutions, the method according to the first aspect of this disclosure can improve encoding and decoding performance.
[0005] In a second aspect, another method for video processing is proposed. This method includes: for a conversion between video units and a video bitstream, determining at least one of the following based on at least one of the following: the presence of a syntax element, or the manner in which a syntax element is signaled; and performing the conversion based on at least one of the following: the presence of a syntax element, or the manner in which a syntax element is signaled. Compared to conventional solutions, the method according to the second aspect of this disclosure can improve encoding and decoding performance.
[0006] In a third aspect, another method for video processing is proposed. This method includes: applying a multi-hypothesis CCP pattern to video blocks of video units for conversion between video units and video bitstreams; and performing the conversion based on the multi-hypothesis CCP pattern. Compared to conventional solutions, the method according to the third aspect of this disclosure can improve encoding and decoding performance by applying a multi-hypothesis CCP pattern to video units.
[0007] In a fourth aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to the first, second, or third aspect of this disclosure.
[0008] In a fifth aspect, a non-transitory computer-readable storage medium is provided. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method according to the first, second, or third aspect of this disclosure.
[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining, based on a decoder-derived method, whether to use at least one of the following: a real-time computed model or a derived model; and generating a bitstream based on at least one of the following: a real-time computed model or a derived model.
[0010] In a seventh aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores video data generated as a bitstream by a method performed by an apparatus for video processing. The method includes: determining at least one of the following based on at least one of the presence of neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using a predictive mode or method: the presence of syntax elements, or the manner in which syntax elements are transmitted via signaling; and generating a bitstream based on at least one of the following: the presence of syntax elements or the manner in which syntax elements are transmitted via signaling.
[0011] In the eighth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: applying a multi-hypothesis CCP mode to video blocks of video units; and generating a bitstream based on the multi-hypothesis CCP mode.
[0012] In a ninth aspect, a method for storing a bitstream of video is proposed. The method includes: determining, based on a decoder-derived method, whether to use at least one of the following: a real-time computed model or a derived model; generating a bitstream based on at least one of the following: a real-time computed model or a derived model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0013] In a tenth aspect, a method for storing a bitstream of video is proposed. The method includes: determining at least one of the following based on at least one of the following: the presence of syntax elements or the manner of transmitting syntax elements via signaling; generating a bitstream based on at least one of the following: the presence of syntax elements or the manner of transmitting syntax elements via signaling; and storing the bitstream in a non-transitory computer-readable recording medium.
[0014] In the eleventh aspect, a method for storing video bitstreams is proposed. The method includes: applying a multi-hypothesis CCP mode to video blocks of video units; generating a bitstream based on the multi-hypothesis CCP mode; and storing the bitstream in a non-transitory computer-readable recording medium.
[0015] This synopsis aims to present, in a simplified form, the concept choices further described below in the detailed embodiments. This synopsis is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description
[0016] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0017] Figure 1 A block diagram illustrating an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram illustrating a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 The diagram illustrates the effect of the slope adjustment parameter "u". Left: Model created using the current CCLM. Right: Updated model as proposed. Figure 5 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown. Figure 6 The neighboring reconstructed sample points used for the DIMD chromaticity mode are shown; Figure 7 The intra-frame template matching search area used is shown; Figure 8 The use of intra-frame TMP block vectors for IBC blocks is shown; Figure 9 The method for dividing angle patterns is shown; Figure 10 An expanded list of MRL candidates is shown; Figure 11 An illustration of the template area is shown; Figure 12 The spatial portion of the convolution filter is shown; Figure 13 The reference region (with its padding) used to derive the filter coefficients is shown. Figure 14 Four Sobel-based gradient modes for GLM are shown; Figure 15 The spatial GPM candidates are shown; Figure 16 The GPM template is shown; Figure 17 GPM mixing is shown; Figure 18 The possible locations of the candidate regions are shown; Figure 19 The locations of adjacent airspace candidates are shown; Figure 20 The transformation selection process for the orientation plane mode is shown; Figure 21 The luminance block used to derive the direct block vector is shown; Figure 22 The diagram shows three types of reconstructed regions, including thirteen columns or rows of reconstructed pixels; Figure 23 The diagram illustrates three types of filter shapes, each with fifteen inputs and producing one output. Figure 24 Examples of predictions for different locations within the current block are shown; Figure 25 The proposed method on the decoder is shown; Figure 26 The luminance samples L0, ..., L5 associated with the chrominance sample C are shown (in a half-pixel luminance grid). Figure 27 The luminance samples L0, ..., L5 associated with chrominance sample C are shown; Figure 28 An example of the current template (colored in dark gray) and reference template (colored in dark gray) involved in CCRM encoding and decoding for the current inter-frame block (filled with dots) is shown; Figure 29An example of the current template (colored in dark gray) and reference template (colored in dark gray) involved in CCRM encoding and decoding for the current IBC block (filled with dots) is shown; Figure 30 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 31 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 32 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 33 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0018] Throughout all the accompanying figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Implementation
[0019] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.
[0020] In the following description and claims, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0021] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.
[0022] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.
[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0024] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0025] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.
[0026] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0027] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.
[0028] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.
[0029] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.
[0030] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0031] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.
[0032] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0033] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2The examples are shown separately.
[0034] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0035] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).
[0036] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.
[0037] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.
[0038] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0039] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 to find a reference video block for the current video block, and can also search for reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0040] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0041] In one example, the motion estimation unit 204 may indicate a value to the video decoder 300 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0042] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0043] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0044] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0045] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0046] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0047] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0048] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0049] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.
[0050] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0051] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0052] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.
[0053] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0054] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.
[0055] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge pattern. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge pattern" may refer to deriving motion information from spatially or temporally neighboring blocks.
[0056] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. Identifiers for the interpolation filters used at sub-pixel precision can be included in the syntax elements.
[0057] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0058] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks for encoding the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.
[0059] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0060] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0061] Some exemplary embodiments of this disclosure will be described in detail below. It should be understood that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section only. Furthermore, although some embodiments are described with reference to multi-functional video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps corresponding to de-encoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.
[0062] 1. Brief Overview This disclosure relates to video codec technology. Specifically, it concerns cross-component coding and decoding in image / video codecs. It can be applied to existing video codec standards such as HEVC, VVC, etc. It can also be applied to future video codec standards or video codecs.
[0063] 2 Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established in 2015 by VCEG and MPEG. JVET meetings are held quarterly, and the new video codec standard was officially named Multifunctional Video Codec (VVC) at the April 2018 JVET meeting, where the first version of the VVC Test Model (VTM) was also released. The VVC working draft and the VTM test model are subsequently updated after each meeting. The VVC project achieved technical completion (FDIS) at a meeting in July 2020.
[0064] 2.1 Intra-frame prediction In intra-frame prediction, the minimum chroma intra-frame prediction unit (SCIPU) constraint in the VVC is removed. Additionally, the VPDU constraint used to reduce CCLM prediction latency is also removed.
[0065] 2.1.1 Multi-model LM (MMLM) The VVC-included CCLM (JVET-D0110) is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, reconstructed neighboring samples are classified into two classes using a threshold that is the average of the reconstructed luminance neighboring samples. A linear model for each class is derived using the least mean square (LMS) method. For the CCLM mode, the linear model is also derived using the LMS method. A slope adjustment is applied to both the cross-component linear model (CCLM) and the multi-model LM predictions. This adjustment tilts the linear function that maps luminance values to chrominance values relative to a center point determined by the average luminance value of the reference samples.
[0066] 2.1.1.1 Slope Adjustment of CCLM CCLM uses a two-parameter model to map luminance values to chrominance values. The slope parameter "a" and the bias parameter "b" define the mapping as follows: chromaVal = a lumaVal + b The slope parameter "u" is adjusted via signal transmission to update the model to the following form: chromaVal = a' lumaVal + b' in a' = a + u b' = b - u y r . Through this selection, the mapping function revolves around a value with brightness y. r The points are tilted or rotated. The average of the reference brightness samples used in model creation is taken as y. r This is to provide meaningful modifications to the model. The following diagram illustrates the process.
[0067] Figure 4 The diagram illustrates the effect of the slope adjustment parameter "u". Left: Model created using the current CCLM. Right: Updated model as proposed.
[0068] Implementation The slope adjustment parameter is provided as an integer between -4 and 4 (inclusive) and is transmitted via signal in the bitstream. The unit of the slope adjustment parameter is 1 / 8 of the chroma sample value per luminance sample value. th (For 10-bit content).
[0069] The CCLM model (“LM_CHROMA_IDX” and “MMLM_CHROMA_IDX”) is adjusted to apply to reference samples used on both the top and left sides of the block, but not to the “one-sided” mode. This choice is based on a trade-off between encoding / decoding efficiency and complexity.
[0070] When slope adjustment is applied to a multi-mode CCLM model, both models can be adjusted, so at most two slope updates are transmitted via signal transmission for a single chroma block.
[0071] Encoder method The proposed encoder method performs a SATD-based search for the optimal slope update for Cr and a similar SATD-based search for Cb. If either parameter results in a non-zero slope adjustment, the combined slope adjustment pair (SATD-based update for Cr, SATD-based update for Cb) is included in the list of RD checks for TU.
[0072] 2.1.2 Gradient PDPC In VVC, PDPC may not be applied in some scenarios due to the unavailability of secondary reference samples. In these cases, gradient-based PDPC (JVET-Q0391), extended from the horizontal / vertical mode, is applied. The PDPC weights (wT / wL) and nScale parameter, used to determine the decay of PDPC weights relative to the distance from the left / top boundary, are set to the corresponding parameters in the horizontal / vertical mode, respectively. Bilinear interpolation is applied when the secondary reference sample is located at the fractional sample position.
[0073] 2.1.3 Secondary MPM As described in JVET-D0114, a secondary MPM list is introduced. The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list includes 16 entries. First, a general MPM list with 22 entries is constructed. Then, the first 6 entries from this general MPM list are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the planar mode. The remaining entries consist of the intra-modes of the left (L), top (A), bottom left (BL), top right (AR), and top left (AL) neighboring blocks, the directional modes with offsets added from the first two available directional modes of the neighboring blocks, and the default mode.
[0074] If the CU block is vertical, the order of the neighboring blocks is A, L, BL, AR, AL; otherwise, the order is L, A, BL, AR, AL.
[0075] Figure 5 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown.
[0076] First, the PMPM flag is parsed. If it is equal to 1, the PMPM index is parsed to determine which entry in the PMPM list to select. Otherwise, the SMPPM flag is parsed to determine whether to parse the SMPM index or the remaining patterns.
[0077] 2.1.4 Reference Sample Interpolation and Smoothing for Intra-Frame Prediction As described in JVET-D0119, a 4-tap cubic interpolation filter is replaced with a 6-tap cubic interpolation filter to derive the predicted samples from the reference samples.
[0078] For reference sample filtering, a 6-tap Gaussian filter is applied to larger blocks (W>= 32 and H>= 32), otherwise the existing VVC 4-tap Gaussian interpolation filter is applied. The extended intra-frame reference samples are derived using a 4-tap interpolation filter instead of nearest-neighbor rounding.
[0079] 2.1.5 Decoder-side Intra-Frame Mode Derivation (DIMD) When applying DIMD, two intra-frame modes are derived from reconstructed neighboring samples, and these two predictions are combined with planar mode predictions as described in JVET-O0449, where weights are derived from gradients. Division operations in weight derivation are performed using the same lookup table (LUT)-based integerization scheme used by CCLM. For example, division operations in direction calculation.
[0080] The following LUT-based method is used for calculation: x = Floor(Log2(Gx)) normDiff=((Gx<<4)>>x)&15 x+=(3 + (normDiff!=0) ? 1 : 0) Orient = (Gy (DivSigTable[ normDiff ]|8) + (1<<(x-1)))>>x in DivSigTable
[16] = { 0, 7, 6, 5 ,5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}. The derived intra-frame modes are included in the main list of most probable intra-frame modes (MPMs), so the DIMD process is performed before the MPM list is built. The main derived intra-frame modes of the DIMD block are stored along with the block and used for the construction of the MPM list of neighboring blocks.
[0081] 2.1.5.1 DIMD Chroma Mode The DIMD chroma mode uses the DIMD derivation method to derive the chroma intra-prediction mode for the current block based on neighboring reconstructed Y, Cb, and Cr samples in the second nearest row and column. Specifically, the horizontal and vertical gradients are calculated for each co-located reconstructed luma sample and the reconstructed Cb and Cr samples of the current chroma block to construct the HoG. The intra-prediction mode with the largest histogram amplitude value is then used to perform chroma intra-prediction for the current chroma block.
[0082] Figure 6 The neighboring reconstructed samples used for the DIMD chromaticity mode are shown.
[0083] When the intra-prediction mode derived from the DIMD chroma mode is the same as the intra-prediction mode derived from the DM mode, the intra-prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode. A CU level flag is transmitted via signal transmission to indicate whether the proposed DIMD chroma mode is applied.
[0084] 2.1.6 Fusion of Chroma Intra-Frame Prediction Modes The DM mode and the four default modes can be merged with the MMLM_LT mode, as shown below:
[0085] in These are predicted values obtained by applying a non-LM model. These are predicted values obtained by applying the MMLM_LT mode, and This is the final predicted value for the current chroma block. Two weights. and Determined by the intra-prediction mode of adjacent chroma blocks, and the shift is set to equal to 2. Specifically, when both the upper adjacent block and the left adjacent block are encoded and decoded in LM mode, { }={1, 3}; When both the upper adjacent block and the left adjacent block are encoded and decoded in non-LM mode, { }={3, 1}; otherwise, { }={2, 2}.
[0086] For syntax design, if a non-LM mode is selected, a flag is transmitted via signaling to indicate whether fusion is applied. This method is only applicable to I-stripes.
[0087] 2.1.7 Intra-frame template matching Intra-frame template matching prediction (Intra-frame TMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed portion of the current frame. The L-shaped template of this best prediction block matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then transmits the use of this mode via signal transmission and performs the same prediction operation on the decoder side.
[0088] By using the L-shaped causal nearest neighbors of the current block with Figure 7 Another block in a predefined search region is matched to generate the predicted signal. This predefined search region includes: R1: Current CTU R2: Top Left CTU R3: Above CTU R4: Left CTU The sum of absolute differences (SAD) is used as the cost function.
[0089] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.
[0090] Set the dimensions (SearchRange_w, SearchRange_h) of all regions to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a BlkW SearchRange_h = a BlkH in' ' is a constant that controls the trade-off between gain and complexity. In practice, ' 'Equals 5.'
[0091] Figure 7 The intra-frame template matching search area used is shown.
[0092] To accelerate the template matching process, the search range of all search regions is downsampled by a factor of 2. This results in a reduction of 4 in the template matching search. After finding the best match, a refinement process is performed. Refinement is accomplished by a second template matching search around the best match with the reduced range. The reduced range is defined as min(BlkW, BlkH) / 2.
[0093] Enable intra-frame template matching for CUs with width and height dimensions less than or equal to 64. The maximum CU size for intra-frame template matching is configurable.
[0094] When DIMD is not used in the current CU, intra-template matching prediction mode is transmitted at the CU level via a dedicated flag.
[0095] 2.1.7.1 Block Vector Candidates Derived from Intra-Frame TMP for IBC In this method, the block vector (BV) derived from intra template matching prediction (intra-TMP) is used for intra-block copying (IBC). The intra-TMP BVs of the stored neighboring blocks, together with the IBC BVs, are used as spatial BV candidates in the construction of the IBC candidate list.
[0096] Intra-frame TMP block vectors are stored in the IBC block vector buffer, and the current IBC block can use both the IBCBV of neighboring blocks and the intra-frame TMP BV as BV candidates for the IBC BV candidate list, such as... Figure 8 As shown. Figure 8The use of intra-frame TMP block vectors for IBC blocks is shown.
[0097] Intra-frame TMP block vectors are added to the IBC block vector candidate list as spatial candidates.
[0098] 2.1.8 Fusion for Template-Based Intra-Frame Mode Inference (TIMD) For each intra-prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The two intra-prediction modes with the smallest SATD are selected as TIMD modes. After applying the PDPC procedure, these two TIMD modes are fused using weights, and this weighted intra-prediction is used for the current CU encoding and decoding. Position-dependent intra-prediction combination (PDPC) is included in the derivation of the TIMD modes.
[0099] The costs of the two selected modes are compared with a threshold, and a cost factor of 2 is applied during the test, as shown below: costMode2<2 costMode1. If the condition is true, then fusion is applied; otherwise, only mode1 is used.
[0100] The weights of the patterns are calculated from their SATD costs as follows: weight1 = costMode2 / (costMode1+ costMode2) weight2 = 1 - weight1 Division is performed using the same lookup table (LUT)-based integerization scheme used by CCLM.
[0101] 2.1.9 Intra-frame prediction fusion This intra-frame prediction method derives the predicted samples as a weighted combination of multiple predicted values generated from different reference rows. In this process, multiple intra-frame predicted values are generated and then fused by a weighted average. The process of deriving the predicted values to be used in the fusion process is described below: • For the intra-angle prediction mode in the single-mode case including TIMD and DIMD, the proposed method improves upon the representation as... Intra-prediction is derived by weighting the intra-prediction obtained from multiple reference rows, where It is an intra-frame prediction from the default reference line, and This is a prediction from the row above the default reference row. The weights are set to... and .
[0102] • For TIMD modes with hybrid patterns, Used in the first mode ( ),and Used in the second mode ( ).
[0103] • For DIMD patterns with a mixture, the number of forecasts selected for weighted averaging increases from 3 to 6.
[0104] When the intra-frame prediction mode has a non-integer slope (requiring reference sample interpolation) and the block size is greater than 16, the intra-frame prediction fusion method is applied to the luma block, used in conjunction with MRL, but not to blocks encoded / decoded by ISP. In the method studied in subtest a, PDPC is applied to the intra-frame prediction mode using the reference line closest to the current block.
[0105] 2.1.10 Combination of CIIP with TIMD and TM Merge In CIIP mode, prediction samples are generated by weighting the inter-frame prediction signal using CIIP-TM Merge candidate prediction and the intra-frame prediction signal using the intra-frame prediction mode derived from TIMD. This method is only applied to codec blocks with an area less than or equal to 1024.
[0106] The TIMD derivation method is used to derive intra-prediction modes in CIIP. Specifically, the intra-prediction mode with the smallest SATD value in the TIMD mode list is selected, and this intra-prediction mode is mapped to one of 67 regular intra-prediction modes.
[0107] Additionally, it is proposed that if the derived intra-prediction mode is an angle mode, the weights (wIntra, wInter) for the two tests should be modified. For near-horizontal mode (2 <= angle mode index < 34), the current block is vertically partitioned; for near-vertical mode (34 <= angle mode index <= 66), the current block is horizontally partitioned.
[0108] For different sub-blocks (wIntra, wInter) such as Figure 9 As shown.
[0109] Table 1. Modified weights used for angle mode
[0110] Using CIIP-TM, a CIIP-TM Merge candidate list is constructed for the CIIP-TM pattern. The Merge candidates are refined through template matching. CIIP-TM Merge candidates are also reordered as regular Merge candidates using the ARMC method. The maximum number of CIIP-TM Merge candidates is 2.
[0111] 2.1.11 Expanded Multi-Reference Row (MRL) List The MRL list in VVC is expanded to include more reference lines for intra-frame prediction. The expanded reference line list consists of line indices {1, 3, 5, 7, 12}. For template-based intra-frame pattern derivation (TIMD), only the first two reference line candidates (i.e., {1, 3}) are used, instead of the complete MRL candidate list. Figure 10 An expanded list of MRL candidates is shown.
[0112] 2.1.12 Template-based multi-reference row intra-frame prediction Template-based multi-reference line intra-frame prediction (TMRL) mode combines reference lines and prediction modes, using template matching to construct a list of candidate combinations. The indexes of the candidate combination list are encoded / decoded to indicate which reference line and prediction mode are used in the current block. Regular multi-reference line (MRL) for non-TIMD portions is replaced by TMRL mode.
[0113] The TMRL mode expands the reference line candidate list and the intra-prediction mode candidate list. The expanded reference line candidate list is {1, 3, 5, 7, 12}. The restrictions on the top CTU line remain unchanged. The size of the intra-prediction mode candidate list is 10. The construction of the intra-prediction mode candidate list is similar to that of MPM, except that planar modes are excluded from the intra-prediction mode candidate list. If the DC mode is not included, the DC mode is added after the modes of the 5 neighboring PUs and the DIMD mode, and has a range from... arrive An incremental angle mode (compared to existing angle modes in the intra-prediction mode candidate list) has been added.
[0114] The TMRL candidates are constructed as follows. There are 5 x 10 = 50 combinations of expanded reference lines and allowed intra-prediction modes for the block. Since the expanded reference lines start from reference line 1, the region covered by reference line 0 is used for template matching. The calculation between prediction (generated from the 50 combinations) and reconstruction is performed within the template region (see...). Figure 11 The SAD cost on the TMRL is calculated. The 20 combinations with the lowest SAD cost are selected in ascending order to form the TMRL candidate list.
[0115] For TMR signaling, instead of directly encoding and decoding the reference line and intra-frame mode, the index of the TMRL candidate list is encoded and decoded to indicate which combination of reference line and prediction mode is used to encode and decode the current block.
[0116] 2.1.13 Convolutional Cross-Component Intra-Frame Prediction Model In this method, a convolutional cross-component model (CCCM) is applied to predict chroma samples from reconstructed luminance samples, similar to what is done by the current CCLM model. As with CCLM, when chroma downsampling is used, the reconstructed luminance samples are downsampled to match a lower-resolution chroma grid. Similar to CCLM, top, left, or top and left reference samples are used as templates for model derivation.
[0117] Furthermore, similar to CCLM, there are options for single-model or multi-model variants of CCCM. The multi-model variant uses two models: one model is derived for samples above the average luminance reference value, and the other model is for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.
[0118] 2.1.13.1 Convolution Filter The convolutional 7-tap filter consists of a 5-tap plus-shaped spatial component, a nonlinear term, and a bias term. The input of the 5-tap spatial component of the filter consists of the center (C) luminance sample that is in the same position as the chrominance sample to be predicted, and its upper / north (N), lower / south (S), left / west (W), and right / east (E) neighbors, as shown below. Figure 12 The spatial portion of the convolution filter is shown.
[0119] The nonlinear term P is represented as the square of the center luminance sample C and scaled to the range of sample values for the content: P = (C C + midVal )>>bitDepth That is, for 10 bits of content, it is calculated as: P = (C C + 512)>>10 The bias term B represents the scalar offset between the input and output (similar to the offset term in CCLM) and is set to an intermediate chroma value (512 for 10-bit content).
[0120] The output of the filter is calculated as the filter coefficients c. i The convolution with the input values is then limited to the range of valid chromaticity samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B 2.1.13.2 Calculation of Filter Coefficients Filter coefficients c i It is calculated by minimizing the MSE between the predicted chromaticity samples and the reconstructed chromaticity samples in the reference region. Figure 13 The reference region is shown, consisting of six rows of chroma samples above and to the left of the PU. The reference region extends one PU width to the right of the PU boundary and one PU height below the PU boundary. The region is adjusted to include only available samples. The expansion of the region shown in blue is necessary to support the "edge samples" of the plus-shaped spatial filter and is filled in unavailable areas.
[0121] MSE minimization is performed by calculating the autocorrelation matrix for the luminance input and the cross-correlation vector between the luminance input and the chrominance output. The autocorrelation matrix is decomposed using LDL, and the final filter coefficients are calculated using inverse substitution. This process largely follows the calculation of ALF filter coefficients in ECM; however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations.
[0122] The autocorrelation matrix is calculated using reconstructed values of luma and chroma samples. These samples are full-range (e.g., between 0 and 1023 for 10-bit content), resulting in relatively large values in the autocorrelation matrix. This requires high-bit-depth operations during model parameter computation. A proposed solution is to remove a fixed offset from the luma and chroma samples in each PU for each model. This reduces the magnitude of the values used in model creation and allows for a reduction in the precision required for fixed-point arithmetic. As a result, a 16-bit decimal precision is proposed instead of the 22-bit precision of the original CCCM implementation.
[0123] For simplicity, reference sample values immediately outside the top-left corner of the PU are used as offsets (offsetLuma, offsetCb, and offsetCr). The sample values used in both model creation and final prediction (i.e., luminance and chromaticity in the reference region and luminance in the current PU) are reduced by these fixed values, as follows: C' = C – offsetLuma N' = N – offsetLuma S' = S – offsetLuma E' = E – offsetLuma W' = W – offsetLuma P' = nonLinear(C') B = midValue = 1<<(bitDepth - 1) Furthermore, the chromaticity values are predicted using the following equations, where offsetChroma for the Cr and Cb components is equal to offsetCr and offsetCb, respectively: predChromaVal = c0C' + c1N' + c2S' + c3E' + c4W' + c5P' + c6B +offsetChroma To avoid any additional sample-level operations, the luminance offset is removed during luminance reference sample interpolation. For example, this can be done by replacing the rounding term used in luminance reference sample interpolation with an updated offset that includes both the rounding term and offsetLuma. The chromaticity offset can be removed by directly subtracting it from the reference chromaticity samples. Alternatively, the effect of the chromaticity offset can be removed from the cross-component vectors, yielding the same result. To add the chromaticity offset back to the output of the convolution prediction operation, it is added to the bias term of the convolutional model.
[0124] The calculation of CCCM model parameters requires division operations. Division operations are not always considered implementation-friendly. Division operations are replaced by multiplication (using scaling factors) and shift operations, where the number of scaling factors and shifts is calculated based on the denominator, similar to the method used in CCLM parameter calculation.
[0125] 2.1.13.3 Gradient Linear Model For the YUV 4:2:0 color format, the Gradient Linear Model (GLM) method can be used to predict chromaticity samples from luminance sample gradients. Two modes are supported: two-parameter GLM mode and three-parameter GLM mode.
[0126] Compared to CCLM, instead of downsampling luminance values, two-parameter GLM uses the gradient of luminance samples to derive a linear model. Specifically, when applying two-parameter GLM, the input to the CCLM process (i.e., downsampled luminance samples) is... Gradient of brightness sample points Replacement. Other parts of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged.
[0127]
[0128] In a three-parameter GLM, chromaticity samples can be predicted based on luminance sample gradients with different parameters and downsampled luminance values. The model parameters of the three-parameter GLM are derived from 6 row and column adjacent samples using an MSE minimization method based on LDL decomposition, as used in CCCM.
[0129]
[0130] For signaling, when CCLM mode is enabled for the current CU, a flag is transmitted via signaling to indicate whether GLM is enabled for both Cb and Cr components; if GLM is enabled, another flag is transmitted via signaling to indicate which of the two GLM modes is selected, and a syntax element is further transmitted via signaling to select one of the four gradient filters for gradient calculation.
[0131] • Enable four gradient filters for GLM, such as Figure 14 As shown.
[0132] 2.1.13.4 Bitstream Signaling The use of the mode utilizes PU-level flags encoded and decoded by CABAC, which are transmitted via signaling. A new CABAC context is included to support this. When signaling is involved, CCCM is considered a submode of CCLM. That is, if the intra-frame prediction mode is LM_CHROMA, the CCCM flag is only transmitted via signaling.
[0133] 2.1.14 Spatial Geometric Partitioning Model (SGPM) SGPM is an intra-frame mode of inter-frame coding / decoding tools similar to GPM, where two prediction components are generated from the intra-frame prediction process. In this mode, a candidate list is constructed, where each entry contains a segmentation partition and two intra-frame prediction modes, such as... Figure 15 As shown, 26 segmentation modes and 3 intra-frame prediction modes were used to form a combination. The length of the candidate list was set to 16. The selected candidate indices were transmitted via signaling.
[0134] The list is reordered using a template. Figure 16 The SAD between the template's prediction and reconstruction was used for sorting. The template size was fixed at 1.
[0135] For each segmentation pattern, the same intra-frame and inter-frame GPM list is used to derive the IPM list for each segment. The IPM list size is set to 3. In the list, the pattern derived by TIMD is replaced by two derived patterns with horizontal and vertical directions.
[0136] The SGPM pattern is applied with limited block sizes: 4 <= width <= 64, 4 <= height <= 64, width < height. 8. Height < Width 8. Width Height >= 32.
[0137] Adaptive blending has also been used in spatial GPM, where Figure 17 The mixing depth τ shown is derived as follows: • If min(width, height) == 4, 1 / 2 τ is selected.
[0138] • Otherwise, if min(width, height) == 8, τ is selected.
[0139] • Otherwise, if min(width, height) == 16, 2τ is selected.
[0140] • Otherwise, if min(width, height) == 32, 4τ is selected.
[0141] • Otherwise, 8τ is selected.
[0142] 2.1.15 Nonlocal cross-component prediction Cross-component prediction (CCP), including CCLM, CCCM, and their variants, is employed in ECM to leverage cross-component correlations. With CCLM or CCCM, training samples are always adjacent to the current block. However, the cross-component relationships of the current block can be more relevant to cross-component relationships in non-local regions.
[0143] A nonlocal cross-component prediction method is proposed to enhance CCP by gaining more advantages from nonlocal regions.
[0144] Method #1: A Non-Adjacent Cross-Component Prediction (NA-CCP) model is proposed. Using the NA-CCP model, samples from regions not adjacent to the current block can be used to derive the CCCM model for the current block. A candidate region list with six candidates is constructed by sequentially examining potential 8×8 regions. If an examined region is available, it is added to the candidate region list. The top-left position of the potential 8×8 region is pre-determined as {(-xStep, 0), (0, -yStep), (xStep, -yStep), (-xStep, yStep), (-xStep, -yStep), (-2...}. xStep, 0), (0, -2 yStep), (-2 xStep, 2 yStep), (2 xStep, -2 yStep), (-2 xStep, yStep), (xStep, -2 yStep),(-2 xStep, -yStep), (-xStep, -2 yStep), (-2 xStep, -2 yStep), (-xStep / 2, 0),(0, -yStep / 2), (xStep / 2, -yStep / 2), (-xStep / 2, yStep / 2), (-xStep / 2, -yStep / 2)}, where xStep = Max(width, 16), yStep = Max(height, 16). Figure 18 Some possible locations of the candidate regions are shown.
[0145] A signal transmission flag is used to indicate whether NA-CCP is applied to the chroma block. If NA-CCP is applied, a signal transmission index is used to indicate which candidate in the candidate region list was used to derive the CCCM model.
[0146] Method #2: A history-based cross-component prediction (H-CCP) model is proposed. H-CCP, similar to the HMVP table, is used to maintain the H-CCLM and H-CCCM tables. After decoding a block encoded / decoded using CCLM or CCCM, the corresponding table is updated. In the H-CCP implementation, the size of either the H-CCLM or H-CCCM table is 6. If the current block is encoded / decoded using CCLM or CCCM mode, a flag is signaled to indicate whether H-CCP is applied. If H-CCP is used, an index is further signaled to indicate which candidate model from the H-CCLM or H-CCCM table is selected.
[0147] 2.1.16 Cross-component Merge Mode for Chroma Intra-Frame Coding / Decoding Cross-component prediction (CCP) models, including Cross-component Linear Model (CCLM), Convolutional Cross-component Model (CCCM), and Gradient Linear Model (GLM), are employed in ECM to leverage cross-component correlations. A new cross-component merge (CCMerge) mode is proposed as a novel CCP mode. The cross-component model parameters of the current chroma block encoded using CCMerge can be inherited from neighboring blocks encoded using CCP. Through CCMerge, CCP can be more efficient and has less signaling overhead.
[0148] In CCMerge, the final cross-component model parameters for the current chroma block can be inherited from its spatially adjacent and non-adjacent neighbors or the default model. A list is created that includes CCP models from spatially adjacent and non-adjacent neighbors encoded and decoded in CCLM, MMLM, CCCM, GLM, chroma blending, and CCMerge modes. After including neighboring CCP models, the default model is further included to fill any remaining empty positions in the list. To avoid including redundant CCP models in the list, a deduplication operation is applied. More details are described below.
[0149] • Airspace adjacent neighbor candidates The positions of adjacent candidates in the airspace are as follows Figure 19 As shown, the airspace candidates are included in the following order: B1 -> A1 -> B0 -> A0 -> B2.
[0150] • Airspace not adjacent to neighboring candidates After examining all spatially adjacent neighbors, spatially non-adjacent neighbor candidates are considered. In the current ECM design, two sets of spatially non-adjacent neighbor candidates are obtained in inter-frame merge mode. In the proposed method, the positions and inclusion order of the first set of spatially non-adjacent neighbor candidates are used.
[0151] • CCLM candidates with default scaling parameters If the list is insufficient after including both spatially adjacent and non-adjacent candidates, CCLM candidates with default scaling parameters are considered. The default scaling parameters are {0, 1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8}, and the offset parameters are derived based on the selected default scaling parameters, the average neighboring reconstructed luminance sample value (Yavg), and the average neighboring reconstructed Cb / Cr sample value (Cavg).
[0152] 2.1.16.1 Merge Model Candidates When merging CCLM candidates, only the scaling parameter is inherited. The offset parameter is derived using the inherited scaling parameter, Yavg, and Cavg.
[0153] When merging MMLM candidates, the scaling parameter and classification threshold are inherited. The offset parameter in each class is derived based on the inherited classification threshold and the Yavg and Cavg values in each class. If no neighboring reconstructed samples are available in a class, the offset parameter is directly inherited from the candidate.
[0154] When merging CCCM candidates, all convolutional parameters, offsets (i.e., offsetLuma, offsetCb, and offsetCr), and classification thresholds are inherited.
[0155] When merging GLM candidates, if the GLM candidate is a three-parameter GLM mode, all gradient mode indices and model parameters are inherited; otherwise, if the GLM candidate is a two-parameter GLM mode, the offset parameters are derived by using the inherited scaling parameters, Yavg, and Cavg.
[0156] When merging chroma blending candidates, the derived MMLM parameters are inherited and used as Merge MMLM candidates.
[0157] For a CCMerge block, if its Merge candidate mode is CCLM, MMLM, CCCM, or GLM, the Merge candidate mode is stored as the propagation mode of the current chroma block; otherwise, if its Merge candidate mode is chroma blending, the propagation mode is set to MMLM. How CCP parameters are inherited or derived when merging CCMerge candidates depends on the propagation mode of the CCMerge candidate, as described in the five paragraphs above.
[0158] 2.1.16.2 Signaling An additional flag indicating whether CCMerge is used is transmitted via semaphore after the `cclm_mode_flag` syntax element. If CCMerge is used, candidate indices are also transmitted via semaphore. The candidate indices transmitted via semaphore are shared across the Cb / Cr color components. Currently, the maximum number of allowed candidates is set to the default value of 6. If the maximum number of allowed candidates is modified to 1, candidate indices do not need to be transmitted via semaphore. Each bit of the candidate index is encoded and decoded using a separate context.
[0159] 2.1.17 Directional Plane Mode Two additional planar modes are used, where either horizontal interpolation only or vertical interpolation only is used to obtain the predicted samples.
[0160] For the planar horizontal pattern, horizontal linear interpolation is performed only based on the left and upper right reference points to predict the current point as:
[0161] For the planar vertical mode, vertical linear interpolation is performed only based on the upper reference point and the lower left reference point to predict the current sample point as:
[0162] Transform kernel selection for horizontal and vertical planar modes, as follows: Figure 20 As shown. If the intra-prediction mode of the current block is planar vertical mode, the horizontal intra-prediction mode is used to derive the transform kernels in the MTS set and LFNST set. Furthermore, if the intra-prediction mode of the current block is planar horizontal mode, the vertical intra-prediction mode is used to derive the transform kernels in the MTS set and LFNST set.
[0163] 2.1.18 Direct block vectors for chroma blocks Direct block vectors are used for chroma blocks in a dual-tree stripe. When the chroma dual-tree is active, a signal transmission flag indicates whether the chroma blocks are being encoded / decoded using IBC mode. Figure 21One of the luma blocks in the five locations shown is encoded and decoded using IBC or intra-frame TMP mode, and its block vector is scaled and used as the block vector for the chroma block. Template matching is used to perform block vector scaling.
[0164] 2.1.19 Intra-frame prediction mode based on extrapolation filter (EFI mode) The proposed intra-frame prediction based on extrapolation filters is processed in two steps. First, using a predetermined template, the extrapolation filter coefficients are obtained from the reconstructed pixels in the neighboring blocks of the current block. Second, extrapolation generates predicted values position by position within the current block, from the top left to the bottom right.
[0165] 2.1.19.1 Search for average, minimum, and maximum values Similar to CCCM mode, the average value should be removed when the input is fed to the EIP filter. The DC mode value of the current block is used as the average value for the EIP prediction. The minimum and maximum values are searched from the reconstructed pixels in the reconstructed region with thirteen columns and thirteen rows.
[0166] 2.1.19.2 Calculation of Filter Coefficients like Figure 22 As shown, three types of reconstructed regions and three filter shapes are proposed. The three types of reconstructed regions defined consist of thirteen columns or rows of reconstructed pixels. When the current block is used for prediction using the proposed EIP pattern, the decoder decodes the relevant syntax elements to determine the selected type of reconstructed region and filter shape for the current block. Figure 23 The diagram illustrates three types of filter shapes, each with fifteen inputs and producing one output.
[0167] The selected filter slides across the selected reconstructed region in a one-pixel step to collect input and output samples for the EIP. The autocorrelation matrix and cross-correlation vector are constructed while removing the mean from the input and output samples. The EIP coefficients are then obtained using the same method as in CCCM.
[0168] 2.1.19.3 Prediction of the current block EIP mode predicts the current block position by position, such as Figure 24 As shown.
[0169] For a position located at the top left of the current block, the input to the EIP filter is the reconstructed sample.
[0170] For a location on the boundary of the current block, part of the input to the EIP filter is the reference sample, and part of the input to the EIP filter is the previously predicted sample.
[0171] For other locations in the current block, the input to the EIP filter is the previously predicted sample.
[0172] To reduce prediction error, the searched minimum and maximum values are applied to limit the output range of each predicted value.
[0173] in It is the predicted value at (x, y) in the current block. It is the minimum and maximum value searched from thirteen reconstructed columns and thirteen reconstructed rows. It is the derived EIP filter. coefficient, It is a reconstructed or predicted value used for the current location. The average value is calculated using the DC forecasting model.
[0174] 2.2 Cross-component residual model (CCRM) for inter-frame prediction (also known as inter-frame CCCM) It is proposed that when blocks use inter-frame prediction or intra-block copy (IBC), a cross-component residual model (CCRM) is applied to predict chrominance samples from reconstructed luminance samples. Figure 25 The decoder side of the method is shown. A cross-component filter is derived using prediction blocks for both luma and chroma. The derived filter is applied to the reconstructed luma block and mixed with the chroma prediction block to produce the final chroma prediction block. During the mixing process, the filtered, reconstructed luma block uses a mixing weight of 0.75, and the chroma prediction block uses a mixing weight of 0.25.
[0175] 2.2.1 Calculation of Convolution Filter and Filter Coefficients The proposed 8-tap filter consists of 6 spatial brightness samples, a nonlinear term, and a bias term. For example... Figure 23 As shown, the spatial luminance samples (L0, ..., L5) are obtained from the luminance grid without downsampling, selecting the 6 luminance samples closest to the chromaticity position C. The predicted chromaticity value is obtained as follows: predChromaVal = c0L0+ c1L1 + c2L2 + c3L3 + c4L4 + c5L5 + c6nonlinear((L0+L3+1)>>1) + c7B, Where nonlinear is the nonlinear operator of CCCM, and B is the bias.
[0176] The filter coefficients were derived using the division-free Gaussian elimination method of ECM, and the necessary offset was applied to the samples prior to the filter derivation.
[0177] For the derivation of the filter coefficients, a maximum of 256 chromaticity samples are used.
[0178] The offset of the ECM's division-free Gaussian elimination method (used to solve the CCRM filter coefficients) is obtained using a four-point average of the luminance and chrominance prediction blocks, where the four points correspond to the top left, top right, bottom left, and bottom right corners of the block.
[0179] 2.2.2 Bitstream Signaling The use of modes utilizes the CABAC-encoded TU level flags transmitted via signaling. A new CABAC context is included to support this. If the TU's luminance Cbf is non-zero and the CU's predMode is MODE_INTER or MODE_IBC, the CCRM flags are transmitted via signaling only.
[0180] 2.2.3 Encoder Operation When the luminance Cbf is non-zero and the predMode of the CU is MODE_INTER or MODE_IBC, the encoder performs the RD decision in the transform selection loop for the chrominance component.
[0181] 2.3 On the cross-component model for residual encoding and decoding in image and video encoding and decoding 2.3.1 Issues related to cross-component models used for residual encoding and decoding There are several problems with existing video encoding and decoding technologies, and they will be further improved in order to achieve higher encoding and decoding gains.
[0182] 1. Several aspects of the video unit of CCRM encoding and decoding (such as filter terms, model type, and applied block type) can be further improved.
[0183] 2. The CCRM-estimated predictions will compete with the original inter-frame predictions. Ultimately, the residual with the lower cost is chosen. However, the concept of fusion can be incorporated to achieve better results.
[0184] 3. CCRM is applied to inter-blocks as long as the luminance component has a non-zero CBF. The CCRM on / off decision can be further designed.
[0185] 4. Currently, the CCRM model is applied, taking luminance reconstruction as input and estimated chromaticity prediction as output. However, estimated chromaticity predictions can be generated by adding residual blocks estimated by CCRM to chromaticity predictions without CCRM.
[0186] 2.3.2 Related Solutions The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0187] The term "video unit" or "code-decoder unit" can refer to a picture, strip, slice, code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0188] The term "block" can refer to code-decode tree block (CTB), code-decode tree unit (CTU), code-decode block (CB), CU, PU, TU, PB, and TB.
[0189] The terms "motion vector" or "block vector" can refer to the vector of horizontal and vertical displacement between the position of a reference block and the position of the current block. The reference block can be a video unit in a reference image within the RPL list. Alternatively, the reference block can be a video unit in the current image.
[0190] The term "LM" can refer to any linear regression-based method, such as CCLM, MMLM, CCCM, GL-CCCM, CCCM without downsampling, GLM, GLM with luminance values, etc. It can also be referred to as "Cross-Component Prediction (CCP)". CCP models can be used for intra-frame prediction, IBC prediction, or inter-frame prediction.
[0191] The term "CCLM" can refer to a single-model LM mode, which can be a single-model CCLM, a single-model CCCM, a single-model GL-CCCM, a single-model CCCM without subsampling, a single-model GLM, a single-model GLM with luminance values, a multi-model CCLM, a MMLM, a multi-model CCCM, a multi-model GL-CCCM, a multi-model CCCM without subsampling, a multi-model GLM, a multi-model GLM with luminance values, etc.
[0192] The term "MMLM" can refer to a multi-model LM mode, which can be multi-model CCLM, MMLM, multi-model CCCM, multi-model GL-CCCM, multi-model CCCM without downsampling, multi-model GLM, multi-model GLM with luminance values, etc.
[0193] The term "MFLM" can refer to a multi-filter LM mode, which can be MF-CCLM, MF-CCCM, MF-GLM, MF-CCRM, MF-CCCM for inter-frame, multi-filter IBC filter, multi-filter intraTMP filter, and / or variants of the mentioned mode, etc.
[0194] The term "CCCM" can refer to regular CCCM mode, GL-CCCM mode, CCCM without downsampling, CCRM, etc.
[0195] The term "GL-CCCM" can refer to a CCCM mode that takes into account the gradient and location of the samples involved.
[0196] The term "CCCM without downsampling" can refer to a CCCM mode that takes into account unsampled luminance samples.
[0197] The term "CCRM" can refer to residual encoding / decoding or derivation based on cross-component models. It can also refer to inter-frame / IBC prediction based on CCCM models (such as inter-frame / IBC CCCM). It can also refer to intra-frame prediction based on CCCM models (such as intra-frame CCCM). It can refer to the generation and application of cross-component models (such as luma-to-chroma prediction). It can also refer to the generation and application of same-component models (such as luma-to-luma prediction).
[0198] In this document, Cross Component Prediction (CCP) can refer to any cross component prediction method, such as any kind of CCLM / CCCM / GLM / GL-CCCM.
[0199] It should be noted that the following terms are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0200] 1) The residuals (and / or predictions) of chroma blocks can be derived based on cross-component models.
[0201] a. For example, the cross-component model can be a specific extrapolation filter (e.g., EIP, etc.).
[0202] b. For example, the cross-component model can be a specific interpolation filter (e.g., GLM, etc.).
[0203] c. For example, the cross-component model can be a specific convolutional filter (CCCM, GL-CCCM, CCCM without downsampling, CCRM, inter-frame CCCM, intra-frame CCCM, etc.).
[0204] d. For example, the cross-component model can be a specific linear filter (e.g., CCLM, MMLM, etc.).
[0205] 2) Cross-component models used for residual coding and decoding (e.g., CCRM) may not contain nonlinear terms.
[0206] a. For example, a cross-component model used for residual encoding and decoding may contain linear terms and / or bias terms, but not nonlinear terms.
[0207] 3) CCRM can be used for intra-frame or IBC blocks.
[0208] a. For example, it can be used for intra-frame or IBC blocks in intra-frame (such as I) stripes.
[0209] b. For example, it can be used in intra-frame or IBC blocks within inter-frame (such as B or P) stripes.
[0210] c. For example, it can also be used for single trees.
[0211] d. For example, it can also be used for two trees.
[0212] e. For example, in a single-tree I-strip, both luminance and chrominance are encoded and decoded by IBC (or IntraTMP). CCRM can be generated based on the reconstructed luminance and chrominance samples within a reference block retrieved / guided by block vectors, and a residual model is applied to estimate the reconstructed values of the chrominance samples in the current block.
[0213] f. For example, in a dual-tree system, luminance is encoded and decoded by IBC (or IntraTMP), while chrominance is encoded and decoded by intra-frame. CCRM can generate reconstructed luminance samples within a reference luminance block retrieved / guided by block vectors, as well as reconstructed chrominance samples that are co-located (e.g., at the same position) within that luminance block, and a residual model is applied to estimate the reconstructed values of the chrominance samples in the current block.
[0214] 4) CCRM can be used for chroma blocks encoded and decoded by DBV.
[0215] a. For example, a reference chroma block and its corresponding luma block can be identified based on the block vector of the chroma block encoded and decoded by DBV. These samples can be used as training samples for computation of the CCRM model.
[0216] b. For example, the derived CCRM model is applied to the reconstructed luminance signal of the DBV chromaticity block to produce the final chromaticity prediction.
[0217] 5) The CCRM model can be generated based on the correlation between the luminance reconstruction values and chrominance reconstruction values from neighboring / non-adjacent samples of the current block.
[0218] a. For example, alternatively, the CCRM model can be generated based on the correlation between the luminance reconstruction values and chrominance reconstruction values in a reference block of a reference image.
[0219] b. For example, alternatively, the CCRM model can be generated based on the correlation between the luminance reconstruction values and chrominance reconstruction values in a reference block in the current image.
[0220] 6) For example, the CCCM used for intra-frame prediction and the CCCM used for inter-frame prediction (e.g., CCRM) can share the same logic.
[0221] a. For example, both can follow the same logic to obtain training samples.
[0222] b. For example, both can follow the same logic to determine the training area.
[0223] 7) CCRM models can be generated based on unsampled luminance samples.
[0224] a. For example, CCRM model coefficients can be solved based on unsampled luminance samples from a reference region used as training samples.
[0225] b. For example, the CCRM model can be applied to chroma blocks, where the chroma prediction of the current chroma block is generated based on the non-downsampled luminance samples of the co-located luminance block.
[0226] 8) More than one CCRM model (e.g., MM-CCRM model, MF-CCRM model, CCRM Merge model) can be generated for blocks.
[0227] a. For example, the training samples of CCRM can be divided into more than one class (e.g., two classes), and each set of samples can contribute to a unique model. In this way, multiple models can be generated, each with its own filter coefficients. Each derived filter is applied to its corresponding set of luminance reconstructed signals to produce a final prediction value for the current chrominance sample belonging to the corresponding class.
[0228] i. For example, according to the multi-model CCRM (e.g., MM-CCRM) mode, training sample pairs of the luminance and chrominance sample pairs of the reference block (e.g., these training samples in the reference frame) can be divided into more than one category.
[0229] ii. For example, alternatively, training sample pairs (e.g., these training samples in the reference frame) of the brightness and chromaticity sample pairs of neighboring samples adjacent / non-adjacent to the reference block can be classified into more than one category, according to a multi-model CCRM (e.g., MM-CCRM) mode.
[0230] iii. For example, alternatively, training sample pairs (e.g., those training samples in the current frame) of the luminance and chrominance sample pairs of neighboring samples that are adjacent / non-adjacent to the current video cell can be classified into more than one class, according to a multi-model CCRM (e.g., MM-CCRM) mode.
[0231] iv. For example, in addition, by following the same criteria (e.g., by a threshold), the luminance samples in the current video unit are divided into more than one group, and for each luminance sample belonging to a category, the corresponding model can be applied to generate the model-estimated chrominance samples belonging to that group.
[0232] b. For example, multiple sets of training samples can be used to derive multiple models.
[0233] i. In one example, there are two groups where the distance between the training sample and the current sample is different.
[0234] c. For example, the threshold used to separate samples into different categories (e.g., category threshold) may depend on the values of samples within or near the training area.
[0235] i. For example, the training region can be a reference block of the current video unit (e.g., these training samples are in the reference frame).
[0236] 1. For example, a reference block can be derived based on a block vector.
[0237] 2. For example, a reference block can be derived based on motion vectors.
[0238] ii. For example, the threshold can be derived based on samples that are adjacent to or not adjacent to the reference block of the current video unit (e.g., these training samples are in the reference frame).
[0239] iii. For example, the threshold can be derived based on samples that are adjacent to or not adjacent to the current video unit (e.g., these training samples are in the current frame).
[0240] iv. For example, the threshold can be derived based on the average / median / intermediate operation of more than one sample point within or adjacent to the training region.
[0241] v. For example, category thresholds can be derived based on unsampled luminance sample values.
[0242] 1. Alternatively, the category threshold can be derived based on the downsampled brightness sample values.
[0243] a. For example, a K-tap (such as K=6) downsampling filter can be used to reduce K surrounding luminance samples to a single downsampled luminance sample value.
[0244] vi. For example, the category threshold can be derived based on the offset removal scheme.
[0245] 1. For example, the offset can be derived based on luminance samples located at fixed positions (such as the upper left or center) within the reference video unit.
[0246] 2. For example, the offset values calculated for category threshold derivation and CCRM model can be the same.
[0247] vii. For example, category thresholds can be derived based on the sub-block level.
[0248] viii. For example, category thresholds can be derived based on CU / PU / TU levels.
[0249] ix. For example, the category threshold can be calculated based on (downsampled or non-downsampled) brightness prediction samples.
[0250] x. For example, category thresholds can be derived based on luminance residual sample values.
[0251] 1. For example, for a second video unit (e.g., a sub-block) that does not have a non-zero residual, the predicted samples of such a video unit may not be included in the calculation of the category threshold for the first video unit.
[0252] a. For example, the second video unit may be a subset of the first video unit.
[0253] b. For example, the second video unit can be equal to the first video unit.
[0254] d. For example, MM-CCRM can be applied at the sub-block level.
[0255] i. For example, the size of the sub-block can be predefined.
[0256] 1. For example, the predefined sub-block size can be 16x16, or 32x32, etc.
[0257] 2. For example, predefined rules can be used to determine the sub-block size of the MM-CCRM for a specific video unit.
[0258] a. For example, the size of a sub-block can be adapted to the block dimensions (width and / or height) of the current video block.
[0259] b. For example, for a sub-block of a video unit encoded and decoded by MM-CCRM, a minimum number of chroma samples can be guaranteed.
[0260] ii. For example, if a video unit is larger than a predefined sub-block size, the video unit can be divided into more than one sub-block and MM-CCRM can be performed.
[0261] iii. For example, at least one sub-block of a video unit may have more than one CCRM model.
[0262] iv. For example, each sub-block (and its associated training region) can have its own category threshold.
[0263] 1. For example, the category threshold for a specific sub-block can be calculated based on the training sample values belonging to that sub-block.
[0264] a. For example, brightness training samples in a reference block can be used to calculate a category threshold.
[0265] v. For example, all sub-blocks (and their associated training regions) can share the same class threshold.
[0266] 1. For example, a category threshold can be calculated and used for all sub-blocks.
[0267] 2. For example, the category threshold for all sub-blocks in the current video unit can be calculated based on the training sample values of the current video unit.
[0268] 3. For example, the category threshold for all applicable sub-blocks in the current video unit can be calculated based on the training sample values of the current video unit.
[0269] a. For example, sub-blocks that do not contain non-zero residuals may not be counted.
[0270] vi. For example, each sub-block of a video unit can have its own training samples, and the training samples of a particular sub-block can be divided into more than one category.
[0271] 1. For example, training samples in the reference video unit of a reference image can be classified based on sub-blocks.
[0272] vii. For example, training samples from the current image can be classified into more than one group, but may not be divided into sub-blocks.
[0273] e. For example, MM-CCRM / MF-CCRM / CCRM Merge can be applied at the TU level (or PU / CU level).
[0274] i. For example, MM-CCRM / MF-CCRM / CCRM Merge can be applied based on TU / CU / PU (e.g., for MM-CCRM applications, TU / CU / PU may not be divided into sub-blocks).
[0275] ii. For example, whether to use a multi-model CCRM / MF-CCRM / CCRM Merge based on TU / PU / CU can be determined at the TU / PU / CU level.
[0276] 1. For example, a video unit (e.g., TU / PU / CU) may choose to use a sub-block-based CCRM (e.g., CCRM / MM-CCRM / MF-CCRM / CCRM Merge) or a TU / PU / CU-based CCRM (e.g., CCRM / MM-CCRM / MF-CCRM / CCRM Merge).
[0277] a. For example, decisions can be made at the TU / PU / CU level.
[0278] f. For example, whether and / or how to apply MM-CCRM (and / or CCRM / MF-CCRM / CCRM Merge) can be deduced based on encoding and decoding information at both the encoder and decoder sides (e.g., not through signal transmission).
[0279] i. In one example, it can be derived on the fly, for example, using information from previously encoded / reconstructed samples.
[0280] ii. For example, the determination of whether to use a sub-block-based CCRM / MM-CCRM / MF-CCRM / CCRM Merge or a TU / CU / PU level CCRM can be based on implicit deduction of codec information (e.g., not through signal transmission).
[0281] iii. For example, the determination of whether to use CCRM based on M1xM2 subblocks or CCRM based on N1xN2 subblocks can be based on the implicit derivation of encoding and decoding information (e.g., without signal transmission).
[0282] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.
[0283] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.
[0284] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.
[0285] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.
[0286] 5. For example, M1 != N1 and / or M2 != N2.
[0287] iv. For example, the determination of whether to use a sub-block-based MM-CCRM / MF-CCRM / CCRM Merge or a TU / CU / PU level MM-CCRM / MF-CCRM / CCRM Merge can be based on implicit deduction of codec information (e.g., not through signal transmission).
[0288] v. For example, the determination of whether to use the MM-CCRM / MF-CCRM / CCRM Merge based on M1xM2 sub-blocks or the MM-CCRM / MF-CCRM / CCRM Merge based on N1xN2 sub-blocks can be based on the implicit derivation of the encoding / decoding information (e.g., not through signal transmission).
[0289] 1. For example, M1 = 16 or 8 or 32 or TU / CU / PU.
[0290] 2. For example, M2 = 16 or 8 or 32 or TU / CU / PU.
[0291] 3. For example, N1 = 16 or 8 or 32 or TU / CU / PU.
[0292] 4. For example, N2 = 16 or 8 or 32 or TU / CU / PU.
[0293] 5. For example, M1 != N1 and / or M2 != N2.
[0294] vi. For example, the determination of whether to use SM-CCRM / MF-CCRM / CCRM Merge or MM-CCRM / MF-CCRM / CCRMMerge can be based on the implicit derivation of codec information (e.g., not through signal transmission).
[0295] vii. For example, determination can be based on a cost derived from the decoder.
[0296] 1. For example, the cost of decoder derivation can be calculated based on minimizing the SAD / SATD / SSE / MSE between the model estimated sample values and the true reconstructed sample values, where a sample can refer to at least one training sample among the training samples.
[0297] 2. For example, a method with lower cost can be chosen as the final method to be applied to the current video unit.
[0298] viii. For example, it can be determined that the information is based on a reference image.
[0299] 1. For example, the determination can be based on the POC distance between the current image and its reference image.
[0300] 2. For example, the determination can be based on a reference index.
[0301] g. Alternatively, whether and / or how MM-CCRM (and / or CCRM / MF-CCRM / CCRM Merge) can be applied can be transmitted via signaling in the bitstream.
[0302] i. For example, syntax elements (such as flags, indices, etc.) can be signaled conditionally based on whether the current block is coded by CCRM.
[0303] 1. For example, if a video unit is coded by CCRM, syntax elements (such as flags, indices, etc.) can be further signaled to indicate whether it is MM-CCRM / MF-CCRM / CCRM Merge.
[0304] ii. For example, syntax elements (such as flags, indices, etc.) can be signaled to indicate whether it is MM-CCRM based on sub-blocks or MM-CCRM / MF-CCRM / CCRM Merge based on TU / CP / PU.
[0305] iii. For example, syntax elements (such as flags, indices, etc.) can be signaled to indicate whether it is CCRM based on sub-blocks or CCRM / MF-CCRM / CCRM Merge based on TU / CP / PU.
[0306] iv. For example, syntax elements can be signaled conditionally based on the block dimensions (width and / or height).
[0307] 1. For example, if W H < T (such as T = 16 or 32), then it may not be signaled.
[0308] v. For example, syntax elements can be signaled depending on the residual / coefficients of the current luminance block.
[0309] 1. For example, whether to signal a syntax element can be conditional based on whether there is a residual (or non-zero coefficient) in the current luminance block.
[0310] 2. For example, whether to signal a syntax element can be conditional based on the distribution / number / value of the residual (or non-zero coefficient) in the current luminance block.
[0311] vi. For example, syntax elements can be signaled conditionally based on the prediction method of neighboring blocks.
[0312] 1. For example, it can be based on whether neighboring blocks (such as the left neighbor and / or the upper neighbor) use the CCRM / MM-CCRM / MF-CCRM / CCRM Merge mode.
[0313] vii. For example, the context model of a syntax element can depend on the coding information of neighboring blocks or the current block.
[0314] 1. For example, the context model can be derived based on whether neighboring blocks (such as the left neighbor and / or the upper neighbor) use the CCRM / MM-CCRM / MF-CCRM / CCRM Merge mode.
[0315] 2. For example, the context model can be derived based on whether the block dimensions of the current block satisfy specific conditions.
[0316] a. For example, if the current block (e.g., TU / PU / CU) is long or wide (e.g., W > a H, and / or H > b W, where W and H are the width and height of the current block, and a and b are predefined constants, such as a = b = 2), then the specified context model can be used.
[0317] h. For example, block restrictions can be applied to indicate the allowance of the MM-CCRM mode.
[0318] i. In one example, assuming that the width and height of the chrominance CU / PU / TU are represented as W and H, then MM-CCRM can be allowed when at least one of the following conditions is satisfied: 1. W H > T0 or W H >= T0 (e.g., T0 = 16 or 32 or 64 or 128) 2. W > T1, or, W >= T1 3. H > T2, or, H >= T2 4. Min(W, H) > T3, or, Min(W, H) >= T3 5. Max(W, H) < T4, or, Max(W, H) <= T4 6. W < T5 H, or, W <= T5 H 7. W > T6 H, or, W >= T6 H 8. H < T7 W, or, H <= T7 W 9. H > T8 W, or, H >= T8 W 10. W H < T9, or W H <= T9 ii. In one example, for a block enabled for a specific tool (e.g., enabling affine motion compensation), MM-CCRM can be prohibited.
[0319] i. For example, video units encoded and decoded by CCRM can always use multi-model CCRM.
[0320] i. Alternatively, video units encoded and decoded by CCRM can use either a single-model CCRM or a multi-model CCRM.
[0321] 9) Chromaticity Cb and Cr can share a single CCRM.
[0322] a. Alternatively, chromaticity Cb and Cr can construct their own CCRM.
[0323] 10) For filter design of CCRM models, sample values and / or gradient and / or location information can be taken into account.
[0324] a. For example, at least one K-tap filter can be used in a CCRM model, which consists of K1 sample terms, K2 gradient terms, K3 localization / position terms, K4 nonlinear terms, K5 bias terms, etc.
[0325] i. For example, K1 = 0 or 1 or 2 or 5 or 6 ii. For example, K2 = 0, 1, 2, or 4 iii. For example, K3 = 0, 1, 2, or 4 iv. For example, K4 = 0, 1, 2, or 4 v. For example, K5 = 0 or 1 vi. For example, K = K1 + K2 + K3 + K4 + K5 vii. For example, the sample item can be calculated based on the luminance sample value.
[0326] viii. For example, the gradient term can be calculated based on more than one sample adjacent to a particular brightness sample.
[0327] ix. For example, the positioning / location item can be calculated based on the horizontal and / or vertical coordinates of a specific brightness sample point, where the coordinates can be relative to the upper left position of a specific reference area.
[0328] x. For example, a nonlinear term could be the square of a specific value (e.g., a bit depth-related intermediate value, such as 512 or 256, or a specific brightness value).
[0329] xi. For example, a nonlinear term can be the square of the gradient value based on a particular gradient term.
[0330] xii. For example, the offset can be subtracted from the terms of the K-tap filter.
[0331] 1. For example, the offset can be derived based on predefined rules such as the value of the training sample in the upper left of the training area, or the average / median value of more than one sample in the training area.
[0332] xiii. For example, the coefficients of a K-tap filter can be solved using a Gaussian elimination solver.
[0333] xiv. For example, the coefficients of a K-tap filter can be solved using the LDL decomposition method.
[0334] i. For example, the coefficients of a K-tap filter can be solved by linear regression.
[0335] ii. For example, the coefficients of a K-tap filter can be solved using linear equations.
[0336] b. For example, more than one filter can be used, and the final prediction can be derived by fusing the filtered outputs of multiple filters together.
[0337] i. For example, the weights that fuse multiple filter values can be solved using a Gaussian elimination solver.
[0338] ii. For example, the weights for fusing multiple filter values can be solved using the LDL decomposition method.
[0339] 11) For example, for a video unit encoded or decoded by CCRM, more than one filter may be allowed, and which filter is ultimately selected may be transmitted through the signal or divided.
[0340] a. For example, syntax elements can be transmitted via signals to indicate which filter (e.g., CCLM or CCCM) is used in CCRM mode.
[0341] b. For example, indicating which filter (e.g., CCLM or CCCM) is used in CCRM mode can be determined based on template costs from both the encoder and decoder.
[0342] c. For example, indicating which filter (e.g., CCLM or CCCM) is used in CCRM mode can be determined based on the cost derived from both the encoder and decoder.
[0343] 12) The filter output can be limited to a certain value.
[0344] a. For example, it can be limited based on the reconstructed values in the training region.
[0345] i. For example, the training region can be derived based on block vectors (or motion vectors).
[0346] ii. For example, the training region can be adjacent to the current block.
[0347] iii. For example, the training region can be the reference region of the current block.
[0348] iv. For example, the filter output can be limited to the minimum and maximum values of the reconstructed (or predicted) luminance sample values in the training region.
[0349] b. For example, it can be limited based on the reconstructed (or predicted) value in the co-located luminance block of the current chroma block.
[0350] i. For example, it can be limited to the minimum and maximum values of the current block brightness reconstruction (or prediction) value.
[0351] c. For example, if the value is outside the valid range, it can be ignored / discarded / not used.
[0352] 13) CCRM parameters can be stored in the cache and used for encoding and decoding of future blocks.
[0353] a. For example, CCRM parameters for video units (e.g., CU, PU, color components, Cb, Cr, etc.) may include model type, model coefficients, whether it is a single model or multiple models, threshold for separating samples into multiple models, etc.
[0354] b. For example, it can be stored in a local cache for encoding and decoding future blocks in the current image.
[0355] c. For example, it can be stored in the temporal domain / picture / frame buffer for encoding and decoding future blocks in a future decoded picture.
[0356] i. For example, the CCRM parameters of the current frame / image can be stored and referenced for the CCP process of future frames / images.
[0357] ii. For example, it can be stored in association with motion and pattern information of the video unit.
[0358] 14) Video blocks can inherit model parameters from previous filter-based codec blocks. In the sub-items below, CCRM can refer to any filter model that includes cross-component models or same-component models.
[0359] a. For example, a cross-component model can refer to a cross-component residual model or cross-component prediction model used within or between frames, where model generation and application are based on the relationship between different color components (such as luminance to chrominance, chrominance Cb to chrominance Cr).
[0360] b. For example, a common component model can refer to an inter-frame / IBC / IBC-LIC / inter-frame-LIC / intraTMP / EIF filter, where model generation and model application are based on the relationship in the common component (such as luminance to luminance S).
[0361] c. For example, video blocks can be encoded and decoded using a type of CCP inheritance pattern.
[0362] d. For example, video blocks can be encoded and decoded using a type of CCP Merge (e.g., CCMerge) mode.
[0363] e. For example, video blocks can be encoded and decoded using a type of filter model inheritance / merge mode. For example, model parameters of previously encoded and decoded blocks using CCRM can be stored in a cache (e.g., local cache, image cache, temporal cache, history-based LUT, etc.).
[0364] f. In one example, parameters could refer to filter information, linear or nonlinear parameters of the model, model index, etc.
[0365] g. For example, at least one syntax element may be signaled at the video unit level (e.g., block level, tu / pu / cu level, etc.) to specify whether and / or how to use CCRM model inheritance mode (e.g., CCRM Merge mode, or CCP Merge mode, or IBC / intraTMP filter Merge mode).
[0366] i. For example, an indicator can be transmitted via signaling at the video unit level to specify whether the current video unit uses the regular CCRM mode or the CCRM model inheritance mode.
[0367] 1. For example, alternatively, based on at least one type of CCRM (e.g., conventional CCRM mode) used for the current video unit, the indicator is conditionally transmitted via signaling.
[0368] 2. For example, a first syntax is transmitted via signal to indicate that the current video unit uses a certain type of CCRM mode, and then a second syntax is further transmitted via signal to indicate which type of CCRM mode is used.
[0369] a. In addition, alternatively, the second syntax may be transmitted via signaling only if at least one available CCRM candidate exists (e.g., at least one valid CCRM Merge candidate exists).
[0370] ii. For example, alternatively, if the CCRM model inheritance pattern is used, another syntax (e.g., index) can be further specified via signaling to indicate which CCRM model candidate is selected to be inherited.
[0371] 1. For example, candidate indexes can be encoded or decoded to indicate CCRM candidates from a candidate list.
[0372] 2. For example, candidate indices can be encoded or decoded using the Rounding Rice (TR) binarization process, or the Rounding Binary (TB) binarization process, or the k-th order Exp-Golomb (EGk) binarization process, or the Fixed Length (FL) binarization process.
[0373] 3. For example, the maximum allowed CCRM candidates (e.g., the maximum length of the candidate list) can be specified in the codec (such as 12 or 8 or 4 or 2, etc.).
[0374] iii. For example, alternatively, an indicator may be transmitted via signaling at the video unit level to specify whether the current video unit uses the CCRM model inheritance mode, and / or which candidate is used for the CCRM model inheritance mode.
[0375] iv. For example, syntax elements can be transmitted via signaling depending on the residual / coefficient of the current luma block.
[0376] 1. For example, whether to transmit a syntax element via signal can be conditional based on the presence of a residual (or non-zero coefficient) in the current luma block.
[0377] 2. For example, whether a signal transmission syntax element is used can be conditional based on the distribution / number / value of the residuals (or non-zero coefficients) in the current luma block.
[0378] v. For example, syntax elements can be conditionally transmitted via signaling based on a prediction method of neighboring blocks.
[0379] 1. For example, it can be based on whether the CCRM / CCRMMerge mode is used based on neighboring blocks (such as left nearest and / or top nearest).
[0380] vi. For example, the context model of a syntax element can depend on the encoding / decoding information of neighboring blocks or the current block.
[0381] 1. For example, the context model can be derived based on whether neighboring blocks (such as left nearest and / or top nearest) use the CCRM / MM-CCRM / MF-CCRM / CCRM Merge pattern.
[0382] 2. For example, the context model can be derived based on whether the block dimension of the current block satisfies specific conditions.
[0383] a. For example, if the current block (e.g., TU / PU / CU) is long or wide (e.g., W>a) H, and / or H>b W, where W and H are the width and height of the current block, and a and b are predefined constants (e.g., a=b=2), then the specified context model can be used.
[0384] h. For example, if the CCRM model inheritance pattern is used, a list of CCRM model candidates can be generated.
[0385] i. For example, the maximum length of the list size can be predefined in the bitstream (e.g., the size is equal to 6, 10, or 12 candidate models).
[0386] 1. In addition, the size of the history table can be predefined (e.g., size equal to 5 or 6).
[0387] ii. For example, CCRM model candidates can be obtained based on previously encoded and decoded CCRM blocks that are spatially adjacent, and / or temporally adjacent, and / or spatially non-adjacent, and / or based on historical CCRM candidates, and / or shifted candidates, and / or default CCRM candidates.
[0388] 1. For example, the candidate insertion order can follow predefined rules, such as spatially adjacent -> temporally adjacent -> spatially non-adjacent -> history -> shift -> default.
[0389] a. Alternatively, the candidate insertion order can follow predefined rules, such as spatially adjacent -> spatially non-adjacent (if applicable) -> history (if applicable) -> shift (if applicable) -> default (if applicable).
[0390] 2. For example, CCRM candidates can be inspected at a sub-block (e.g., 4x4) granularity.
[0391] a. For example, each consecutive sub-block within a predefined area can be inspected; for instance, all 4x4 sub-blocks above and to the left of the current video unit can be inspected.
[0392] b. Alternatively, a predefined scatter check order can be used.
[0393] 3. For example, the position of a non-adjacent neighboring block can be based on the block dimension of the current video unit, such as a certain distance from the current video unit, where the distance is proportional to the width and / or height of the current video unit.
[0394] 4. For example, motion shifts (e.g., zero vectors or non-zero vectors) can be used to locate temporal candidates.
[0395] a. For example, motion shift can be based on the motion vector of neighboring blocks.
[0396] b. For example, temporal candidates can come from co-located images.
[0397] c. Alternative locations: Temporal candidates can be derived from reference images, which do not necessarily have to be in the same position as the image.
[0398] 5. For example, history-based CCRM candidates can come from a first-in-first-out history table.
[0399] a. For example, history tables can be initialized at the slice / ctu row / strip / image level.
[0400] 6. For example, time-domain candidates can be derived based on motion displacement.
[0401] a. For example, the temporal candidate of a block encoded by IBC / IntraTMP / inter-frame coding can be derived based on the MV / BV of neighboring blocks.
[0402] iii. For example, it can be based on the MV / BV of the current block. For example, deduplication / redundancy / similarity checks can be applied to CCRM candidate list construction.
[0403] 1. For example, if the candidate to be inserted is different from the specified candidate already in the list, the candidate to be inserted is inserted into the list.
[0404] a. For example, specifying a candidate could refer to all available CCRM candidates in the list.
[0405] b. Alternatively, a specified candidate may refer to one or more specified CCRM candidates in a list (e.g., the last one, and / or the last one in the list – X, where X is a predefined constant).
[0406] 2. For example, the same deduplication / redundancy / similarity check rules can be applied to all types of CCRM candidates.
[0407] a. Alternatively, different deduplication / redundancy / similarity check rules can be applied to different types of CCRM candidates. iv. For example, CCRM candidate reordering can be applied.
[0408] 1. For example, CCRM candidates in the list can be sorted based on the cost inferred by the decoder (e.g., template cost).
[0409] 2. For example, the cost can be derived based on applying CCRM candidates to the reference region / block of the current video unit.
[0410] a. For example, a reference region / block can be identified by the motion vector of the current block.
[0411] b. For example, for each CCRM candidate, the model is first applied to the reference luminance to obtain the predicted reference chromaticity, and then the cost is calculated as the absolute difference between the true reference chromaticity and the predicted reference chromaticity.
[0412] 3. For example, the cost can be derived based on applying CCRM candidates to the neighboring regions / blocks of the current video unit.
[0413] a. For example, a neighboring region / block can be the current block's upper / left nearest neighbor.
[0414] b. For example, for each CCRM candidate, the model is first applied to neighboring luminance to obtain the predicted neighboring chromaticity, and then the cost is calculated as the absolute difference between the true neighboring chromaticity and the predicted neighboring chromaticity.
[0415] 4. For example, based on the cost derived from the decoder above, CCRM candidates can be sorted from lowest cost to highest cost, and the one with the lowest cost is sorted first in the list.
[0416] i. For example, if the CCRM model inheritance mode is used, the specified CCRM candidate model is directly applied to the current video unit without model estimation.
[0417] i. For example, the first candidate in the CCRM list can always be used in the CCRM model inheritance pattern.
[0418] ii. Alternatively, in the bitstream, it can be signaled which candidate from the CCRM list was used for the block encoded / decoded via the CCRM model inheritance mode.
[0419] iii. For example, for RRIBC blocks encoded and decoded using the CCRM model inheritance pattern, the inherited CCRM model can be applied based on the inherited RRIBC flip type.
[0420] 1. For example, if the inherited RRIBC flip type indicates that the inherited CCRM model comes from an RRIBC codec block (e.g., the inherited RRIBC flip type is non-zero), then the CCRM filter taps for the current block can be flipped according to the inherited RRIBC flip type.
[0421] a. For example, when generating CCRM filter taps from unsampled luminance samples, the filter taps of the unsampled luminance samples can be swapped / flipped.
[0422] 2. For example, suppose an 8-tap CCRM model consists of 6 spatial luminance samples, a nonlinear term, and a bias term. Without downsampling, the 6 luminance samples closest to the chromaticity position C are selected from the luminance grid to obtain the spatial luminance samples (L0, ..., L5). The predicted chromaticity value is obtained as: predChromaVal=c0L0+c1L1+c2L2+c3L3+c4L4+c5L5+c6nonlinear((L0+L3+1)>>1)+c7B, where nonlinear is the nonlinear operator of CCRM, and B is the bias.
[0423] a. For example, if the inherited CCRM model comes from a horizontally flipped RRIBC codec block, then when the inherited CCRM model is applied to the current block (e.g., regardless of whether the current block is RRIBC codec), luma samples L1 and L2 can be swapped. And luma samples L4 and L5 can be swapped.
[0424] b. For example, if the inherited CCRM model comes from a vertically flipped RRIBC codec block, then when the inherited CCRM model is applied to the current block (e.g., regardless of whether the current block is RRIBC codec), luma samples L1 and L4 can be swapped. Furthermore, luma samples L0 and L3 can be swapped. Additionally, luma samples L2 and L5 can be swapped. Figure 27 The luminance samples L0, ..., L5 are shown relative to the chromaticity sample C.
[0425] j. For example, the CCRM model of a block encoded and decoded by CCRM can be stored in a cache.
[0426] i. For example, the stored CCRM model information may include the following information.
[0427] 1. CCRM model coefficients / tap / parameters for Cb and Cr components.
[0428] 2. Intermediate values of blocks encoded / decoded by CCRM 3. Bit depth of the block encoded / decoded by CCRM 4. Offset of the Y component of the block encoded / decoded by CCRM 5. Offset values of the U and / or V components of the block encoded / decoded by CCRM. 6. RRIBC flip type of blocks encoded / decoded by CCRM ii. Alternatively, model offsets may not be stored in the cache.
[0429] 1. For example, the model offset of a neighboring block may not be reused in the current block / inherited for use in the current block.
[0430] 2. For example, the model offset of the current block is recalculated based on specific available samples.
[0431] k. Alternatively, filter model information (e.g., model coefficients / tap / parameters / offsets for the Y component) can be stored in a cache, where the filter coefficients are calculated based on the relationship between samples in the same component (e.g., all training samples are in the luminance component domain).
[0432] i. Alternatively, model offsets may not be stored in the cache.
[0433] 1. For example, the model offset of a neighboring block may not be reused in the current block / inherited for use in the current block.
[0434] 2. For example, the model offset of the current block is recalculated based on specific available samples.
[0435] 15) The final prediction can be generated from a weighted sum of multiple fusion / hybrid hypotheses, where at least one hypothesis is based on predictions of CCPs (e.g., CCRM, CCRM Merge, CCP Merge, CCLM, LM, CCCM, GLM, etc.).
[0436] a. In one example, the final prediction of a block can be generated based on multiple prediction candidates from different CCPs (e.g., CCRM, CCRMMerge, CCP Merge, CCLM, LM, CCCM, GLM, etc.).
[0437] i. For example, more than one CCP prediction can be merged together.
[0438] ii. For example, the weights / coefficients of different fusion terms can be solved based on the Gaussian elimination method.
[0439] iii. For example, the weights / coefficients of different fusion terms can be solved based on the LDL decomposition method.
[0440] iv. For example, bias terms can be involved for fusion.
[0441] v. For example, nonlinear terms can be involved in fusion.
[0442] b. In one example, multiple CCP models can be derived to obtain a fused prediction.
[0443] i. Fusion predictions can refer to predictions generated through weighted summation.
[0444] ii. In one example, P0 in luminance and chrominance can be used to derive CCP model M0, P1 in luminance and chrominance can be used to derive CCP model M1, and the final chrominance prediction can be derived as wc0×Pc0+wc1×Pc1, where Pc0 and Pc1 are chrominance predictions obtained using M0 and M1, and wc0 and wc1 are weighting factors.
[0445] iii. P0 and P1 can be predictions from two different directions in a bidirectional prediction.
[0446] iv. For example, predictions for two hypotheses can be generated based on the top two candidates in the CCP Merge pattern list, and the final prediction can be derived based on a weighted sum of the predictions for the two hypotheses.
[0447] c. In one example, the chromaticity prediction obtained through CCP can be fused with other predictions.
[0448] i. For example, chromaticity predictions obtained through CCP can be fused with intra-frame angular predictions.
[0449] ii. For example, chromaticity predictions obtained through CCP can be fused with CCLM predictions.
[0450] iii. For example, chromaticity predictions obtained through CCP can be fused with CCCM predictions.
[0451] iv. For example, chroma predictions obtained through CCP can be fused with original predictions (e.g., intra-frame, inter-frame, or IBC predictions) without CCP.
[0452] 1. For example, suppose the final chromaticity prediction can be derived as (w0×P0 + w1×P1 + offset) >> shift, where P0 represents the chromaticity prediction obtained through CCRM, and P1 represents the chromaticity prediction without CCP. a. The fusion weights w0 and w1 can be fixed and / or predefined, for example, w0=3 and w1=1, or w0=2 and w1=2.
[0453] b. The shift value can be a constant value that can be derived from w0 and w1, for example, shift = log2(w0 + w1).
[0454] c. The offset can be a constant value that can be derived based on the shift and / or fusion weights, for example, offset = shift >> 1, or offset = log2(w0+w1) >> 1.
[0455] d. Alternatively, the fusion weights w0 and w1 can be adaptively determined based on encoding and decoding information (e.g., block dimension, prediction patterns of nearest neighbors, etc.).
[0456] d. In one example, the weighting factors for different hypotheses in the fusion / mixing process can be derived based on predefined rules. The final prediction P is assumed to be derived as P = w0 × P0 + w1 × P1 + w2 × P2 + ..., where w0, w1, and w2 are weighting factors.
[0457] i. For example, fixed values can be assigned to w0, w1, w2… ii. For example, block-based w0, w1, w2... can be assigned.
[0458] iii. For example, for each prediction, a sample-based weighting factor can be assigned (e.g., for different samples in a hypothetical prediction block, at least two different weights can be applied).
[0459] iv. For example, w0, w1, w2 can be derived on the fly (e.g., based on decoded neighboring samples, nearest neighbor prediction patterns, and / or template costs).
[0460] e. In one example, indicators of the weighting factors for different assumptions in the fusion / mixing process can be transmitted via signals in the bitstream.
[0461] i. For example, a lookup table containing multiple sets of weighting factors can be defined, and the index can be signaled to look up the corresponding weight.
[0462] f. For example, the sample values of the final fused / mixed prediction can be clipped to a predefined range, e.g., it can be required to be no less than T1 and no greater than T2, where T2 can depend on the bit depth. For example, Clip1(x) = Clip3(0, (1< <BitDepth) 1,x).
[0463] g. For example, the proposed method can be applied to fuse more than one hypothesis, where at least one hypothesis is based on predictions from a filtering model.
[0464] i. For example, filter coefficients can be calculated based on the relationship between two sets of samples in the same component domain (e.g., one set consists of luminance samples adjacent to the current block, and the other set consists of luminance samples adjacent to the reference block).
[0465] ii. For example, prediction based on a filter model can refer to predictions generated based on applying filters to MC-compensated video units.
[0466] 1. For example, filters can be based on IBC filters, intra-frame TMP filters, LICs for inter-frame, LICs for IBC, EIFs, etc.
[0467] iii. For example, the prediction based on the filter model can be generated based on filter-based merge / inheritance patterns (e.g., IBC filter-based merge / inheritance pattern, intra-frame TMP filter-based merge / inheritance pattern, inter-frame LIC-based merge / inheritance pattern, IBC LIC-based merge / inheritance pattern, EIF-based merge / inheritance pattern).
[0468] 1. For example, a filter-based Merge pattern can refer to a pattern in which the filter model is inherited / derived from the candidate model (e.g., derived from a list of candidate models).
[0469] iv. For example, a prediction based on a filtering model (e.g., encoded or decoded by prediction type A) can be fused with another prediction (e.g., encoded or decoded by prediction type B).
[0470] 1. For example, prediction type B may not be based on a filtering model.
[0471] 2. For example, prediction type B can be a prediction based on a filter model, while the filter types in A and B are different.
[0472] 3. Alternatively, prediction type B can be a prediction based on a filter model, and the filter types in A and B are the same.
[0473] a. For example, predictions for two hypotheses can be generated based on the top two candidates in a filter-based Merge pattern list, and the final prediction can be derived from a weighted sum of the predictions for the two hypotheses.
[0474] 16) Whether CCRM predictions are fused with another prediction can be transmitted via signaling in the bitstream.
[0475] a. For example, a flag can be transmitted via signaling at the video unit level (e.g., TU / PU / CU / strip header / picture header / SPS / PPS level) to indicate such a CCRM fusion mode.
[0476] b. Alternatively, CCRM predictions can always be merged with another prediction without the need for signal transmission.
[0477] 17) For CCRM, bidirectional forecasts can be managed in a different way than unidirectional forecasts. In the following discussion, assume that the forecasts from the two directions are P0 and P1, and that bidirectional forecasts are expressed as Pb = w0 × P0 + w1 × P1, where w0 and w1 are weighting factors.
[0478] a. In one example, Pb in luminance and chrominance can be used to derive the CCRM model.
[0479] b. In one example, P0 or P1 in luminance and chrominance can be used to derive the CCRM model.
[0480] c. In one example, which prediction was used to deduce that the CCRM model could be transmitted via signaling?
[0481] 18) The permission for the CCRM model may depend on at least one of the following: a. Prediction mode of video unit (e.g., MODE_INTRA, MODE_INTER, MODE_IBC, MODE_PLT, etc.); b. Transformation type of video unit (e.g., ACT, color transformation, transformation skip, etc.). c. SBT (e.g., whether SBT is applied to the current video unit); d. The number of non-zero coefficients in a video unit; e. Luminance coefficients (e.g., luminance coefficient values, sum of absolute values of all luminance coefficients, last scan position of non-zero luminance coefficients, AC value, DC value, etc.) and segmentation tree type (e.g., single tree, dual tree). f. Strip type (e.g., I, B, P stripes); g. Color format (e.g., whether it is 4:0:0); h. Availability of chromaticity components; i. For example, CCRM may not be allowed for ACT and / or 4:0:0 color formats.
[0482] j. For example, CCRM on / off can be determined based on the last scan position of a non-zero luminance coefficient.
[0483] i. For example, if the last scan position is less than a threshold, CCRM can be presumed to be disabled for the current chroma unit, and therefore no syntax element is signaled for CCRM use.
[0484] 1. For example, the threshold can be a fixed constant (such as 1).
[0485] 2. For example, the threshold can be a variable based on encoding / decoding information, such as block dimensions.
[0486] ii. For example, if the brightness is not transformed and the encoding / decoding is skipped, such a condition can be checked.
[0487] iii. For example, conditions such as skipping encoding / decoding regardless of whether brightness is transformed can be checked.
[0488] k. For example, CCRM on / off can be determined based on the absolute value of the non-zero luminance coefficient.
[0489] i. For example, it can be determined based on the absolute values of all luminance coefficients (e.g., both AC and DC).
[0490] ii. For example, it can be determined based on the absolute values of all luminance AC coefficients.
[0491] iii. For example, it can be determined based on the luminance DC coefficient value.
[0492] iv. For example, it can be determined based on at least one luminance coefficient value (e.g., DC and / or AC).
[0493] v. For example, if the absolute values are less than a threshold, CCRM can be presumed to be disabled for the current chroma unit, and therefore no syntax element is signaled for CCRM use.
[0494] 1. For example, the threshold can be a fixed constant value.
[0495] 2. For example, the threshold can be a variable based on encoding / decoding information, such as block dimensions.
[0496] vi. For example, if the brightness is not transformed and encoding / decoding is skipped, such a condition can be checked.
[0497] vii. For example, conditions such as skipping encoding / decoding regardless of whether brightness is transformed can be checked.
[0498] viii. For example, such conditions can be checked together with conditions based on block size (e.g., TU width and / or width).
[0499] For example, if a transform skip is used on the luminance component, CCRM may not be applied to the chrominance component.
[0500] i. For example, if a transform skip is used on the luminance component, CCRM can be presumed to be disabled for the current chromaticity unit (e.g., Cb and / or Cr).
[0501] 1. Furthermore, in this case, no syntax elements are used by CCRM on this video unit via signal transmission.
[0502] ii. Alternatively, if transform skip is used on the luminance component, CCRM can be presumed to be always enabled for the current chromaticity unit (e.g., Cb and / or Cr).
[0503] 1. Furthermore, in this case, no syntax elements are used by CCRM on this video unit via signal transmission.
[0504] 19) The application of CCRM can depend on template information.
[0505] a. For example, whether CCRM is allowed for use in video units may depend on the template cost.
[0506] i. For example, if it is determined by a template cost-based approach that CCRM is disabled for the current video unit (i.e., CCRM on / off is presumed rather than signaled), then no syntax element is signaled for CCRM use on that video unit.
[0507] b. For example, such as Figure 28 As shown, assuming the current block is inter-frame encoded / decoded, two costs (e.g., SAD) can be calculated: the first cost is calculated based on the absolute difference between the current template predicted by the CCRM model and the actual reconstruction of the current template, and the second cost is calculated based on the difference between the reference template and the actual reconstruction of the current template. If the first cost is lower than the second cost, CCRM is presumed to be used for the current chroma unit; otherwise, the current chroma unit is encoded / decoded without CCRM.
[0508] i. For example, the CCRM model can be calculated based on the relationship between a reference luminance block (colored in light gray) and a reference chrominance block (colored in light gray).
[0509] ii. For example, if it is determined that CCRM is to be used, the CCRM model can be applied to the current luminance reconstruction block (i.e., the input of the CCRM model) and generate the current chromaticity prediction predicted by the CCRM model (i.e., the output of the CCRM model).
[0510] iii. For example, in this case (i.e., CCRM on / off is presumed rather than transmitted via signaling), no syntax element is transmitted via signaling for CCRM use on that video unit.
[0511] iv. For example, such a template cost method can be applied to blocks that have undergone inter-frame encoding and decoding.
[0512] c. For example, such as Figure 29 As shown, assuming the current block is encoded / decoded using IBC, two costs (e.g., SAD) can be calculated: the first cost is calculated based on the absolute difference between the current template predicted by the CCRM model and the actual reconstruction of the current template, and the second cost is calculated based on the difference between the reference template and the actual reconstruction of the current template. If the first cost is lower than the second cost, CCRM is presumed to be used for the current chroma unit; otherwise, the current chroma unit is encoded / decoded without CCRM.
[0513] i. For example, the CCRM model can be calculated based on the relationship between a reference luminance block (colored in light gray) and a reference chrominance block (colored in light gray).
[0514] ii. For example, if it is determined that CCRM is to be used, the CCRM model can be applied to the current luminance reconstruction block (i.e., the input of the CCRM model) and generate the current chromaticity prediction predicted by the CCRM model (i.e., the output of the CCRM model).
[0515] iii. For example, in this case (i.e., CCRM on / off is presumed rather than transmitted via signaling), no syntax element is transmitted via signaling for CCRM use on that video unit.
[0516] iv. For example, such a template cost method can be applied to blocks encoded and decoded by IBC.
[0517] d. For example, whether a template cost-based approach is used to determine CCRM on / off may depend on whether the current video unit (e.g., TU) has residual / non-zero coefficients and / or SBT usage.
[0518] i. For example, for the zero-residual portion of the current CU encoded and decoded by SBT, the CCRM decision based on template cost may not be applied.
[0519] ii. For example, for the portion of the current CU with residuals after SBT encoding and decoding, the CCRM decision based on template cost may not be applied.
[0520] iii. For example, if the CBF flag of the current luminance TU is false, then the CCRM decision based on template cost may not be applied.
[0521] e. For example, for a TU generated from a CU encoded and decoded by SBT (e.g., the TU size is smaller than the CU size), the template can be constructed from neighboring samples outside the entire CU.
[0522] f. For example, in sub-block / sub-segmentation-based inter-frame / IBC modes, since each sub-block can have its own motion vector, the motion vectors of predefined sub-blocks can be used to locate the reference template.
[0523] i. For example, for a TU encoded with affine / sbTMVP, the MV of a specific sub-block (e.g., the top left corner or the center) can be used.
[0524] ii. For example, for a TU that has been encoded and decoded by GPM inter-frame-to-inter-frame, a specific segment of the MV (e.g., part 0 or part 1) can be used.
[0525] iii. For example, for a TU that has been encoded and decoded via GPM inter-intra-frame, the MV of the inter-frame portion can be used.
[0526] iv. For example, for a TU encoded and decoded by GPM, the MV after TM / MMVD can be used.
[0527] 1. As an alternative, MV prior to TM / MMVD can be used.
[0528] v. Alternatively, if the current TU is encoded and decoded in a sub-block / sub-segmentation-based inter-frame / IBC mode, then the template-based approach may not be applied to the CCRM on / off decision.
[0529] g. For example, if sub-block-based CCRM is applied, the samples of the current template predicted by the CCRM model can be constructed based on sub-blocks.
[0530] i. For example, the sample points of the current template predicted by the CCRM model can be constructed by applying multiple CCRM models with boundary sub-blocks (e.g., upper and / or left boundary sub-blocks).
[0531] ii. For example, if the boundary sub-block does not have a valid CCRM model, the corresponding template samples may not be calculated for cost calculation.
[0532] 1. For example, alternatively, it can be filled with real reconstructed template points.
[0533] h. For example, the samples of the current template predicted by the CCRM model can be constructed from the same CCRM model.
[0534] 20) The CCRM model can be applied to the luminance residual block and output CCRM estimated chrominance residual block.
[0535] a. For example, the final chromaticity prediction block can be generated by adding the first candidate to the second candidate.
[0536] i. For example, the first candidate can be based on the CCRM to estimate the chromaticity residual block, and the second candidate can be based on the chromaticity prediction block before / before the CCRM.
[0537] ii. Alternatively, the two candidates can be mixed / fused based on a weighted sum method.
[0538] 1. For example, the weights of the two candidates can be fixed and / or based on predefined rules.
[0539] b. For example, a model can be derived from reference reconstructed samples and applied to the current residual samples.
[0540] i. For example, CCRM model coefficients can be solved / derived based on a set of training samples, where the training samples can refer to luminance (unsampled or downsampled) and chrominance samples in a reference block.
[0541] ii. For example, the derived model coefficients can be applied to the luminance residual block (unsampled or downsampled) and output the chrominance residual block estimated by CCRM.
[0542] c. For example, different offset values can be used during the CCRM model coefficient derivation and CCRM model application processes. Assume an 8-tap CCRM model consists of 6 spatial luminance samples, a nonlinear term, and an offset term; for model coefficient derivation, the estimated reference chromaticity value is obtained as estChromaVal. ref =c0(L0 ref –offset ref )+c1(L1 ref –offset ref )+c2(L2 ref –offset ref )+c3(L3 ref –offset ref )+c4(L4 ref –offset ref )+c5(L5 ref –offset ref )+c6nonlinear((L0 ref +L3 ref +1)>>1)+c7B ref , of which (L0 ref ,…,L5 ref ) represents the six brightness reconstruction samples in the reference block, nonlinear is the nonlinear operator of CCCM, and B ref It is a bias, and offset ref These are block-based variables; for model applications, the estimated current chromaticity residual value is obtained as estChromaResiVal. cur =c0(L0 cur –offset cur )+c1(L1 cur –offset cur )+c2(L2 cur –offset cur )+c3(L3 cur –offset cur )+c4(L4 cur –offset cur )+c5(L5 cur –offset cur )+c6nonlinear((L0 cur +L3 cur +1)>>1)+c7Bcur , of which (L0 cur ,…,L5 cur ) represents the six brightness residual samples in the current block, nonlinear is the nonlinear operator of CCCM, and B cur It is a bias, and offset cur These are block-based variables.
[0543] i. For example, the offset value can be derived based on at least one training sample.
[0544] 1. For example, the offset can be derived based on specific brightness training samples (e.g., samples at fixed positions in the brightness reference block, such as the upper left sample or the center sample value).
[0545] 2. For example, the offset can be derived based on the average / median of at least two training samples (e.g., all brightness training samples).
[0546] ii. For example, the offset can be defined as a fixed constant (e.g., 0).
[0547] iii. For example, the first offset (e.g., offset) ref ) can be used in the derivation of model coefficients (e.g., c0…c7).
[0548] 1. For example, a Gaussian elimination solver can be used to minimize the difference between the reference chromaticity block estimated by the CCRM model (e.g., the model input could be a real reference luminance reconstruction block) and the real reference chromaticity reconstruction block.
[0549] 2. For example, the first offset (e.g., offset) ref The average value of all sample points in the reconstructed block can be derived based on the real reference brightness.
[0550] iv. For example, the second offset (e.g., offset) cur () can be used in the application process of the model.
[0551] 1. For example, the CCRM model associated with the derived model coefficients can be applied to the current luminance residual block and output an estimated chrominance residual block for the current block.
[0552] 2. For example, the second offset (e.g., offset) cur ) can be fixed to be equal to 0.
[0553] d. For example, different bias values can be used during the CCRM model coefficient derivation process and the CCRM model application process.
[0554] i. For example, the bias value can be derived based on the bit depth of the luminance and chrominance prediction / reconstruction samples in the bitstream (e.g., it can be equal to 1 << (bit depth - 1)).
[0555] 1. Alternatively, it can be equal to a fixed constant (e.g., 0).
[0556] ii. For example, the first bias (e.g., B) ref () can be used in the derivation of model coefficients.
[0557] 1. For example, B ref It can be equal to 1 << (bit depth - 1).
[0558] iii. For example, the second offset (e.g., B) cur () can be used in the model application process.
[0559] 1. For example, B cur It can be fixed to be equal to 0.
[0560] e. For example, whether the CCRM model is used to predict the current chromaticity prediction or the current chromaticity residual can be transmitted via signaling in the bitstream.
[0561] i. Alternatively, it can be implicitly inferred based on decoder information.
[0562] ii. Alternatively, the CCRM model can always be applied to predict the current chromaticity residual.
[0563] 21) The disclosed CCRM pattern can be based on one of the following filters: a. CCLM and / or its variants; b. MMLM and / or its variants; c. CCCM and / or its variants (e.g., GL-CCCM, non-subsampled CCCM, BVG-CCCM, inter-frame CCCM, intra-frame CCCM, etc.). d. GLM and / or its variants; e. Any cross-component prediction that uses information from one channel / component to predict information from another channel / component; f. Any filter-based prediction, wherein the filter coefficients are solved based on the correlation between the prediction and / or reconstruction information.
[0564] 22) Block restrictions can be applied to limit the application of specific types of CCP patterns.
[0565] a. For example, CCP mode can be allowed only when the block size meets predefined rules.
[0566] b. For example, a syntax element can be signaled only if the CCP mode is applicable.
[0567] c. For example, if the use of the CCP mode is not allowed, the syntax element can be presumed to indicate a specific value that no such CCP mode is used for such a block.
[0568] d. For example, at least one of the following block restrictions can be applied to the CCRM mode (assuming W represents the block width and H represents the block height): i. W < T1, or, W <= T1 ii. H < T2, or, H <= T2 iii. Min(W, H) > T3, or, Min(W, H) >= T3 iv. Max(W, H) < T4, or, Max(W, H) <= T4 v. W < T5 H, or, W <= T5 H vi. W > T6 H, or, W >= T6 H vii. H < T7 W, or, H <= T7 W viii. H > T8 W, or, H >= T8 W ix. W H < T9, or W H <= T9 x. For example, T1, T2,... T9 can be predefined integer constants.
[0569] e. For example, the CCRM mode can be allowed only for small blocks.
[0570] i. For example, it can be allowed for blocks smaller than 4x4, or 8x8, or 16x16, or 32x32.
[0571] ii. For example, it can be allowed for blocks with a sample count less than 32, or 64, or 128.
[0572] iii. For example, it can be allowed for blocks with a sample count less than 32, or 64, or 128.
[0573] iv. For example, it can be not allowed for 2xN blocks, where N can be greater than 4 or 8 or 16.
[0574] v. For example, it can be not allowed for Nx2 blocks, where N can be greater than 4 or 8 or 16.
[0575] 23) The disclosed method can be used in a single tree.
[0576] 24) The disclosed method can be used in two trees.
[0577] 25) The disclosed method can be used in inter-frame (such as B or P) stripes.
[0578] 26) The disclosed method can be used in intra-frame (such as I) stripes.
[0579] 27) The “block vector” in the disclosed method can be a “motion vector”.
[0580] 28) The training / reference samples in the disclosed method may refer to the predicted samples and / or reconstructed samples in the training / reference region.
[0581] 29) Whether and / or how the methods disclosed above can be applied can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0582] 30) Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel.
[0583] 31) Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / double tree segmentation, color components, and stripe / picture type.
[0584] 3 questions 1) The permission of a CCP mode may depend on decoding information, such as temporal layer, luma factor, block dimension, etc.
[0585] 2) Signaling for a CCP mode can depend on decoding information, such as proximity prediction mode, block dimension, etc.
[0586] 3) The determination of a certain CCP mode can be based on the cost assessment derived from the decoder.
[0587] a. In addition, bias factors can be introduced for cost assessment.
[0588] 4) Currently, for CCP model calculation and threshold calculation in multi-model CCP mode, all samples in the corresponding luma block are used. However, for inter-frame CCCM mode, a maximum of 256 luma samples are limited to data collection for single-model inter-frame CCCM model parameter calculation. Whether or how such a limitation is applied can be redesigned.
[0589] 5) When the intra-frame CCP Merge mode is used, a candidate ranking / re-ranking process is applied to rank CCP candidates in ascending order based on template cost. However, diversity criteria can be applied during the ranking process.
[0590] a. If applicable, the diversity criterion can be extended to inter-frame CCP Merge mode.
[0591] b. If applicable, the diversity criterion can be extended to the inter-frame CCCM Merge mode.
[0592] 4 Detailed Solutions The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.
[0593] The term "video unit" or "code-decoder unit" can refer to a picture, strip, slice, code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, TU, PB, TB.
[0594] The term "block" can refer to a codec tree block (CTB), a codec tree unit (CTU), or a codec block (CB).
[0595] The term "CCCM" can refer to intra-frame CCCM mode, inter-frame CCCM mode, IBC CCCM mode, intra-frame TMP CCCM mode, etc. It can be a regular CCCM mode or its variants (e.g., GL-CCCM, CCCM without downsampling, CCRM, etc.).
[0596] The term "CCP" can refer to any cross-component prediction method, such as any kind of LM / CCLM / CCCM / GLM / GL-CCCM. It can be an inter-frame CCP, an intra-frame CCP, or a BV-guided CCP.
[0597] The term "CCP mode" can refer to either CCP mode or CCP Merge mode. Examples of CCP Merge modes include inter-frame CCCM Merge mode, intra-frame CCCM Merge mode, inter-frame CCP Merge mode, intra-frame CCP Merge mode, etc. CCP modes can be based on a single model or multiple models. CCP modes can be based on a single filter or multiple filters.
[0598] The term "CCP model" in CCP mode can be calculated based on decoded information such as neighboring samples or reference samples. It can also be inherited / copied / converted / generated from previously encoded / decoded blocks such as CCP Merge mode, inter-frame CCCM Merge mode, etc. The CCP model can be based on linear, non-linear, or convolutional models.
[0599] In this context, when discussing block dimensions (such as block width and block height), "W" is used to indicate block width and "H" is used to indicate block height, and the block can be TU / PU / TU.
[0600] It should be noted that the terminology mentioned below is not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply.
[0601] 1) Whether a certain CCP mode is allowed can depend on the decoding information.
[0602] a. For example, a CCP mode can be inter-frame CCCM Merge mode, intra-frame CCCM Merge mode, inter-frame CCP Merge mode, intra-frame CCP Merge mode, etc.
[0603] i. For example, the CCP model can be inherited / copied / converted / generated from previously encoded / decoded blocks.
[0604] b. For example, a certain CCP mode can be inter-frame CCCM mode, intra-frame CCCM mode, or LM / CCLM / CCCM / GLM / GL-CCCM, etc.
[0605] i. For example, the CCP model can be calculated based on decoded information such as neighboring samples or reference samples.
[0606] c. For example, a CCP pattern can be based on a single model or multiple models.
[0607] d. For example, a CCP mode can be based on a single filter or multiple filters.
[0608] e. For example, whether a certain CCP mode is allowed can depend on the dimensions / width / height of the block (e.g., TU / PU / CU, etc.).
[0609] i. For example, the CCP mode is only allowed for block sizes that meet predefined rules.
[0610] ii. For example, at least one of the following block restrictions can be applied to the CCP mode: 1. a0 W < b0 H, or, a0 W <= b0 H 2. a1 W > b1 H, or, a1 W >= b1 H 3. a2 H < b2 W, or, a2 H <= b2 W 4. a3 H > b3 W, or, a3 H >= b3 W 5. Min(W, H) > T0, or, Min(W, H) >= T0 6. Max(W, H) < T1, or, Max(W, H) <= T1 7. W H < T2, or W H <= T2 8. For example, a0, a1, a2, a3, b0, b1, b2, b3 can be predefined integers.
[0611] 9. For example, T0, T1, T2 can be predefined values.
[0612] iii. For example, if the current TU / PU / CU meets the condition of W H <= T1 (such as T1 = 8 or 16), the inter-frame CCCM (and / or inter-frame CCCM Merge mode, and / or multi-model inter-frame CCCM mode) may not be allowed to be used.
[0613] iv. For example, if the current TU / PU / CU meets the condition of "W >= T2 H" or "H >= T3 W" (such as T2 = T3 = 16 or 8), the inter-frame CCCM (and / or inter-frame CCCM Merge mode, and / or multi-model inter-frame CCCM mode) may not be allowed to be used.
[0614] f. For example, whether a certain CCP mode is allowed can depend on the time-domain layer.
[0615] i. For example, a CCP mode can be allowed only if the temporal layer ID of the video unit meets a condition (e.g., less than a certain value, or greater than a certain value).
[0616] ii. Alternatively, for a predefined time-domain layer, a certain CCP mode may be allowed.
[0617] g. For example, whether a certain CCP mode is allowed can depend on the luminance coefficient.
[0618] i. For example, whether a certain CCP mode is allowed can depend on the sum of the luminance coefficients.
[0619] ii. For example, whether a certain CCP mode is allowed may depend on the scan position with a non-zero luminance coefficient.
[0620] iii. For example, whether a certain CCP mode is allowed can depend on the number of non-zero luminance coefficients.
[0621] iv. For example, whether a multi-model CCP mode is allowed can depend on the luminance coefficient.
[0622] v. For example, whether a single-model CCP mode is allowed can depend on the luminance coefficient.
[0623] h. For example, syntax elements associated with a CCP mode can be transmitted via signals only if the CCP mode is permitted.
[0624] i. For example, if the CCP pattern is not allowed, the syntax element associated with that CCP pattern can be presumed to indicate that no such CCP pattern is allowed / applicable / used for a certain value of such a block.
[0625] 2) Whether and / or how CCP mode can be transmitted via signal can be based on decoded information.
[0626] a. For example, CCP mode can be based on inter-frame CCCM mode, or intra-frame CCCM mode, or inter-frame CCCM Merge mode, or intra-frame / inter-frame CCP Merge mode, or multi-model inter-frame / intra-frame CCP / CCCM mode.
[0627] b. For example, encoding / decoding information can refer to prediction mode / method, block width / height, luminance sample value, temporal layer of the current block, etc.
[0628] c. For example, whether or how CCP mode is transmitted via signaling may depend on the encoding and decoding information of certain previously encoded and / or current blocks.
[0629] i. For example, some previously encoded / decoded blocks may refer to at least one neighboring block of the current block.
[0630] 1. For example, neighboring blocks can be adjacent to and / or not adjacent to the current block.
[0631] 2. For example, a neighboring block can be the TU / PU / CU in the M rows above and / or the N columns to the left of the current block.
[0632] a. For example, M and N can be based on CTU size and / or maximum TU / PU / CU size.
[0633] b. Alternatively, M and N can be fixed values, such as 32 or 16 or 8 or 4 or 2.
[0634] 3. For example, neighboring blocks can be defined based on a predefined inspection order (such as LUTs or rules).
[0635] 4. For example, a neighboring block may be temporally co-located with the current block and / or in the same picture as the current block.
[0636] 5. For example, neighboring blocks can be defined based on a history-based table, such as a FIFO table that stores predictions of previously encoded or decoded blocks.
[0637] ii. For example, whether a syntax element of mode B is transmitted via signaling can be based on whether certain previously encoded or decoded blocks of mode A are utilized.
[0638] 1. For example, mode A can be inter-frame CCCM, and mode B can be a sub-mode of inter-frame CCCM (such as inter-frame CCCM Merge mode or multi-model inter-frame CCCM, etc.).
[0639] 2. For example, mode A can be intra / inter-frame CCP, and mode B can be a sub-mode of intra / inter-frame CCP (such as intra / inter-frame CCP Merge mode, multi-model intra / inter-frame CCP, etc.).
[0640] 3. For example, if at least K blocks (e.g., adjacent to or preceding the current block) are encoded or decoded using pattern A, then syntax elements of pattern B can be signaled. Otherwise, pattern B can be presumed to be unused.
[0641] iii. For example, the maximum length of the candidate list size for mode B can be based on whether some previously encoded blocks have been utilized by mode A encoding / decoding.
[0642] 1. For example, mode A can be inter-frame CCCM, and mode B can be a sub-mode of inter-frame CCCM (such as inter-frame CCCM Merge mode, etc.).
[0643] 2. For example, mode A can be intra / inter-frame CCP, and mode B can be a sub-mode of intra / inter-frame CCP (such as intra / inter-frame CCP Merge mode, etc.).
[0644] 3. For example, if the number of some previously encoded / decoded blocks meets a condition (such as being less than a certain value), the smaller value can be assigned as the maximum length of the candidate list size.
[0645] a. Alternatively, if the number of some previously encoded / decoded blocks meets a condition (such as being greater than a certain value), the larger value can be assigned as the maximum length of the candidate list size.
[0646] iv. For example, how to transmit a signal through a candidate list of modes B can be based on whether certain previously encoded blocks have been utilized by mode A encoding / decoding.
[0647] 1. For example, mode A can be inter-frame CCCM, and mode B can be a sub-mode of inter-frame CCCM (such as inter-frame CCCM Merge mode, etc.).
[0648] 2. For example, mode A can be intra / inter-frame CCP, and mode B can be a sub-mode of intra / inter-frame CCP (such as intra / inter-frame CCP Merge mode, etc.).
[0649] 3. For example, if the number of some previously encoded / decoded blocks meets a condition (such as being less than a certain value), fewer bits can be allocated to encode / decode candidate indices.
[0650] a. Alternatively, if the number of previously encoded / decoded blocks meets a condition (such as being greater than a certain value), more bits can be allocated to encode / decode candidate indices.
[0651] d. For example, whether or how CCP mode is transmitted via signaling can depend on the encoding / decoding information of the current block.
[0652] i. For example, whether a syntax element of signal transmission mode B can be conditionally determined based on whether the current block uses mode A encoding / decoding.
[0653] 1. For example, pattern B can be a subpattern of pattern A.
[0654] 2. For example, whether the flag for inter-frame CCCM Merge mode is transmitted via signal transmission can be conditionally based on the presence of the inter-frame CCCM mode flag.
[0655] a. For example, the inter-frame CCCM Merge mode can be considered a sub-mode of the inter-frame CCCM mode.
[0656] b. For example, if the inter-frame CCCM mode flag is true, the inter-frame CCCM Merge flag can be transmitted via signaling.
[0657] i. For example, if the inter-frame CCCM Merge mode flag is equal to 0 and the inter-frame CCCM mode flag is true, it indicates that regular inter-frame CCCM is used instead of inter-frame CCCM Merge mode (e.g., the inter-frame CCCM model parameters are calculated from training samples, rather than inherited).
[0658] 3. For example, whether the flag for the inter-frame CCCM mode of multi-mode signal transmission is used can be conditionally based on the presence of the inter-frame CCCM mode flag.
[0659] a. For example, multi-model inter-frame CCCM mode can be regarded as a sub-mode of inter-frame CCCM mode.
[0660] b. For example, if the inter-frame CCCM mode flag is true, the multi-model inter-frame CCCM flag can be transmitted via signaling.
[0661] i. For example, if the multi-model inter-frame CCCM mode flag is equal to 0 and the inter-frame CCCM mode flag is true, it indicates that a single-model inter-frame CCCM is used.
[0662] ii. Alternatively, the flag of mode B may be transmitted via signaling independently of the presence of mode A.
[0663] 1. For example, even if the inter-frame CCCM flag is equal to 0, the inter-frame CCCM Merge mode can still be used / transmitted via signaling.
[0664] e. For example, whether CCP mode can be transmitted via signaling can be determined at the CU / PU / TU level, CTU level, slice level, sub-picture level, strip (header) level, picture (header) level, picture group level, sequence level, etc.
[0665] f. For example, syntax elements of CCP mode can be encoded and decoded in context.
[0666] i. For example, at least one context model can be used.
[0667] ii. For example, more than one context model can be used.
[0668] 1. For example, which context model to use may depend on the width and / or height of the block (e.g., TU / PU / CU).
[0669] a. For example, if the block size satisfies the condition W>a H or H>b W (where a and b are constants, e.g., a=b=2), a certain context model can be used.
[0670] 2. For example, which context model to use can depend on the prediction patterns of neighboring blocks.
[0671] a. For example, the context model of pattern B may depend on whether certain neighboring blocks (to the left and / or above) are encoded or decoded using pattern A.
[0672] i. For example, which context model is used for the inter-frame CCCM Merge mode flag can depend on whether the left and / or upper neighbors are being encoded or decoded using the inter-frame CCCM Merge mode.
[0673] ii. For example, which context model is used for the multi-model inter-frame CCCM mode flag can depend on whether the left neighbor and / or the top neighbor are being encoded or decoded using the multi-model inter-frame CCCM mode.
[0674] b. For example, it may depend on whether the left and / or upper neighbors are encoded or decoded using mode B.
[0675] i. For example, which context model is used for the inter-frame CCCM Merge mode flag can depend on whether the left neighbor and / or the top neighbor are being encoded or decoded using the inter-frame CCCM mode.
[0676] ii. For example, which context model is used for the multi-model inter-frame CCCM mode flag can depend on whether the left neighbor and / or the top neighbor are being encoded or decoded using inter-frame CCCM mode.
[0677] 3) Bias factors can be used to determine the use of a particular CCP mode.
[0678] a. For example, assuming the determination of the use of a certain CCP mode is based on a cost comparison derived from the decoder, a bias factor can be introduced for cost evaluation.
[0679] i. For example, the decoded value of a pattern can be derived based on the difference between the real reconstructed sample values and the pattern-predicted sample values in a predetermined region of the sample points.
[0680] 1. For example, the predefined region of a sample point can be the neighboring sample points surrounding the current block.
[0681] 2. Alternatively, the predefined region of a sample point can be a reference sample point in a reference block in the current image (e.g., a reference block guided by BV).
[0682] 3. Alternatively, the predefined region of the sample point can be a reference sample point in a reference block in a reference image (e.g., a reference block guided by MV).
[0683] ii. For example, whether to use the first mode or the second mode can be determined by the cost evaluation derived by the decoder.
[0684] 1. For example, two decoded derivation costs can be calculated for each mode, such as denoted by cost A and cost B, and whether to choose the first mode can be determined by whether the following condition is true.
[0685] a. ((a Cost A + Offset) >> Displacement) < Cost B, where a / offset / displacement is the bias factor.
[0686] 2. For example, the first mode can be a multi-model CCP mode, while the second mode can be a single-model CCP.
[0687] a. For example, in addition, CCP can be inter-frame CCCM.
[0688] 3. For example, the first mode can be a multi-filter CCP mode, while the second mode can be a single-filter CCP.
[0689] a. For example, in addition, CCP can be inter-frame CCCM.
[0690] b. For example, in addition, CCP can be an intra-frame CCCM.
[0691] iii. For example, the value of the bias factor can depend on the decoded information.
[0692] 1. For example, the value of the bias factor can be based on the time-domain layer.
[0693] 2. Alternatively, the bias factor can be set to a predefined constant.
[0694] 4) The luminance samples used in CCP model calculations can be downsampled based on the block size.
[0695] a. For example, CCP mode can be inter-frame CCCM mode, inter-frame CCP mode, or intra-frame CCP mode.
[0696] b. For example, the downsampling factor can depend on the block size.
[0697] i. For example, the downsampling factor may not depend on the chroma format.
[0698] ii. For example, a larger downsampling factor can be assigned to a larger block.
[0699] iii. For example, the downsampling factor can be predefined based on the block width and / or height (e.g., based on LUT).
[0700] c. For example, downsampled luminance blocks can be used for CCP model calculations.
[0701] i. For example, downsampled blocks can be used for data collection to compute CCP model parameters.
[0702] ii. For example, downsampled blocks can be used to calculate thresholds for multi-model sample point classification / categorization.
[0703] d. For example, the maximum allowed number of samples calculated for a certain CCP model can be equal to X (where X is a predefined constant, such as X = 256 or 1024 or 4096).
[0704] i. For example, regardless of the original block size (e.g., even for 256x256, 128x128, or 64x64 blocks), the number of samples used for a certain CCP model calculation can not exceed the value of X.
[0705] ii. For example, alternatively, different X values can be used for different CCP modes.
[0706] 1. For example, X1 can be used for single-model inter-frame CCCM mode; while X2 can be used for multi-model inter-frame CCCM mode, where X2>= X1.
[0707] 2. For example, X1 can be used for single-model inter-frame CCP mode; while X2 can be used for multi-model inter-frame CCP mode, where X2>= X1.
[0708] iii. Alternatively, there may be no limit to the maximum allowed number of samples that can be applied to the calculation of a particular CCP model.
[0709] 1. For example, the complete sample points of a block (without downsampling) can be used for CCP model calculations.
[0710] 2. For example, X1 can be used for single-model inter-frame CCCM mode; while for multi-model inter-frame CCCM mode, there are no restrictions on its use.
[0711] 3. For example, X1 can be used for single-model inter-frame CCP mode; while for multi-model inter-frame CCP mode, there are no restrictions on its use.
[0712] 4. For example, there can be no limit to the maximum allowed number of samples that can be applied to a CCP model calculation, whether it is a multi-model or single-model CCP / CCCM model calculation.
[0713] 5) The diversity criterion can be applied to reorder / rank multiple candidates for CCP patterns.
[0714] a. For example, CCP model candidates in the candidate list can be reordered / ranked according to diversity criteria.
[0715] b. For example, diversity criteria could be based on the cost difference between the current candidate and its earlier candidates in the list.
[0716] i. For example, cost can refer to the cost derived by the decoder based on decoding information (e.g., template cost, reference block prediction cost, etc.).
[0717] ii. For example, diversity candidates can be placed at the beginning of the list.
[0718] 1. In addition, for example, redundant candidates can be placed after non-redundant candidates.
[0719] iii. For example, if the minimum cost difference between the current candidate and its predecessor is less than a threshold, the current candidate can be considered redundant.
[0720] 1. For example, the threshold can depend on the Lagrange parameter.
[0721] 2. For example, thresholds can be predefined according to rules.
[0722] 3. For example, the threshold can be a constant.
[0723] 4. For example, the threshold can depend on the decoding information (such as the temporal layer, block dimension, etc.).
[0724] c. For example, the final CCP model can be selected from a reordered / sorted candidate list of codecs for the current block.
[0725] i. In addition, for example, indicators (e.g., indices) can be transmitted via signals in the bitstream to specify the final selected candidate in the entire candidate list.
[0726] d. For example, CCP Merge modes (e.g., intra-frame CCP Merge mode, inter-frame CCP Merge mode, inter-frame CCCM Merge mode, etc.) can be applied based on a reordered / sorted candidate list.
[0727] 6) Whether to use a model that has been computed on the fly or a model that has been derived can be determined based on the method of derivation by the decoder (rather than the encoder selection).
[0728] a. For example, the CCP model can be calculated on the fly (e.g., the CCP model can be calculated from reference luminance samples and reference chrominance samples), or it can be derived / inherited from the CCP model associated with a previously encoded / decoded CCP block.
[0729] b. For example, whether to use a computed CCP model or a deduced / inherited CCP model (or which CCP mode to use) for the current block may depend on the cost evaluation method deduced by the decoder.
[0730] i. For example, cost evaluation methods derived from decoder derivation can be applied based on decoder-side information such as reference samples.
[0731] ii. For example, given a candidate model, predicted chromaticity samples can be obtained by applying the candidate model to reference luminance samples, and the cost can be calculated by accumulating the differences between the predicted and reconstructed chromaticity samples. In this way, each candidate model can calculate its own cost.
[0732] iii. For example, comparisons can be made based on comparing the cost of the model computed on the fly with the cost of the model derived / inherited, and the one with the lower cost can be determined as the final model to be used for the current block.
[0733] iv. For example, the comparison can be based on comparing the cost of the first derived / inherited model with the cost of the second derived / inherited model, and the one with the lower cost can be determined as the final model used for the current block.
[0734] c. For example, no syntax element may be transmitted via signaling to indicate whether a computed CCP model or a deduced / inherited CCP model (or which CCP mode to use) is used for the current CCP-encoded block.
[0735] i. For example, it can be determined on the decoder side.
[0736] ii. For example, encoder search based on rate-distortion optimization (RDO) may not be necessary.
[0737] d. For example, the CCP model for a block encoded and decoded by inter-frame CCCM can be calculated from a reference sample or derived / inherited from a previously encoded and decoded block by inter-frame CCCM, but it does not need to know where the signal transmission model comes from.
[0738] i. For example, only one flag (e.g., the inter-frame CCCM flag) can be signaled to specify whether the current block is inter-frame CCCM encoded or decoded. However, it is not necessary to signal additional flags to specify whether the model is computed on the fly or inherited / derived.
[0739] 7) The presence (or how it is transmitted via signaling) of syntax elements (e.g., pattern flags, candidate indices, etc.) may depend on whether there are (or how many) nearest neighbors that are encoded and decoded using a particular prediction pattern / method.
[0740] a. For example, whether or not a signal transmission mode flag is used can depend on this.
[0741] i. For example, if there are no nearest neighbors that utilize prediction mode / method A for encoding / decoding, the mode flag for prediction mode / method B for the current block may not be transmitted via signaling (e.g., it is presumed to be equal to 0).
[0742] ii. For example, if the number of the upper (and / or left) nearest neighbors encoded using prediction mode / method A is less than a certain value, the mode flag for prediction mode / method B for the current block may not be transmitted via signaling (e.g., it is presumed to be equal to 0).
[0743] b. For example, whether a prediction pattern / method is allowed for the current block may depend on this.
[0744] i. For example, if there are no nearest neighbors that utilize prediction mode / method A for encoding / decoding, then prediction mode / method B may not be applied to the current block.
[0745] ii. For example, if the number of the top (and / or left) nearest neighbors encoded using prediction mode / method A is less than a certain value, then prediction mode / method B may not be applied to the current block.
[0746] c. For example, how candidate indices are transmitted via signals may depend on this.
[0747] i. For example, the maximum allowed number of Merge candidates can depend on this.
[0748] ii. For example, if the number of upper (and / or left) nearest neighbors encoded using a specific prediction mode / method is equal to K0, then the maximum allowed number of merge candidates for the current block can be equal to L0. Otherwise, if the number of upper (and / or left) nearest neighbors encoded using a specific prediction mode / method is equal to K1, then the maximum allowed number of merge candidates for the current block can be equal to L1, where K0, K1, L0, and L1 are predefined values.
[0749] iii. For example, the binarization process may depend on this.
[0750] d. For example, the context model for syntax signaling can depend on this.
[0751] i. For example, which context model is used can depend on how many nearest neighbors are encoded / decoded using prediction pattern / method A.
[0752] e. For example, the size of the Merge list can depend on this.
[0753] i. For example, if the number of above (and / or left) nearest neighbors encoded using a specific prediction mode / method is equal to K0, then the size of the Merge list for the current block can be equal to L0. Otherwise, if the number of above (and / or left) nearest neighbors encoded using a specific prediction mode / method is equal to K1, then the size of the Merge list for the current block can be equal to L1, where K0, K1, L0, and L1 are predefined values.
[0754] f. For example, the nearest neighbors to be checked can be based on the current TU / PU / CU above M rows and / or to the left N columns.
[0755] g. For example, the nearest neighbors to be examined can be based on spatial proximity and / or non-proximity and / or candidates based on the temporal domain.
[0756] i. For example, nearest neighbors can be checked based on the candidate check order of the Merge list.
[0757] 1. For example, deduplication may not be applied.
[0758] 2. For example, no reconstructed samples may be inspected.
[0759] 3. For example, no candidates may be checked.
[0760] h. For example, the existence of syntax flags for pattern B can depend on the encoding and decoding information of neighboring blocks (such as whether the nearest neighbor is encoded and decoded using pattern A).
[0761] i. For example, pattern B can be a subpattern of pattern A.
[0762] ii. For example, pattern B can be equivalent to pattern A.
[0763] iii. For example, the presence of the CCP Merge / submode flag may depend on whether neighboring blocks are encoded or decoded using CCP mode.
[0764] iv. For example, the presence of the inter-frame CCCM Merge / submode flag can depend on whether neighboring blocks are encoded or decoded using inter-frame CCCM mode.
[0765] v. For example, the presence of the CIIP TM / submode flag can depend on whether neighboring blocks are encoded or decoded using inter-frame / CIIP / TM mode.
[0766] vi. For example, the presence of the DIMD Merge / submode flag can depend on whether neighboring blocks are encoded or decoded using intra-frame / DIMD / TIMD modes.
[0767] vii. For example, the presence of the TIMD Merge / submode flag can depend on whether neighboring blocks are encoded or decoded using intra-frame / DIMD / TIMD modes.
[0768] 8) Multiple assumptions: CCP can be applied to video blocks.
[0769] a. For example, the prediction of chroma blocks can be generated based on the mixing / fusion of the first CCP prediction and the second CCP prediction.
[0770] i. For example, the first CCP prediction and the second CCP prediction can be different.
[0771] b. For example, chroma block predictions can be generated based on a mixture / fusion of regular inter-frame (or intra-frame) predictions and CCP predictions.
[0772] c. For example, CCP prediction can be generated based on inter-frame CCCM mode, inter-frame CCP mode, intra-frame CCP mode, etc.
[0773] d. For example, CCP predictions can be generated based on CCP models of previously encoded and decoded CCP blocks (such as spatially adjacent, non-adjacent, history-based, and time-based CCP candidates).
[0774] e. For example, CCP predictions can be generated based on a CCP model calculated using neighboring sample information (or reference sample information).
[0775] f. For example, a weighted sum of multiple hypotheses can be considered as the final prediction of a video block.
[0776] i. For example, average weighting (e.g., equal weights) can be applied.
[0777] ii. For example, fixed weighting factors (e.g., predefined unequal weights) can be used.
[0778] iii. For example, weights can be determined based on reconstruction / decoding information for neighboring samples / blocks.
[0779] 1. For example, if the energy of the residual signal in the decoded region is low, a higher weight can be assigned to the hypothesis (e.g., the energy of the residual signal in the decoded region can be calculated based on the absolute difference between neighboring sample regions with and without the hypothesis).
[0780] iv. For example, weights can be derived based on syntactic information.
[0781] 9) The disclosed method can be used in a single tree.
[0782] 10) The disclosed method can be used in two trees.
[0783] 11) The disclosed method can be used in inter-frame (such as B or P) stripes.
[0784] 12) The disclosed method can be used in intra-frame (such as I) stripes.
[0785] 13) Whether and / or how the methods disclosed above are applied can be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.
[0786] 14) Whether and / or how the methods disclosed above are applied can be used to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas including more than one sample point or pixel.
[0787] 15) Whether and how to apply the methods disclosed above may depend on the encoded / decoded information, such as block size, color format, single / dual tree segmentation, color components, and stripe / picture type.
[0788] Figure 30 A flowchart of a method 3000 for video processing according to an embodiment of the present disclosure is shown. Method 3000 is implemented during the conversion between video units of a video and a bitstream of a video.
[0789] At box 3010, it is determined whether at least one of the following methods based on decoder derivation is used for the conversion between video units and video bitstreams of a video: a real-time computed model or a derived model.
[0790] At box 3020, the conversion is performed based on at least one of a real-time computed model or a derived model. In some embodiments, the conversion may include encoding video units into a bitstream. Alternatively, the conversion may include decoding video units from a bitstream.
[0791] Method 3000 improves encoding and decoding performance by determining whether to use at least one of a real-time computed model or a derived model.
[0792] In some embodiments, the cross-component prediction (CCP) model can be computed in real time from a CCP model associated with a previously encoded CCP block. For example, the CCP model can be computed from reference luminance samples and reference chrominance samples. In some other embodiments, the CCP model can be derived from a CCP model associated with a previously encoded CCP block. Alternatively, the CCP model can be inherited from a CCP model associated with a previously encoded CCP block.
[0793] In some embodiments, whether at least one of a computed CCP model, a derived model, or an inherited CCP model is used for the current block may depend on the cost evaluation method derived by the decoder. In some other embodiments, which CCP mode is used may depend on the cost evaluation method derived by the decoder. In some embodiments, the cost evaluation method derived by the decoder may be applied based on decoder-side information. For example, decoder-side information may include reference samples. In some other embodiments, predicted chroma samples based on a given candidate model can be obtained by applying the candidate model to reference luma samples. In some examples, the cost value can be calculated by accumulating the difference between the predicted chroma samples and the reconstructed chroma samples. In this way, each candidate model can calculate its own cost value.
[0794] In some embodiments, the comparison may be based on comparing the cost of a model computed in real time with the cost of a derived model or an inherited model. In this case, the model with the lower cost is determined as the final model for the current block. In some other embodiments, the comparison may be based on comparing the cost of a first derived model or a first inherited model with the cost of a second derived model or a second inherited model. In this case, the model with the lower cost is determined as the final model for the current block.
[0795] In some embodiments, no syntax element may be signaled to indicate whether at least one of a computed CCP model, a derived model, or an inherited CCP model is used for the current block. Alternatively, no syntax element may be signaled to indicate which CCP mode is used. In some examples, whether at least one of a computed CCP model, a derived model, or an inherited CCP model is used for the current block, or which CCP mode is used, may be determined at the decoder side. In some embodiments, rate-distortion optimization (RDO) based encoder search may not be required.
[0796] In some embodiments, the CCP model for a block encoded / decoded using the inter-convolutional cross-component model (CCCM) can be computed from reference samples. In some other embodiments, the CCP model for a block encoded / decoded using the inter-convolutional cross-component model (CCCM) can be derived from a previous block encoded / decoded using the inter-convolutional cross-component model (CCCM). Alternatively, the CCP model for a block encoded / decoded using the inter-convolutional cross-component model (CCCM) can be inherited from a previous block encoded / decoded using the inter-convolutional cross-component model (CCCM). However, it is not necessary to signal where the model comes from. In some embodiments, only one flag may be signaled to indicate whether the current block is encoded / decoded using the inter-convolutional cross-component model (CCCM). For example, there may be only one flag, which is the inter-convolutional CCCM flag. However, it is not necessary to signal additional flags to indicate whether the model is computed, inherited, or derived in real time.
[0797] In some embodiments, determining whether to use at least one of a real-time computed model or a derived model may involve using at least one of a single-tree or dual-tree architecture. In some embodiments, determining whether to use at least one of a real-time computed model or a derived model may involve using an inter-frame stripe. For example, an inter-frame stripe may be a B-strip or a P-strip. In some other embodiments, determining whether to use at least one of a real-time computed model or a derived model may involve using an intra-frame stripe. In some embodiments, an intra-frame stripe may be an I-strip.
[0798] In some embodiments, an indication of whether to determine whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method and / or how to determine whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method may be indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, an indication of whether to determine whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method and / or how to determine whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method may be indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0799] In some embodiments, an indication of whether to determine whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method and / or how to determine whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method may be included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[0800] In some embodiments, method 3000 may further include: determining, based on the encoded and decoded information of the target block, whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method and / or how to determine whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method, wherein the encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0801] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining, based on a decoder-derived method, whether to use at least one of the following: a real-time computed model or a derived model; and generating a bitstream based on at least one of the following: a real-time computed model or a derived model.
[0802] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining, based on a decoder-derived method, whether to use at least one of the following: a real-time computed model or a derived model; generating a bitstream based on at least one of the following: a real-time computed model or a derived model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0803] Figure 31 A flowchart of a method 3100 for video processing according to an embodiment of the present disclosure is shown. Method 3100 is implemented during the conversion between video units of a video and a bitstream of a video.
[0804] At box 3110, at least one of the following is determined based on whether there are neighboring blocks encoded or decoded using a predictive mode or method, or the number of neighboring blocks encoded or decoded using a predictive mode or method: the conversion between video units and video bitstreams, the presence of syntax elements, or the manner of transmitting syntax elements via signals.
[0805] At box 3120, the conversion is performed based on at least one of the following: the presence of a syntax element, or the manner in which a syntax element is signaled. In some embodiments, the conversion may include encoding video units into a bitstream. Alternatively, the conversion may include decoding video units from a bitstream.
[0806] Method 3100 improves encoding / decoding performance by determining the presence of a syntax element or by signaling a syntax element.
[0807] In some embodiments, a neighboring block may be a previously encoded block that is adjacent to the current block in the current image. In some embodiments, a neighboring block may be a previously encoded block that is not adjacent to the current block in the current image. Alternatively, a neighboring block may be a previously encoded block that is adjacent to a co-occurring block of the current block in a reference image. In some other embodiments, a neighboring block may be a previously encoded block that is not adjacent to a co-occurring block of the current block in a reference image.
[0808] In some embodiments, whether a mode flag is transmitted via signaling may depend on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using a predictive mode or method. In some examples, if there are no neighbors encoded using a first predictive mode or method, the mode flag for a second predictive mode or method for the current block may not be transmitted via signaling. For example, the mode flag may be presumed to be equal to 0. In some embodiments, if the number of at least one of the above neighbor or left neighbor encoded using the first predictive mode or method is less than a value, the mode flag for a second predictive mode or method for the current block may not be transmitted via signaling. For example, the mode flag may be presumed to be equal to 0.
[0809] In some embodiments, whether a prediction mode or method is allowed for the current block may depend on at least one of the following: whether there are nearest neighbors encoded using the prediction mode or method, or the number of nearest neighbors encoded using the prediction mode or method. In some examples, if there are no nearest neighbors encoded using the first prediction mode or method, the second prediction mode or method may not be applied to the current block. In some other examples, if the number of at least one of the above nearest neighbors or left nearest neighbors encoded using the first prediction mode or method is less than a certain value, the second prediction mode or method may not be applied to the current block.
[0810] In some embodiments, how candidate indices are transmitted via signaling may depend on at least one of the following: whether there are neighboring blocks encoded using a prediction mode or method, or the number of neighboring blocks encoded using a prediction mode or method. In some embodiments, the maximum allowed number of merge candidates may depend on at least one of the following: whether there are neighboring blocks encoded using a prediction mode or method, or the number of neighboring blocks encoded using a prediction mode or method. In some other embodiments, if the number of at least one of the above nearest neighbor or left nearest neighbor encoded using a prediction mode or method is equal to a first number, then the maximum allowed number of merge candidates for the current block may be equal to a second number. Alternatively, if the number of at least one of the above nearest neighbor or left nearest neighbor encoded using a prediction mode or method is equal to a third number, then the maximum allowed number of merge candidates for the current block may be equal to a fourth number. In this case, the first number, the second number, the third number, and the fourth number are predetermined values. In some embodiments, the binarization process may depend on at least one of the following: whether there are neighboring blocks encoded using a prediction mode or method, or the number of neighboring blocks encoded using a prediction mode or method.
[0811] In some embodiments, the context model used for syntax signaling may depend on at least one of the following: whether there are neighboring blocks encoded using a predictive pattern or method, or the number of neighboring blocks encoded using a predictive pattern or method. For example, which context model is used may depend on the number of nearest neighbors encoded using a predictive pattern or method.
[0812] In some embodiments, the Merge list size may depend on at least one of the following: whether there are neighboring blocks encoded using a prediction mode or method, or the number of neighboring blocks encoded using a prediction mode or method. In some examples, if the number of at least one of the above nearest neighbor or left nearest neighbor encoded using a prediction mode or method is equal to a first number, then the Merge list size for the current block may be equal to a second number. In some other examples, if the number of at least one of the above nearest neighbor or left nearest neighbor encoded using a prediction mode or method is equal to a third number, then the Merge list size for the current block may be equal to a fourth number. In this case, the first number, the second number, the third number, and the fourth number are predetermined values.
[0813] In some embodiments, the nearest neighbors to be checked may be based on at least one of the following: a first number of rows above the current transform unit (TU), a first number of rows above the current prediction unit (PU), a first number of rows above the current codec unit (CU), a second number of columns to the left of the current TU, a second number of columns to the left of the current PU, or a second number of columns to the left of the current CU. In some other embodiments, the nearest neighbors to be checked may be based on at least one of the following: spatially adjacent candidates, spatially non-adjacent candidates, or temporally based candidates. For example, nearest neighbors may be checked based on the checking order of candidates in the Merge list. In some embodiments, deduplication may not be applied. In some other embodiments, reconstructed samples may not be checked. Alternatively, candidates may not be checked.
[0814] In some embodiments, the presence of the syntax flag for the second mode may depend on the encoding / decoding information of neighboring blocks. For example, the encoding / decoding information of neighboring blocks may include whether the neighboring blocks are encoded / decoded using the first mode. In some embodiments, the second mode may be a submode of the first mode. In some other embodiments, the second mode may be the same as the first mode. In some embodiments, the presence of the CCP Merge or submode flag may depend on whether the neighboring blocks are encoded / decoded using the CCP mode. Alternatively, the presence of the inter-frame CCCM Merge or submode flag may depend on whether the neighboring blocks are encoded / decoded using the inter-frame CCCM mode. In some embodiments, the presence of the intra-frame inter-frame joint prediction (CIIP) template matching (TM) or submode flag may depend on whether the neighboring blocks are encoded / decoded using at least one of the following: inter-frame mode, CIIP mode, or TM mode. In some other embodiments, the presence of the decoder-side intra-frame mode derivation (DIMD) Merge or submode flag may depend on whether the neighboring blocks are encoded / decoded using at least one of the following: intra-frame mode, DIMD mode, or template-based intra-frame mode derivation (TIMD) mode. Alternatively, the presence of a template-based intra-mode derivation (TIMD) merge or sub-mode flag may depend on whether neighboring blocks are encoded or decoded using at least one of the following: intra-mode, DIMD mode, or template-based intra-mode derivation (TIMD) mode.
[0815] In some embodiments, it is determined that at least one of the following can be used in at least one of single-tree or dual-tree configurations: the presence of a syntax element, or how the syntax element is transmitted via signaling. In some embodiments, it is determined that at least one of the following can be used in inter-frame stripes: the presence of a syntax element, or how the syntax element is transmitted via signaling. For example, an inter-frame stripe can be a B-strip or a P-strip. In some other embodiments, it is determined that at least one of the following can be used in intra-frame stripes: the presence of a syntax element, or how the syntax element is transmitted via signaling. In some embodiments, an intra-frame stripe can be an I-strip.
[0816] In some embodiments, the indication of whether the existence of a syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using a predictive mode or method, or the manner of signaling a syntax element, and / or how the existence of a syntax element is determined based on at least one of the following: sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, the indication of whether the presence of a syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using a predictive mode or method, or the manner of signaling a syntax element, and / or how to determine the presence of a syntax element based on at least one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0817] In some embodiments, an indication of whether the presence of a syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a prediction mode or method, or the number of neighboring blocks encoded using a prediction mode or method, or the manner of signaling a syntax element, and / or how the presence of a syntax element is determined based on at least one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a strip, a slice, a sub-picture, or a region comprising more than one sample point or pixel.
[0818] In some embodiments, method 3100 may further include: determining, based on the encoded information of the target block, whether the existence of a syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using a predictive mode or method, or the manner of signaling a syntax element, and / or how the existence of a syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using a predictive mode or method, or the manner of signaling a syntax element, wherein the encoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0819] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining at least one of the following based on at least one of the presence of a syntax element, or the manner in which a syntax element is signaled, based on whether there is a neighboring block encoded using a predictive mode or method, or the number of neighboring blocks encoded using a predictive mode or method; and generating a bitstream based on at least one of the following: the presence of a syntax element or the manner in which a syntax element is signaled.
[0820] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining at least one of the following based on at least one of the presence of neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using a predictive mode or method: the presence of syntax elements or the manner of signaling syntax elements; generating a bitstream based on at least one of the following: the presence of syntax elements or the manner of signaling syntax elements; and storing the bitstream in a non-transitory computer-readable recording medium.
[0821] Figure 32 A flowchart of a method 3200 for video processing according to an embodiment of the present disclosure is shown. Method 3200 is implemented during the conversion between video units of a video and a bitstream of a video.
[0822] At box 3210, for the conversion between video units and video bitstreams, it is assumed that CCP mode is applied to the video blocks of the video units.
[0823] At box 3220, the conversion is performed based on a multiple assumption CCP mode. In some embodiments, the conversion may include encoding video units into a bitstream. Alternatively, the conversion may include decoding video units from a bitstream.
[0824] Method 3200 enables the application of multi-hypothesis CCP modes to video blocks. Encoding and decoding performance is improved compared to conventional solutions.
[0825] In some embodiments, the prediction of chroma blocks may be generated based on a mixture or fusion of a first CCP prediction and a second CCP prediction. For example, the first CCP prediction and the second CCP prediction may be different. In some other embodiments, the prediction of chroma blocks may be generated based on a mixture or fusion of conventional inter-frame prediction or conventional intra-frame prediction and CCP prediction.
[0826] In some embodiments, CCP prediction may be generated based on at least one of the following: inter-frame CCCM mode, inter-frame CCP mode, or intra-frame CCP mode. Alternatively, CCP prediction may be generated based on a CCP model of a previously encoded / decoded CCP block. For example, a CCP model of a previously encoded / decoded CCP block may include at least one of the following: spatially adjacent CCP candidates, non-adjacent CCP candidates, history-based CCP candidates, or temporal-based CCP candidates. In some other embodiments, CCP prediction may be generated based on calculating a CCP model using at least one of neighboring sample information or reference sample information.
[0827] In some embodiments, the weighted sum of multiple hypothetical CCPs can be considered as the final prediction of a video block. In some embodiments, average weighting can be applied. For example, average weighting can include equal weights. In some other embodiments, fixed weighting factors can be used. For example, fixed weighting factors can include predetermined unequal weights. In some embodiments, weights can be determined based on at least one of: reconstruction information for neighboring samples, reconstruction information for neighboring blocks, decoding information for neighboring samples, or decoding information for neighboring blocks. In some examples, higher weights can be assigned to hypothetical CCPs if the energy of the residual signal in the decoded region is low. For example, the energy of the residual signal in the decoded region can be calculated based on the cumulative absolute difference between neighboring sample regions with a hypothetical prediction method or hypothetical prediction pattern and neighboring sample regions without a hypothetical prediction method or hypothetical prediction pattern. In some embodiments, weights can be derived based on syntactic information.
[0828] In some embodiments, the multiple hypothesis CCP mode can be used in at least one of single-tree or dual-tree configurations. In some embodiments, the multiple hypothesis CCP mode can be used in inter-frame stripes. For example, the inter-frame stripe can be a B-strip or a P-strip. In some other embodiments, the multiple hypothesis CCP mode can be used in intra-frame stripes. In some embodiments, the intra-frame stripe can be an I-strip.
[0829] In some embodiments, an indication of whether and / or how to apply a multi-hypothesis CCP mode may be given at one of the following: sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, an indication of whether and / or how to apply a multi-hypothesis CCP mode may be given at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0830] In some embodiments, an indication of whether and / or how to apply a multi-hypothesis CCP mode may be included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[0831] In some embodiments, method 3200 may further include: determining whether and / or how to apply a multi-hypothesis CCP mode based on encoded and decoded information of the target block, wherein the encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0832] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: applying multiple hypothetical CCP modes to video blocks of video units in the video; and generating a bitstream based on the multiple hypothetical CCP modes.
[0833] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: applying a multiple hypothetical CCP pattern to video blocks of video units of the video; generating a bitstream based on the multiple hypothetical CCP pattern; and storing the bitstream in a non-transitory computer-readable recording medium.
[0834] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.
[0835] Item 1. A method for video processing, comprising: conversion between video units of a video and a bitstream of the video; determining, based on a decoder-derived method, whether to use at least one of the following: a real-time computed model or a derived model; and performing the conversion based on at least one of the following: the real-time computed model or the derived model.
[0836] Item 2. The method according to Item 1, wherein the cross-component prediction (CCP) model is computed in real time from a CCP model associated with a previously encoded CCP block, or wherein the CCP model is derived from the CCP model associated with the previously encoded CCP block, or wherein the CCP model is inherited from the CCP model associated with the previously encoded CCP block.
[0837] Item 3. The method according to Item 2, wherein the CCP model is calculated from reference luminance samples and reference chromaticity samples.
[0838] Item 4. The method according to Item 1, wherein whether at least one of the computed CCP model, the derived model, or the inherited CCP model is used for the current block depends on the cost evaluation method derived by the decoder.
[0839] Item 5. According to the method described in Item 4, which CCP mode is used depends on the cost evaluation method derived by the decoder.
[0840] Item 6. The method according to Item 4, wherein the cost evaluation method derived by the decoder is applied based on decoder-side information.
[0841] Item 7. The method according to Item 6, wherein the decoder-side information includes reference samples.
[0842] Item 8. The method according to Item 4, wherein the predicted chromaticity samples are obtained by applying the candidate model to reference luminance samples based on a given candidate model.
[0843] Item 9. The method according to Item 8, wherein the cost is calculated by summing the differences between the predicted chromaticity samples and the reconstructed chromaticity samples.
[0844] Item 10. The method according to Item 4, wherein the comparison is based on comparing the cost of the real-time computed model with the cost of the derived model or the cost of the inherited model, wherein the one with the lower cost is determined as the final model for the current block.
[0845] Item 11. The method according to Item 4, wherein the comparison is based on comparing the cost of a first derived model or a first inherited model with the cost of a second derived model or a second inherited model, wherein the one with the lower cost is determined as the final model for the current block.
[0846] Item 12. The method according to Item 1, wherein no syntax element is transmitted by signal to indicate whether at least one of a computed CCP model, a derived model, or an inherited CCP model is used for the current block.
[0847] Item 13. The method described in Item 12, wherein no syntax element is transmitted by signal to indicate which CCP mode is used.
[0848] Item 14. The method according to Item 13, wherein whether at least one of the computed CCP model, the derived model, or the inherited CCP model, or which CCP mode is used for the current block, is determined on the decoder side.
[0849] Item 15. The method according to Item 13, wherein rate-distortion optimization (RDO) based encoder search is not required.
[0850] Item 16. The method according to Item 1, wherein the CCP model for a block encoded and decoded by inter-frame convolutional cross-component model (CCCM) is calculated from reference samples, or wherein the CCP model for the block encoded and decoded by inter-frame convolutional cross-component model (CCCM) is derived from a previous block encoded and decoded by inter-frame CCCM, or wherein the CCP model for the block encoded and decoded by inter-frame convolutional cross-component model (CCCM) is inherited from the previous block encoded and decoded by inter-frame CCCM.
[0851] Item 17. The method described in Item 16, wherein only one flag is transmitted via signaling to indicate whether the current block is encoded or decoded via inter-frame CCCM.
[0852] Item 18. The method according to Item 17, wherein only one of the flags is an inter-frame CCCM flag.
[0853] Item 19. The method according to Item 1, wherein determining whether at least one of a real-time computed model or a derived model is used in at least one of a single tree or a dual tree.
[0854] Item 20. The method according to Item 1, wherein it is determined whether at least one of a real-time computed model or a derived model is used in the inter-frame stripe.
[0855] Item 21. The method according to Item 20, wherein the inter-frame stripe is a B stripe or a P stripe.
[0856] Item 22. The method according to Item 1, wherein it is determined whether at least one of a real-time computed model or a derived model is used in the intra-frame strip.
[0857] Item 23. The method according to Item 22, wherein the intra-frame stripe is an I-strip.
[0858] Item 24. The method according to any one of Items 1 to 23, wherein an indication of whether to use the real-time computed model or at least one of the derived models based on the decoder-derived method and / or how to determine whether to use the real-time computed model or at least one of the derived models based on the decoder-derived method is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0859] Item 25. The method according to any one of Items 1 to 23, wherein an indication of whether to use the real-time computed model or at least one of the derived models based on the decoder-derived method and / or how to determine whether to use the real-time computed model or at least one of the derived models based on the decoder-derived method is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0860] Item 26. The method according to any one of Items 1 to 23, wherein an indication of whether to use the real-time computed model or at least one of the derived models based on the decoder-derived method and / or how to determine whether to use the real-time computed model or at least one of the derived models based on the decoder-derived method is included in one of: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a strip, a slice, a sub-picture, or a region comprising more than one sample point or pixel.
[0861] Item 27. The method according to any one of items 1 to 23 further comprises: determining, based on the encoded and decoded information of the target block, whether to determine whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method and / or how to determine whether to use the real-time computed model or at least one of the derived model based on a decoder-derived method, wherein the encoded and decoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0862] Item 28. A method for video processing, comprising: a conversion between video units of a video and a bitstream of the video; determining, based on at least one of the following: the presence of a syntax element, or the manner in which the syntax element is signaled, whether there is a neighboring block encoded using a prediction mode or method, or the number of neighboring blocks encoded using the prediction mode or method; and performing the conversion based on at least one of the following: the presence of the syntax element, or the manner in which the syntax element is signaled.
[0863] Item 29. According to the method of Item 28, the neighboring block is one of the following: a decoded block adjacent to the current block in the current image, a decoded block not adjacent to the current block in the current image, a decoded block adjacent to a co-occurring block of the current block in the reference image, or a decoded block not adjacent to the co-occurring block of the current block in the reference image.
[0864] Item 30. The method according to Item 28, wherein whether or not a signal transmission mode flag is used depends on at least one of the following: whether there are neighboring blocks encoded or decoded using the prediction mode or method, or the number of neighboring blocks encoded or decoded using the prediction mode or method.
[0865] Item 31. The method according to Item 30, wherein if there is no nearest neighbor encoded or decoded using the first prediction mode or method, the mode flag for the second prediction mode or method for the current block is not transmitted via signaling, wherein the mode flag is presumed to be equal to 0.
[0866] Item 32. The method according to Item 30, wherein if the number of at least one of the upper nearest neighbor or left nearest neighbor encoded using the first prediction mode or method is less than a value, the mode flag of the second prediction mode or method for the current block is not transmitted via signaling.
[0867] Item 33. The method according to Item 32, wherein the mode flag is presumed to be equal to 0.
[0868] Item 34. The method according to Item 28, wherein whether a prediction mode or method is allowed for the current block depends on at least one of the following: whether there are nearest neighbors encoded using the prediction mode or method, or the number of nearest neighbors encoded using the prediction mode or method.
[0869] Item 35. The method according to Item 34, wherein if there are no nearest neighbors encoded or decoded using the first prediction mode or method, the second prediction mode or method is not applied to the current block.
[0870] Item 36. The method according to Item 34, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded using the first prediction mode or method is less than a value, then the second prediction mode or method is not applied to the current block.
[0871] Item 37. The method according to Item 28, wherein how candidate indices are transmitted via signaling depends on at least one of the following: whether there are neighboring blocks encoded using the prediction mode or method, or the number of neighboring blocks encoded using the prediction mode or method.
[0872] Item 38. The method according to Item 37, wherein the maximum allowed number of Merge candidates depends on at least one of the following: whether there are neighboring blocks encoded using the prediction mode or method, or the number of neighboring blocks encoded using the prediction mode or method.
[0873] Item 39. The method according to Item 37, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded using a prediction mode or method is equal to a first number, then the maximum allowed number of Merge candidates for the current block is equal to a second number.
[0874] Item 40. The method according to Item 39, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded using the prediction mode or method is equal to the third number, then the maximum allowed number of Merge candidates for the current block is equal to the fourth number.
[0875] Item 41. The method according to Item 40, wherein the first number, the second number, the third number, and the fourth number are predetermined values.
[0876] Item 42. The method according to Item 37, wherein the binarization process depends on at least one of the following: whether there are neighboring blocks encoded or decoded using the prediction mode or method, or the number of neighboring blocks encoded or decoded using the prediction mode or method.
[0877] Item 43. The method according to Item 28, wherein the context model for syntax signaling depends on at least one of the following: whether there are neighboring blocks encoded or decoded using the prediction mode or method, or the number of neighboring blocks encoded or decoded using the prediction mode or method.
[0878] Item 44. The method described in Item 43, wherein which context model is used depends on the number of nearest neighbors encoded or decoded using the prediction pattern or method.
[0879] Item 45. The method according to Item 28, wherein the size of the Merge list depends on at least one of the following: whether there are neighboring blocks encoded using the prediction mode or method, or the number of neighboring blocks encoded using the prediction mode or method.
[0880] Item 46. The method according to Item 45, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded using a prediction mode or method is equal to a first number, then the size of the Merge list for the current block is equal to a second number.
[0881] Item 47. The method according to Item 46, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded using a prediction mode or method is equal to a third number, then the size of the Merge list for the current block is equal to a fourth number.
[0882] Item 48. The method according to Item 47, wherein the first number, the second number, the third number, and the fourth number are predetermined values.
[0883] Item 49. The method according to Item 28, wherein the nearest neighbors to be examined are based on at least one of the following: a first number of rows above the current transform unit (TU), a first number of rows above the current prediction unit (PU), a first number of rows above the current codec unit (CU), a second number of columns to the left of the current TU, a second number of columns to the left of the current PU, or a second number of columns to the left of the current CU.
[0884] Item 50. The method according to Item 28, wherein the nearest neighbor to be examined is based on at least one of the following: spatially adjacent candidates, spatially non-adjacent candidates, or temporally based candidates.
[0885] Item 51. The method according to Item 50, wherein the neighbor-based Merge list candidate inspection order is checked.
[0886] Item 52. The method described in Item 51, wherein no deduplication is applied.
[0887] Item 53. The method described in Item 51, wherein no reconstructed samples are checked.
[0888] Item 54. The method described in Item 51, wherein no filling candidate is checked.
[0889] Item 55. The method according to Item 28, wherein the presence of the syntax flags of the second mode depends on the encoding / decoding information of neighboring blocks.
[0890] Item 56. The method according to Item 55, wherein the encoding / decoding information of neighboring blocks includes whether the neighboring blocks are encoded / decoded using a first mode.
[0891] Item 57. The method according to Item 56, wherein the second mode is a sub-mode of the first mode.
[0892] Item 58. The method according to Item 56, wherein the second mode is the same as the first mode.
[0893] Item 59. The method according to Item 56, wherein the presence of the CCP Merge or sub-mode flag depends on whether neighboring blocks are encoded or decoded using CCP mode.
[0894] Item 60. The method according to Item 56, wherein the presence of the inter-frame CCCM Merge or sub-mode flag depends on whether neighboring blocks are encoded or decoded using inter-frame CCCM mode.
[0895] Item 61. The method according to Item 56, wherein the presence of an intra-inter-frame joint prediction (CIIP) template matching (TM) or submode flag depends on whether a neighboring block is encoded or decoded using at least one of the following: inter-frame mode, CIIP mode, or TM mode.
[0896] Item 62. The method according to Item 56, wherein the presence of the decoder-side intra-mode derivation (DIMD) merge or sub-mode flag depends on whether the neighboring block is encoded or decoded using at least one of the following: intra-mode, DIMD mode, or template-based intra-mode derivation (TIMD) mode.
[0897] Item 63. The method according to Item 56, wherein the presence of a template-based intra-mode derivation (TIMD) Merge or sub-mode flag depends on whether a neighboring block is encoded or decoded using at least one of the following: intra-mode, DIMD mode, or template-based intra-mode derivation (TIMD) mode.
[0898] Item 64. The method according to Item 28, wherein at least one of the following is determined to be used in at least one of a single tree or a dual tree: the presence of a syntax element, or how the syntax element is transmitted via a signal.
[0899] Item 65. The method according to Item 28, wherein at least one of the following is determined to be used in the inter-frame stripe: the presence of a syntax element, or how the syntax element is transmitted via signaling.
[0900] Item 66. The method according to Item 65, wherein the inter-frame stripe is a B stripe or a P stripe.
[0901] Item 67. The method according to Item 28, wherein at least one of the following is determined to be used in the intra-frame stripe: the presence of a syntax element, or how the syntax element is transmitted via signaling.
[0902] Item 68. The method according to Item 67, wherein the intra-frame stripe is an I-strip.
[0903] Item 69. The method according to any one of items 28 to 68, wherein the indication of whether the existence of the syntax element is determined based on at least one of the presence or number of neighboring blocks encoded using a predictive mode or method, or the manner of signaling the syntax element, and / or how the presence or manner of signaling the syntax element is determined based on at least one of the presence or number of neighboring blocks encoded using a predictive mode or method, is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0904] Item 70. The method according to any one of Items 28 to 68, wherein the indication of whether the existence of the syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using the predictive mode or method, or the manner in which the syntax element is signaled, and / or how the presence of the syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using the predictive mode or method, or the manner in which the syntax element is signaled, is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0905] Item 71. The method according to any one of Items 28 to 68, wherein the indication of whether the existence of the syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a prediction mode or method, or the number of neighboring blocks encoded using the prediction mode or method, or the manner in which the syntax element is transmitted by signaling, and / or how the existence of the syntax element is determined based on at least one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[0906] Item 72. The method according to any one of items 28 to 68, further comprising: determining, based on the encoded information of the target block, whether the existence of the syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using the predictive mode or method; and / or how the existence of the syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using the predictive mode or method; and / or how the existence of the syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded using a predictive mode or method, or the number of neighboring blocks encoded using the predictive mode or method; and the manner in which the syntax element is transmitted via signaling, wherein the encoded information includes at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0907] Item 73. A method for video processing, comprising: a conversion between a video unit of a video and a bitstream of the video; applying a multiple hypothesis CCP mode to a video block of the video unit; and performing the conversion based on the multiple hypothesis CCP mode.
[0908] Item 74. The method according to Item 73, wherein the prediction of the chroma block is generated based on a mixture or fusion of the first CCP prediction and the second CCP prediction.
[0909] Item 75. The method according to Item 74, wherein the first CCP prediction is different from the second CCP prediction.
[0910] Item 76. The method according to Item 73, wherein the prediction of chroma blocks is generated based on a mixture or fusion of conventional inter-frame prediction or conventional intra-frame prediction and CCP prediction.
[0911] Item 77. The method according to Item 73, wherein CCP prediction is generated based on at least one of: inter-frame CCCM mode, inter-frame CCP mode, or intra-frame CCP mode.
[0912] Item 78. The method according to Item 73, wherein CCP prediction is generated based on a CCP model of previously encoded and decoded CCP blocks.
[0913] Item 79. The method according to Item 78, wherein the CCP model of the previously encoded / decoded CCP block includes at least one of the following: spatially adjacent CCP candidates, non-adjacent CCP candidates, history-based CCP candidates, or time-based CCP candidates.
[0914] Item 80. The method according to Item 73, wherein the CCP prediction is generated based on calculating a CCP model using at least one of neighboring sample information or reference sample information.
[0915] Item 81. The method according to Item 73, wherein the weighted sum of the multiple hypothetical CCPs is regarded as the final prediction of the video block.
[0916] Item 82. The method according to Item 81, wherein an average weighting is applied.
[0917] Item 83. The method according to Item 82, wherein the average weighting comprises equal weights.
[0918] Item 84. The method described in Item 81, wherein a fixed weighting factor is used.
[0919] Item 85. The method according to Item 84, wherein the fixed weighting factor comprises predetermined unequal weights.
[0920] Item 86. The method according to Item 81, wherein the weights are determined based on at least one of: reconstruction information for neighboring samples, reconstruction information for neighboring blocks, decoding information for neighboring samples, or decoding information for neighboring blocks.
[0921] Item 87. The method according to Item 86, wherein if the energy of the residual signal in the decoded region is low, a higher weight is assigned to the assumed CCP.
[0922] Item 88. The method according to Item 87, wherein the energy of the residual signal of the decoded region is calculated based on the cumulative absolute difference between neighboring sample regions having a hypothetical prediction method or a hypothetical prediction pattern and neighboring sample regions not having the hypothetical prediction method or the hypothetical prediction pattern.
[0923] Item 89. The method described in Item 81, wherein the weights are derived based on syntactic information.
[0924] Item 90. The method according to Item 73, wherein the multiple assumption CCP pattern is used in at least one of a single tree or a dual tree.
[0925] Item 91. The method according to Item 73, wherein multiple assumptions CCP mode is used in inter-frame stripes.
[0926] Item 92. The method according to Item 91, wherein the inter-frame stripe is a B stripe or a P stripe.
[0927] Item 93. The method according to Item 73, wherein multiple assumptions CCP mode is used in intra-frame stripes.
[0928] Item 94. The method according to Item 93, wherein the intra-frame stripe is an I-strip.
[0929] Item 95. The method according to any one of Items 73 to 94, wherein an indication of whether and / or how a multi-hypothesis CCP mode is applied is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0930] Item 96. The method according to any one of Items 73 to 94, wherein the indication of whether and / or how to apply the multi-hypothesis CCP mode is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.
[0931] Item 97. The method according to any one of Items 73 to 94, wherein an indication of whether and / or how to apply a multi-hypothesis CCP mode is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region comprising more than one sample point or pixel.
[0932] Item 98. The method according to any one of items 73 to 94 further comprises: determining, based on the encoded and decoded information of the target block, whether and / or how to apply the multi-hypothesis CCP mode, said encoded and decoded information including at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color components, stripe type, or picture type.
[0933] Item 99. The method according to any one of items 1 to 98, wherein the conversion includes encoding the video unit into the bitstream.
[0934] Item 100. The method according to any one of items 1 to 98, wherein the conversion includes decoding the video unit from the bitstream.
[0935] Item 101. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1 to 100.
[0936] Item 102. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1 to 100.
[0937] Item 103. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining, based on a decoder-derived method, whether to use at least one of: a real-time computed model or a derived model; and generating the bitstream based on at least one of: the real-time computed model or the derived model.
[0938] Item 104. A method for storing a bitstream of video, comprising: determining, based on a decoder-derived method, whether to use at least one of the following: a real-time computed model or a derived model; generating the bitstream based on at least one of the following: the real-time computed model or the derived model; and storing the bitstream in a non-transitory computer-readable recording medium.
[0939] Item 105. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining at least one of the following based on at least one of the presence of a syntax element, or the manner in which the syntax element is signaled, based on whether there is a neighboring block encoded using a predictive mode or method, or the number of neighboring blocks encoded using the predictive mode or method; and generating the bitstream based on at least one of the following: the presence of the syntax element or the manner in which the syntax element is signaled.
[0940] Item 106. A method for storing a bitstream of video, comprising: determining at least one of the following based on at least one of the presence of a syntax element or the manner in which the syntax element is signaled, based on whether there is a neighboring block encoded using a predictive mode or method or the number of neighboring blocks encoded using the predictive mode or method; generating the bitstream based on at least one of the following: the presence of the syntax element or the manner in which the syntax element is signaled; and storing the bitstream in a non-transitory computer-readable recording medium.
[0941] Item 107. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: applying a multiple hypothetical CCP pattern to video blocks of video units of the video; and generating the bitstream based on the multiple hypothetical CCP pattern.
[0942] Item 108. A method for storing a bitstream of video, comprising: applying a multiple hypothetical CCP pattern to video blocks of video units of the video; generating the bitstream based on the multiple hypothetical CCP pattern; and storing the bitstream in a non-transitory computer-readable recording medium.
[0943] Example device Figure 33 A block diagram of a computing device 3300 in which various embodiments of the present disclosure may be implemented is shown. The computing device 3300 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0944] It should be understood that, Figure 33 The computing device 3300 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.
[0945] like Figure 33As shown, computing device 3300 includes general-purpose computing device 3300. Computing device 3300 may include at least one or more processors or processing units 3310, memory 3320, storage unit 3330, one or more communication units 3340, one or more input devices 3350, and one or more output devices 3360.
[0946] In some embodiments, the computing device 3300 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 3300 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).
[0947] Processing unit 3310 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 3320. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 3300. Processing unit 3310 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0948] Computing device 3300 typically includes various computer storage media. Such media can be any media accessible by computing device 3300, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 3320 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 3330 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 3300.
[0949] The computing device 3300 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 33 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0950] Communication unit 3340 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 3300 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 3300 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0951] Input device 3350 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 3360 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 3340, computing device 3300 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 3300 can also communicate with one or more devices that enable a user to interact with computing device 3300, or any device that enables computing device 3300 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0952] In some embodiments, some or all components of computing device 3300 may be arranged in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be provided remotely and may work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing is provided via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications accessible over a WAN via a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they may appear as a single access point for users. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided from conventional servers or installed directly or otherwise on client devices.
[0953] In embodiments of this disclosure, computing device 3300 can be used to implement video encoding / decoding. Memory 3320 may include one or more video codec modules 3325 having one or more program instructions. These modules are accessible and executable by processing unit 3310 to perform the functions of the various embodiments described herein.
[0954] In an example embodiment of performing video encoding, input device 3350 may receive video data as input 3370 to be encoded. The video data may be processed, for example, by video codec module 3325 to generate an encoded bitstream. The encoded bitstream may be provided as output 3380 via output device 3360.
[0955] In an example embodiment of performing video decoding, input device 3350 may receive an encoded bitstream as input 3370. The encoded bitstream may be processed, for example, by a video codec module 3325 to generate decoded video data. The decoded video data may be provided as output 3380 via output device 3360.
[0956] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between video units and the video bitstream, a decoder-derived method determines whether to use at least one of the following: a real-time computed model or a derived model; and The transformation is performed based on at least one of the following: the real-time calculated model or the derived model.
2. The method of claim 1, wherein the cross-component prediction (CCP) model is computed in real time from a CCP model associated with a previously encoded / decoded CCP block, or The CCP model mentioned above is derived from the CCP model associated with the previously encoded / decoded CCP block, or The CCP model mentioned therein is inherited from the CCP model associated with the previously encoded / decoded CCP block.
3. The method of claim 2, wherein the CCP model is calculated from reference luminance samples and reference chrominance samples.
4. The method of claim 1, wherein whether at least one of a computed CCP model, a derived model, or an inherited CCP model is used for the current block depends on the cost evaluation method derived by the decoder.
5. The method of claim 4, wherein which CCP mode is used depends on the cost evaluation method derived by the decoder.
6. The method of claim 4, wherein the cost evaluation method derived by the decoder is applied based on decoder-side information.
7. The method of claim 6, wherein the decoder-side information includes reference samples.
8. The method of claim 4, wherein the predicted chromaticity samples are obtained by applying the candidate model to reference luminance samples based on a given candidate model.
9. The method of claim 8, wherein the cost is calculated by summing the differences between the predicted chromaticity samples and the reconstructed chromaticity samples.
10. The method of claim 4, wherein the comparison is based on comparing the cost of the real-time computed model with the cost of the derived model or the cost of the inherited model, wherein the one with the lower cost is determined as the final model for the current block.
11. The method of claim 4, wherein the comparison is based on comparing the cost of a first derived model or a first inherited model with the cost of a second derived model or a second inherited model, wherein the one with the lower cost is determined as the final model for the current block.
12. The method of claim 1, wherein no syntax element is transmitted by signaling to indicate whether at least one of a computed CCP model, a derived model, or an inherited CCP model is used for the current block.
13. The method of claim 12, wherein no syntax element is transmitted via signaling to indicate which CCP mode is used.
14. The method of claim 13, wherein whether at least one of the computed CCP model, the derived model, or the inherited CCP model, or which CCP mode is used for the current block, is determined on the decoder side.
15. The method of claim 13, wherein rate-distortion optimization (RDO) based encoder search is not required.
16. The method of claim 1, wherein the CCP model for a block encoded and decoded using inter-frame convolutional cross-component model (CCCM) is calculated from reference samples, or The CCP model for the block encoded and decoded using the inter-frame convolutional cross-component model (CCCM) is derived from the previous block encoded and decoded using the inter-frame CCCM model. The CCP model for the block encoded and decoded by the inter-frame convolutional cross-component model (CCCM) is inherited from the previous block encoded and decoded by the inter-frame CCCM.
17. The method of claim 16, wherein only one flag is transmitted via signaling to indicate whether the current block is encoded or decoded via inter-frame CCCM.
18. The method of claim 17, wherein only one of the flags is an inter-frame CCCM flag.
19. The method of claim 1, wherein determining whether at least one of a real-time computed model or a derived model is used in at least one of a single tree or a dual tree.
20. The method of claim 1, wherein at least one of a real-time computed model or a derived model is used in the inter-frame stripe.
21. The method of claim 20, wherein the inter-frame stripe is a B-strip or a P-strip.
22. The method of claim 1, wherein determining whether at least one of a real-time computed model or a derived model is used in the intra-frame stripe.
23. The method of claim 22, wherein the intra-frame stripe is an I-strip.
24. The method according to any one of claims 1 to 23, wherein an indication of whether to use at least one of the real-time computed model or the derived model based on the decoder-derived method and / or how to determine whether to use at least one of the real-time computed model or the derived model based on the decoder-derived method is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.
25. The method according to any one of claims 1 to 23, wherein an indication of whether to use at least one of the real-time computed model or the derived model based on the decoder-derived method and / or how to determine whether to use at least one of the real-time computed model or the derived model based on the decoder-derived method is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
26. The method according to any one of claims 1 to 23, wherein an indication of whether to determine whether to use at least one of the real-time computed model or the derived model based on the decoder-derived method and / or how to determine whether to use at least one of the real-time computed model or the derived model based on the decoder-derived method is included in one of the following: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
27. The method according to any one of claims 1 to 23, further comprising: Based on the encoded and decoded information of the target block, it is determined whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method and / or how to determine whether to use at least one of the real-time computed model or the derived model based on a decoder-derived method, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.
28. A method for video processing, comprising: For the conversion between video units and the bitstream of the video, based on at least one of the following: the presence of a syntax element, or the manner in which the syntax element is transmitted via signaling; and The transformation is performed based on at least one of the following: the presence of the syntax element, or the manner in which the syntax element is transmitted via signaling.
29. The method of claim 28, wherein the neighboring block is one of the following: Encoded blocks adjacent to the current block in the current image. Encoded blocks that are not adjacent to the current block in the current image. The encoded / decoded block adjacent to the current block in the reference image, or A decoded block that is not adjacent to the current block in the reference image.
30. The method of claim 28, wherein whether or not a signal transmission mode flag is used depends on at least one of the following: whether there are neighboring blocks encoded or decoded using the prediction mode or method, or the number of neighboring blocks encoded or decoded using the prediction mode or method.
31. The method of claim 30, wherein if there is no nearest neighbor encoded or decoded using the first prediction mode or method, the mode flag for the second prediction mode or method for the current block is not transmitted via signaling, wherein the mode flag is presumed to be equal to 0.
32. The method of claim 30, wherein if the number of at least one of the upper nearest neighbor or left nearest neighbor encoded using the first prediction mode or method is less than a value, the mode flag of the second prediction mode or method for the current block is not transmitted via signaling.
33. The method of claim 32, wherein the mode flag is presumed to be equal to 0.
34. The method of claim 28, wherein whether a prediction mode or method is allowed for the current block depends on at least one of the following: whether there are nearest neighbors encoded using the prediction mode or method, or the number of nearest neighbors encoded using the prediction mode or method.
35. The method of claim 34, wherein if there are no nearest neighbors encoded or decoded using the first prediction mode or method, the second prediction mode or method is not applied to the current block.
36. The method of claim 34, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded and decoded using the first prediction mode or method is less than a value, then the second prediction mode or method is not applied to the current block.
37. The method of claim 28, wherein how candidate indices are transmitted via signaling depends on at least one of the following: whether there are neighboring blocks encoded using the prediction mode or method, or the number of neighboring blocks encoded using the prediction mode or method.
38. The method of claim 37, wherein the maximum allowed number of Merge candidates depends on at least one of the following: whether there are neighboring blocks encoded using the prediction mode or method, or the number of neighboring blocks encoded using the prediction mode or method.
39. The method of claim 37, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded and decoded using the prediction mode or method is equal to a first number, then the maximum allowed number of Merge candidates for the current block is equal to a second number.
40. The method of claim 39, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded using the prediction mode or method is equal to a third number, then the maximum allowed number of Merge candidates for the current block is equal to a fourth number.
41. The method of claim 40, wherein the first number, the second number, the third number, and the fourth number are predetermined values.
42. The method of claim 37, wherein the binarization process depends on at least one of the following: whether there are neighboring blocks encoded or decoded using the prediction mode or method, or the number of neighboring blocks encoded or decoded using the prediction mode or method.
43. The method of claim 28, wherein the context model for syntax signaling depends on at least one of the following: whether there are neighboring blocks encoded or decoded using the prediction mode or method, or the number of neighboring blocks encoded or decoded using the prediction mode or method.
44. The method of claim 43, wherein which context model is used depends on the number of nearest neighbors encoded or decoded using the prediction pattern or method.
45. The method of claim 28, wherein the size of the Merge list depends on at least one of the following: whether there are neighboring blocks encoded using the prediction mode or method, or the number of neighboring blocks encoded using the prediction mode or method.
46. The method of claim 45, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded using a prediction mode or method is equal to a first number, then the size of the Merge list for the current block is equal to a second number.
47. The method of claim 46, wherein if the number of at least one of the upper nearest neighbors or left nearest neighbors encoded using the prediction mode or method is equal to a third number, then the size of the Merge list for the current block is equal to a fourth number.
48. The method of claim 47, wherein the first number, the second number, the third number, and the fourth number are predetermined values.
49. The method of claim 28, wherein the nearest neighbors to be checked are based on at least one of the following: The first row above the current transform unit (TU), The first row above the current prediction unit (PU), The first row above the current codec unit (CU), The second column to the left of the current TU, The second column to the left of the current PU, or The second column to the left of the current CU.
50. The method of claim 28, wherein the nearest neighbor to be examined is based on at least one of the following: spatially adjacent candidates, spatially non-adjacent candidates, or temporally based candidates.
51. The method of claim 50, wherein the neighbor-based Merge list candidate inspection order is checked.
52. The method of claim 51, wherein no deduplication is applied.
53. The method of claim 51, wherein no reconstructed samples are inspected.
54. The method of claim 51, wherein no filling candidate is checked.
55. The method of claim 28, wherein the presence of the syntax flags of the second mode depends on the encoding / decoding information of neighboring blocks.
56. The method of claim 55, wherein the encoding / decoding information of neighboring blocks includes whether the neighboring blocks are encoded / decoded using a first mode.
57. The method of claim 56, wherein the second mode is a sub-mode of the first mode.
58. The method of claim 56, wherein the second mode is the same as the first mode.
59. The method of claim 56, wherein the presence of the CCP Merge or sub-mode flag depends on whether neighboring blocks are encoded or decoded using CCP mode.
60. The method of claim 56, wherein the presence of an inter-frame CCCM Merge or sub-mode flag depends on whether neighboring blocks are encoded or decoded using inter-frame CCCM modes.
61. The method of claim 56, wherein the presence of an intra-frame inter-frame joint prediction (CIIP) template matching (TM) or sub-mode flag depends on whether a neighboring block is encoded or decoded using at least one of the following: inter-frame mode, CIIP mode, or TM mode.
62. The method of claim 56, wherein the presence of a decoder-side intra-mode derivation (DIMD) merge or sub-mode flag depends on whether a neighboring block is encoded or decoded using at least one of the following: intra-mode, DIMD mode, or template-based intra-mode derivation (TIMD) mode.
63. The method of claim 56, wherein the presence of a template-based intra-mode derivation (TIMD) merge or sub-mode flag depends on whether a neighboring block is encoded or decoded using at least one of the following: intra-mode, DIMD mode, or template-based intra-mode derivation (TIMD) mode.
64. The method of claim 28, wherein determining at least one of the following is used in at least one of a single tree or a dual tree: the presence of a syntax element, or how the syntax element is transmitted via a signal.
65. The method of claim 28, wherein at least one of the following is determined to be used in the inter-frame stripe: the presence of a syntax element, or how the syntax element is transmitted via signaling.
66. The method of claim 65, wherein the inter-frame stripe is a B-strip or a P-strip.
67. The method of claim 28, wherein at least one of the following is determined to be used in the intra-frame stripe: the presence of a syntax element, or how the syntax element is transmitted via signaling.
68. The method of claim 67, wherein the intra-frame stripe is an I-strip.
69. The method of any one of claims 28 to 68, wherein the presence of the syntax element or the manner of signaling the syntax element is determined based on at least one of the presence or number of neighboring blocks encoded using a predictive mode or method, and / or the indication of how to determine the presence or manner of signaling the syntax element based on at least one of the presence or number of neighboring blocks encoded using a predictive mode or method is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.
70. The method of any one of claims 28 to 68, wherein the presence of the syntax element or the manner of signaling the syntax element is determined based on at least one of the presence or number of neighboring blocks encoded using a predictive mode or method, and / or the indication of how to determine the presence or manner of signaling the syntax element based on at least one of the presence or number of neighboring blocks encoded using a predictive mode or method is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
71. The method of any one of claims 28 to 68, wherein the indication of whether the existence of the syntax element is determined based on at least one of the presence or number of neighboring blocks encoded using a predictive mode or method, or the manner of signaling the syntax element, and / or how the presence or manner of signaling the syntax element is determined based on at least one of the presence or number of neighboring blocks encoded using a predictive mode or method, is included in one of the following: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
72. The method according to any one of claims 28 to 68, further comprising: Based on the encoded and decoded information of the target block, it is determined whether the existence of the syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded and decoded using a prediction mode or method, or the number of neighboring blocks encoded and decoded using the prediction mode or method; or at least one of the following: the manner of signaling the syntax element; and / or how the existence of the syntax element is determined based on at least one of the following: whether there are neighboring blocks encoded and decoded using a prediction mode or method, or the number of neighboring blocks encoded and decoded using the prediction mode or method; or at least one of the following: the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.
73. A method for video processing, comprising: For the conversion between video units and the bitstream of the video, a multi-hypothesis CCP mode is applied to the video blocks of the video units; as well as The transformation is performed based on the aforementioned multiple hypothesis CCP model.
74. The method of claim 73, wherein the prediction of the chroma block is generated based on a mixture or fusion of the first CCP prediction and the second CCP prediction.
75. The method of claim 74, wherein the first CCP prediction is different from the second CCP prediction.
76. The method of claim 73, wherein the prediction of the chroma block is generated based on a mixture or fusion of conventional inter-frame prediction or conventional intra-frame prediction and CCP prediction.
77. The method of claim 73, wherein CCP prediction is generated based on at least one of: inter-frame CCCM mode, inter-frame CCP mode, or intra-frame CCP mode.
78. The method of claim 73, wherein the CCP prediction is generated based on a CCP model of previously encoded and decoded CCP blocks.
79. The method of claim 78, wherein the CCP model of the previously encoded / decoded CCP block includes at least one of the following: spatially adjacent CCP candidates, non-adjacent CCP candidates, history-based CCP candidates, or time-based CCP candidates.
80. The method of claim 73, wherein the CCP prediction is generated based on calculating a CCP model using at least one of neighboring sample information or reference sample information.
81. The method of claim 73, wherein the weighted sum of the multiple hypothetical CCPs is regarded as the final prediction of the video block.
82. The method of claim 81, wherein an average weighting is applied.
83. The method of claim 82, wherein the average weighting comprises equal weights.
84. The method of claim 81, wherein a fixed weighting factor is used.
85. The method of claim 84, wherein the fixed weighting factor comprises predetermined unequal weights.
86. The method of claim 81, wherein the weights are determined based on at least one of: reconstruction information for neighboring samples, reconstruction information for neighboring blocks, decoding information for neighboring samples, or decoding information for neighboring blocks.
87. The method of claim 86, wherein if the energy of the residual signal in the decoded region is low, a higher weight is assigned to the assumed CCP.
88. The method of claim 87, wherein the energy of the residual signal in the decoded region is calculated based on the cumulative absolute difference between neighboring sample regions having a hypothetical prediction method or a hypothetical prediction pattern and neighboring sample regions not having the hypothetical prediction method or the hypothetical prediction pattern.
89. The method of claim 81, wherein the weights are derived based on syntactic information.
90. The method of claim 73, wherein the multiple hypothesis CCP pattern is used in at least one of a single tree or a dual tree.
91. The method of claim 73, wherein a multiple hypothesis CCP mode is used in inter-frame stripes.
92. The method of claim 91, wherein the inter-frame stripe is a B-strip or a P-strip.
93. The method of claim 73, wherein the multiple hypothesis CCP mode is used in intra-frame stripes.
94. The method of claim 93, wherein the intra-frame stripe is an I-strip.
95. The method according to any one of claims 73 to 94, wherein the indication of whether and / or how to apply the multiple hypothesis CCP model is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.
96. The method according to any one of claims 73 to 94, wherein the indication of whether and / or how to apply the multiple hypothesis CCP model is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
97. The method according to any one of claims 73 to 94, wherein an indication of whether and / or how to apply the multiple hypothesis CCP model is included in one of the following: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU). Code-decode tree unit (CTU) CTU line, strip, piece, Sub-images, or This includes regions containing more than one sample point or pixel.
98. The method according to any one of claims 73 to 94, further comprising: Based on the encoded and decoded information of the target block, determine whether and / or how to apply the multi-hypothesis CCP mode, wherein the encoded and decoded information includes at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color components, Strip type, or Image type.
99. The method according to any one of claims 1 to 98, wherein the conversion comprises encoding the video unit into the bitstream.
100. The method according to any one of claims 1 to 98, wherein the conversion comprises decoding the video unit from the bitstream.
101. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 100.
102. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 100.
103. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Decoder-derived methods determine whether to use at least one of the following: a model computed in real time or a derived model; as well as The bitstream is generated based on at least one of the following: the real-time computed model or the derived model.
104. A method for storing a bitstream of video, comprising: Decoder-derived methods determine whether to use at least one of the following: a model computed in real time or a derived model; The bitstream is generated based on at least one of the following: the real-time computed model or the derived model; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
105. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Based on at least one of the following: the presence of a syntax element, or the manner in which the syntax element is signaled; and the existence of a syntax element, or the manner in which the syntax element is signaled, determined by the presence or absence of a neighboring block encoded or decoded using a predictive pattern or method; and The bitstream is generated based on at least one of the following: the presence of the syntax element or the manner in which the syntax element is transmitted via signaling.
106. A method for storing a bitstream of video, comprising: Based on at least one of the following: the presence of a syntax element or the manner in which the syntax element is signaled; The bitstream is generated based on at least one of the following: the presence of the syntax element or the manner in which the syntax element is transmitted via signaling; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
107. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Apply the multiple hypothesis CCP mode to the video blocks of the video unit of the video; as well as The bitstream is generated based on the aforementioned multiple hypothesis CCP mode.
108. A method for storing a bitstream of video, comprising: Apply the multiple hypothesis CCP mode to the video blocks of the video unit of the video; The bitstream is generated based on the aforementioned multiple hypothesis CCP mode; as well as The bitstream is stored in a non-transitory computer-readable recording medium.