Selection of reference samples in block vector-induced cross-component prediction

The improved method for selecting reference samples based on local properties in block vector-guided cross-component prediction addresses performance degradation by ensuring accurate and efficient reference sample selection, enhancing prediction accuracy in video codecs.

JP2026515713APending Publication Date: 2026-05-19NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NOKIA TECHNOLOGIES OY
Filing Date
2024-03-11
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing block vector-guided cross-component prediction methods in video codecs face performance degradation due to inappropriate selection of reference samples, especially when multiple blocks with associated block vectors are located in the same region, leading to decreased prediction accuracy.

Method used

An improved method for selecting reference samples based on local properties, such as size and shape of prediction units, discarding overlapping areas, using weighted averages of block vectors, and applying template-matching refinement mechanisms to derive robust cross-component prediction models, which involves determining a reference sample area using an identified subset of available block vectors in luma encoding or prediction units.

Benefits of technology

Enhances the accuracy and efficiency of cross-component prediction by ensuring the selection of appropriate reference samples, reducing redundancy and improving prediction performance in video codecs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026515713000001_ABST
    Figure 2026515713000001_ABST
Patent Text Reader

Abstract

According to exemplary embodiments of the present invention, there is at least one method and apparatus for performing the following: determining a reference sample area of ​​a video clip sample for deriving a cross-component model by a video decoder, wherein determining is using an identified subset of available block vectors in one of a colocation lumane coding unit or a colocation lumane prediction unit, the subset being identified based on local properties; obtaining a cross-component prediction model based on the determination; and using the cross-component prediction model to decode the video clip sample, as shown in Figure 17.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The teachings according to exemplary embodiments of the present invention generally relate to improving block vector-guided cross-component prediction, and more specifically to improving block vector-guided cross-component prediction, which involves at least the selection of reference samples, processing of multiple block vectors in the same-location area, processing of overlapping reference samples, and selection of robust reference samples based on co-location luminance sample values. [Background technology]

[0002] This section is intended to provide background or context to the invention described in the claims. The descriptions herein may include concepts that can be pursued, but are not necessarily previously invented or pursued. Therefore, unless otherwise indicated herein, the contents of this section are not prior art to the description and claims of this application, nor will they be deemed prior art by being included in this section.

[0003] Certain abbreviations found in descriptions and / or drawings are defined as follows: AMVR Adaptive Motion Vector Resolution BV Block Vector CC Cross-Component CCCM Cross-Component Linear Model CCLM Intersecting Component Linear Model Intra Prediction CTU Encoding Tree Unit CU Central Unit IBC Intrablock Copy ISP Intra-subpartition LM Linear Model LMS (Least Squares Mean) MRL (Multiple Reference Lines) MMLM Multi-Model LM MVD motion vector difference TM Template Matching VVC Multipurpose Video Codec

[0004] A brief explanation of the development process so far Block vector-guided cross-component prediction in video codecs uses non-local regions to improve the prediction performance of cross-component models when intra-block copying is used. When identical blocks in a reference channel are encoded using an intra-block copy (IBC) scheme, the associated block vector (BV) is used to indicate the reference region for calculating the parameters of the cross-component prediction model.

[0005] When multiple blocks with associated block vectors are located in the same region, the selection of a reference sample can be performed in several ways. The reference sample must also reflect the intensity distribution of samples within the same block. Inappropriately selected reference samples directly lead to decreased performance in cross-component prediction.

[0006] Exemplary embodiments of the present invention propose at least an improved operation for block vector-induced cross-component prediction. [Overview of the project]

[0007] This section includes, but does not limit, examples of possible implementations.

[0008] In an exemplary aspect of the present invention, a device such as a user equipment side device includes at least one processor and, when executed by the at least one processor, causes the device to determine a reference sample area of a video clip sample for deriving a cross-component model by a video decoder, the determining being using an identified subset of available block vectors in one of an in-same-position luma encoding unit or an in-same-position luma prediction unit, the subset being identified based on local properties, and stores instructions for at least performing, based on the determining, obtaining a cross-component prediction model and using the cross-component prediction model to decode a video clip sample, at least one non-transitory memory.

[0009] In yet another exemplary aspect of the present invention, there is a method comprising determining, by a video decoder, a reference sample area of a video clip sample for deriving a cross-component model, the determining being using an identified subset of available block vectors in one of an in-same-position luma encoding unit or an in-same-position luma prediction unit, the subset being identified based on local properties, obtaining a cross-component prediction model based on the determining, and using the cross-component prediction model to decode a video clip sample.

[0010] Further exemplary embodiments include apparatus and methods comprising the apparatus and methods of the preceding paragraph, wherein local properties include at least one of the size or shape of a prediction unit, the size or shape of the prediction unit being predetermined by a video decoder or received by a video encoder, the reference sample area is an intrablock copy reference area, the reference sample area is determined based on a plurality of block areas available in at least one prediction unit of a colocation luma, the determination comprising deriving the reference sample area such that overlapping or redundant areas are discarded, the average of the available block vectors is used as a block vector pointing to the reference sample area, the block vector points to the reference sample area if the size difference of the block vectors is less than or equal to a threshold, the average may be calculated as a weighted average based on the size of one of the colocation luma coding units or colocation luma prediction units, and the block vectors of at least one prediction unit are at least one prediction unit Weights are assigned based on the size of the knit prediction unit, the mean is calculated based on a fixed grid of block vectors, at least one of linear or nonlinear estimation methods is used to derive block vectors that point to the reference sample area, and based on multiple block areas, the maximum and minimum coordinate values ​​of a composite reference area defined by the colocation vector and colocation prediction unit area are used to determine the reference sample area, and based on the fact that the dimensions of the reference sample area differ from the dimensions of at least one prediction unit, the device adjusts the dimensions of the reference sample area to match the dimensions of at least one prediction unit, and may use temporary block vectors to identify temporary block vectors of colocation prediction units that do not have associated block vectors, and to determine at least one of the block vectors that can be directly obtained from the reference sample area or colocation prediction unit, and the spatial position of the block vector in one colocation area of ​​the colocation lumern coding unit or colocation lumern prediction unit isFor the derivation of the cross-component model, a template-matching refinement mechanism is used to interpolate or estimate block vectors pointing to reference samples within a reference sample area, and based on the fact that multiple block vectors are available in the same-location prediction unit area, to find at least one matching reference area block or corresponding block vector from an identified subset of available block vectors, and the template-matching refinement mechanism uses at least one sample of the same-location reference area block or at least one of at least one adjacent reference sample within the current block, and if multiple block vectors are available in the same-location prediction unit area, the chroma prediction unit is divided into multiple cross-component models based on the spatial location of the same-location prediction unit, with block vector 0 used to derive the cross-component model of the upper half of the chroma prediction unit marked with 0, and block vector 1 used to derive the cross-component model of the lower half of the chroma prediction unit marked with 1, and based on the fact that multiple block vectors are available in the same-location prediction unit area, multiple cross-component models are divided using multiple block vectors. Dell makes multiple predictions, and the final prediction is obtained by combining the multiple predictions, and the weights for combining the multiple predictions can be those defined in the codec specification, weight identifier indices signaled to the video decoder, or calculated on the decoder side based on block information with reconstructed sample and block size, and statistics of blocks at the same location in multiple block areas are used to reduce the training samples in the reference area block of the reference sample area pointed to by at least one block vector, the minimum and maximum intensity values ​​of samples in blocks at the same location are determined, and samples within the minimum intensity range and maximum intensity range are parameterized,Alternatively, to be used to reduce the minimum and maximum intensity values ​​to be considered suitable for training at least one of the cross-component prediction models, and to provide a wider range for training, the minimum and maximum intensity values ​​are extended by delta values ​​from the upper and lower limits of the maximum intensity range, the delta values ​​being one of a fixed set or determined by the minimum and maximum intensity values, the type of cross-component model is determined based on the distribution of samples from blocks and / or reference blocks at the same location pointed to by block vectors, and if the sample ratios of the models are distributed based on classification parameters, then multiple model variations of the cross-component prediction model are determined, and the sample ratios are predetermined or among those signaled to the decoder One model is determined by a minimum and maximum value, or by other parameters, where the model type is determined by the mean value of the samples used, the mean value has upper and lower limits based on the intensity distance used in parameter calculations, the intensity distance is predetermined or one of those signaled to the decoder, and based on the availability of multiple block vectors in the same-location prediction unit area, the reference block for model derivation is determined by finding an area that coincides with a block at the same location in the reference channel, and / or one or more correlation indices are used to select a sample from the reference sample area pointed to by the block vector of the same-location prediction unit.

[0011] A non-temporary computer-readable medium storing program code, the program code being executed by at least one processor in order to perform at least the methods described in the above paragraph.

[0012] In yet another exemplary aspect of the present invention, there is an apparatus comprising means for determining, by a video decoder, a reference sample area of a video clip sample for deriving an inter-component model, wherein determining is by using an identified subset of available block vectors in one of an in-same-position luma encoding unit or an in-same-position luma prediction unit, the subset being identified based on local properties, and based on the determination, means for obtaining an inter-component prediction model, and means for using the inter-component prediction model for decoding a video clip sample.

[0013] According to the exemplary embodiments described in the above paragraph, at least, the means for determining and obtaining comprises a network interface and computer program code stored in a computer-readable medium and executed by at least one processor.

[0014] A communication system comprises a network-side device and a user equipment-side device that execute the above-described operations.

[0015] The above and other aspects, features, and advantages of various embodiments of the present disclosure will be more clearly understood by reference to the following detailed description and the accompanying drawings. In the drawings, the same reference numerals are used to indicate the same or equivalent elements. The drawings are shown to facilitate a better understanding of the embodiments of the present disclosure and are not necessarily drawn to actual scale.

Brief Description of the Drawings

[0016] [Figure 1A] A diagram showing the positions of samples used for deriving α and β. [Figure 1B] A diagram showing the derivation of a chroma prediction mode from a luma mode when cclm is enabled. [Figure 1C] A diagram showing a unified binarization table for a chroma prediction mode. [Figure 2]This figure shows the lumens of two classes used to derive two sets of α and β (top), the sample region (top), and the spatial region (bottom). [Figure 3] This figure shows the sample positions used for deriving the CCCM filter when six reference lines are used. [Figure 4] The diagram shows, from left to right, three taps arranged vertically, three taps arranged horizontally, five taps arranged in a cross shape, and 25 taps arranged in a diamond shape. [Figure 5] This figure shows an example of four reference lines adjacent to a prediction block. [Figure 6] This figure shows the matrix-weighted intra-prediction process. [Figure 7] This figure shows the HoG calculation from pixels of a template with a width of 3. [Figure 8] This figure shows the low-frequency non-separable transform (LFNST) process. [Figure 9A] This is a diagram showing a table for selecting the conversion method. [Figure 9B] This diagram shows the intra-template matching search area used. [Figure 10] This figure shows the IBC reference region corresponding to the current CU position. [Figure 11] This diagram shows the IBC reference area when CTU(m,n) is encoded. [Figure 12A] (a) This figure shows an example of BV adjustment in the case of horizontal inversion. [Figure 12B] (b) This figure shows an example of BV adjustment in the case of vertical inversion. [Figure 13] This diagram shows the chroma PU and the luma PU in the same position. [Figure 14A] This figure shows two block vectors from coding units C and TL at the same location, pointing to different reference sample areas. Overlapping areas (shown in shaded areas) are considered only once. [Figure 14B] This figure shows two block vectors from coding units C and TL at the same location, indicating that the maximum and minimum coordinate values ​​of a composite reference area defined by a block vector at the same location and a PU area at the same location are used to determine the reference area. [Figure 15] This figure shows two block vectors from predictive units C and TL at the same location, pointing to different reference sample areas. The chroma PU can be divided into two cross-component models (0 and 1) based on the spatial location of the PU to which the block vectors (bv0 and bv1) belong. [Figure 16] This is a block diagram of one possible, non-limiting, exemplary system in which exemplary embodiments may be implemented. [Figure 17] This figure shows a method according to an exemplary embodiment of the present invention, which can be carried out by an apparatus such as the apparatus shown in Figure 16. [Modes for carrying out the invention]

[0017] In exemplary embodiments of the present invention, at least a method and apparatus are proposed for improving block vector-induced cross-component prediction. The improvements address the selection of reference samples, the processing of multiple block vectors within the same location area, the processing of overlapping reference samples, and robust selection of reference samples based on identical luminance sample values.

[0018] As similarly stated above, hybrid video codecs, such as ITU-T H.263, H.264 / AVC, and HEVC, can encode video information in two stages. First, the pixel values ​​of a certain picture area (or "block") are predicted, for example, by motion compensation means (finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being encoded) or by spatial means (using the pixel values ​​around the block to be encoded in a specified manner). In this first stage, predictive coding can be applied, for example, as so-called sample prediction and / or so-called syntax prediction.

[0019] This sample prediction predicts the pixel values ​​or sample values ​​of a given picture area or "block." These pixel values ​​or sample values ​​can be predicted using, for example, one or more motion compensation mechanisms or intra-prediction mechanisms.

[0020] Motion compensation mechanisms (sometimes called interpretation, temporal prediction, or motion-corrected temporal prediction, or motion-corrected prediction or MCP) involve finding and indicating an area in one of the previously encoded video frames that closely corresponds to the block being encoded. Interpretation can reduce temporal redundancy.

[0021] Intra prediction can predict pixel values ​​or sample values ​​using spatial mechanisms. Intra prediction involves finding and demonstrating spatial domain relationships. It leverages the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transformation domains, i.e., it can predict sample values ​​or transformation coefficients. Intra prediction can be used in intra coding where interpretation is not typically applied.

[0022] Syntax prediction, sometimes called parameter prediction, predicts syntax elements, and / or syntax element values ​​and / or variables derived from syntax elements, based on previously encoded (decoded) syntax elements and / or previously derived variables. A non-restrictive example of syntax prediction will be provided later.

[0023] Motion vector prediction allows for the differential encoding of motion vectors, such as motion vectors for inter-prediction and / or interview prediction, relative to the predicted motion vector of a particular block. In many video codecs, the predicted motion vector is generated in a predetermined manner, for example, by calculating the median of the encoded or decoded motion vectors of adjacent blocks. Another method for creating motion vector predictions, sometimes called Advanced Motion Vector Prediction (AMVP), generates a list of candidate predictions from adjacent blocks and / or blocks at the same location as the temporal reference picture, and signals the selected candidates as motion vector predictors. In addition to predicting motion vector values, it is possible to predict the reference index of a previously encoded / decoded picture. The reference index is usually predicted from adjacent blocks and / or blocks at the same location as the temporal reference picture. Differential encoding of motion vectors is usually prohibited across slice boundaries.

[0024] For example, it is possible to predict the block partitioning from CTU to CU, and then to PU.

[0025] Filter parameter prediction can predict filtering parameters, such as filtering parameters for sample adaptation offsets.

[0026] Prediction methods that use image information from previously encoded images are sometimes called interpretation methods, and these methods are sometimes referred to as temporal prediction and motion compensation.

[0027] Prediction methods that use image information from the same image are sometimes called intraprediction methods.

[0028] In the second stage, the prediction error, i.e., the difference between the predicted pixel block and the original pixel block, is encoded. This can be done by transforming the difference in pixel values ​​using a specified transformation (e.g., the discrete cosine transform (DCT) or a variation thereof), quantizing the coefficients, and entropy encoding the quantized coefficients. By changing the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (picture quality) and the size of the resulting encoded video representation (file size at the transmission bitrate).

[0029] In many video codecs, including H.264 / AVC and HEVC, motion information is indicated by motion vectors associated with each motion-compensated image block. Each of these motion vectors represents the displacement between the image block of the picture being encoded (by the encoder) or decoded (by the decoder) and one of the previously encoded or decoded prediction source blocks of the image (or picture). In H.264 / AVC and HEVC, as with many other video compression standards, a single picture is divided into multiple rectangular meshes, each of which is indicated for interpretation by one similar block from the reference picture. The position of the prediction block is encoded as a motion vector that indicates the position of the prediction block relative to the encoded block.

[0030] The Multipurpose Video Codec (VVC) under development includes the following new encoding tools (detailed explanations will be added to the final patent draft later, as needed): ●Intra prediction: - 67 intra modes with wide-angle mode extension - Block size and mode-dependent 4-tap interpolation filter - Location-dependent intra-predictive combination (PDPC) - Cross-Component Linear Model Intra-Prediction (CCLM) - Multiple reference line intra prediction - Intra subpartition - Weighted intra prediction using matrix multiplication ●Picture-to-picture prediction: - Block motion copying using spatial, temporal, history-based, and pairwise average merge candidates - Affine motion interface prediction - Prediction of temporal motion vectors based on subblocks - Adaptive motion vector resolution - Motion compression based on 8x8 blocks for temporal motion prediction - High-precision (1 / 16 pixel) motion vector storage and motion compensation using an 8-tap interpolation filter for the lumens component and a 4-tap interpolation filter for the chromens component. - Triangular partition - Combined intra-prediction and inter-prediction - Merging using MVD (MMVD) - Symmetric MVD coding - Bidirectional optical flow - Refinement of decoder-side motion vector - Bidirectional prediction using CU-level weights ●Transformation, quantization, and coefficient coding: - Multiple linear transformation selection using DCT2, DST7, and DCT8 - Secondary conversion for the low-frequency zone - Subblock transformation for predicted inter-residuals - Dependent quantization using maximum QP increased from 51 to 63 - Transformation coefficient coding using sign data hiding - Conversion skip residual coding ● Entropic coding: - Arithmetic coding engine using adaptive double-window probability updates ● In-loop filter: - In-loop reshaping - Deblocking filter using a more powerful and longer filter - Sample-adaptive offset - Adaptive loop filter ●Screen content encoding: - Current picture referencing using reference area restrictions ● 360-degree video encoding: - Horizontal wrap-around motion compensation ● High-level syntax and parallel processing: - Reference picture management using direct reference picture list signaling - Tile group containing rectangular tile group

[0031] Partitioning in VVC In VVC, each picture is divided into coding tree units (CTUs) similar to those in HEVC. Pictures can also be divided into slices, tiles, bricks, and subpictures. A quaternary tree structure can be used to further divide CTUs into smaller CUs. Each CU can be further divided using nested multi-type trees, including quadtrees, ternary, and binary partitions.

[0032] There are specific rules for inferring picture boundary partitioning.

[0033] Redundant partition patterns are not permitted in nested multi-type partitioning.

[0034] Cross-Component Linear Model Prediction (CCLM) In VVC, a cross-component linear model (CCLM) prediction mode is used to reduce cross-component redundancy. In this mode, the following linear model is used to predict chroma samples based on reconstructed luma samples of the same CU. Nod c (i,j)=α·rec L ’ (i,j)+β In the above formula, pred c(i,j) represents the predicted chroma sample of CU, and rec L ’ (i,j) represents the downsampled and reconstructed lumens sample of the same CU.

[0035] The CCLM parameters (α and β) are derived using up to four adjacent chromatic samples and their corresponding downsampled chromatic samples. If the dimensions of the current chromatic block are W × H, then W' and H' are set as follows: - When LM mode is applied, W'=W, H'=H - When LM-A mode is applied, W'=W+H - When LM-L mode is applied, H'=H+W

[0036] The upper adjacent positions are represented as S[0,-1]...S[W'-1,-1], and the left adjacent positions are represented as S[-1,0]...S[-1,H'-1]. Then, four samples are selected as follows. - When LM mode is applied and both the upper adjacent sample and the left adjacent sample are available, S[W' / 4,-1], S[3*W' / 4,-1], S[-1,H' / 4], S[-1,3*H' / 4] - When LM-A mode is applied, or when only the upper adjacent sample is available, S[W' / 8,-1], S[3*W' / 8,-1], S[5*W' / 8,-1], S[7*W' / 8,-1] - When LM-L mode is applied, or when only left-neighbor samples are available, use S[-1,H' / 8], S[-1,3*H' / 8], S[-1,5*H' / 8], S[-1,7*H' / 8] - Downsample four adjacent chroma samples at the selected location and compare them four times to find two smaller values ​​x0A and x1A and two larger values ​​x0B and x1B. Their corresponding chroma sample values ​​are represented as y0A, y1A, y0B, and y1B. Then derive xA, xB, yA, and yB as shown below. -X a=(x 0 A +x 1 A +1)>>1;X b =(x 0 B +x 1 B +1)>>1;Y a =(y 0 A +y 1 A +1)>>1;Y b =(y 0 B +y 1 B +1)>>1 - Finally, the linear model parameters α and β are obtained according to the following formula.

[0037]

Equation

[0038] Figure 1 shows an example of the positions of the left and upper samples and the sample of the current block involved in the CCLM mode. - The division operation for calculating the parameter α is performed using a lookup table. To reduce the memory required to store this table, the diff value (the difference between the maximum and minimum values) - and the parameter α are expressed in exponential notation. For example, diff is approximated using 4 bits of significant digits and an exponent part. Therefore, the table of 1 / diff is reduced to 16 elements for 16 values of the mantissa as follows. - DivTable[] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0} - This has the advantages of both reducing the computational complexity and reducing the memory size required to store the necessary tables.

[0039] Alternatively, in addition to being able to calculate linear model coefficients together using the upper and left templates, these templates can also be used in two other LM modes called LM_A mode and LM_L mode.

[0040] In LM_A mode, the linear model coefficients are calculated using only the upper template. To obtain more samples, the upper template is extended to (W+H). In LM_L mode, the linear model coefficients are calculated using only the left template. To obtain more samples, the left template is extended to (H+W).

[0041] For non-square blocks, the top template is extended to W+W, and the left template is extended to H+H.

[0042] To align the chroma sample positions for a 4:2:0 video sequence, two types of downsampling filters are applied to the chroma samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "Type 0" and "Type 2" content, respectively.

[0043]

number

[0044] Note that when the upper reference line is on the CTU boundary, only one lumern line (a common line buffer in intra-prediction) is used to create a downsampled lumern sample.

[0045] This parameter calculation is performed as part of the decoding process, not solely as an encoder search operation. Consequently, no syntax is used to communicate the α and β values ​​to the decoder.

[0046] For chroma intra-mode coding, a total of eight intra-modes are allowed. These modes include five traditional intra-modes and three cross-component linear model modes (CCLM, LM_A, and LM_L). The chroma mode signaling and derivation processes are shown in Table 1 below. Chroma mode coding directly depends on the intra-prediction mode of the corresponding rumor block. Since separate block partitioning structures for rumor and chroma components are made available in I-slices, one chroma block may correspond to multiple rumor blocks. Therefore, for chroma DM modes, the intra-prediction mode of the corresponding rumor block covering the central position of the current chroma block is directly inherited.

[0047] Figure 1C shows a unified binarization table for chroma prediction modes. In Figure 1C, the first binary (bin) indicates whether it is normal mode (0) or LM mode (1). If it is LM mode, the next binary indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next binary indicates whether it is LM_L (0) or LM_A (1). In this case, when sps_cclm_enabled_flag is 0, the first binary in the corresponding intra_chroma_pred_mode binarization table may be discarded before entropy coding. Or, in other words, the first binary is inferred to be 0 and therefore not coded. This single binarization table is used for both cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two binaries in the table are context-coded using their own context models, while the other binaries are bypass-coded.

[0048] Furthermore, to reduce rumor-chroma latency in dual trees, when a 64x64 rumor-encoded tree node is partitioned with Not Split (ISP is not used for 64x64 CUs) or QT, the chroma CU of a 32x32 / 32x16 chromor-encoded tree node is allowed to use CCLM in the following manner. - If a 32x32 chroma node is not partitioned or not partitioned with QT partitioning, all chroma CUs on the 32x32 node can use CCLM. - If a 32x32 chroma node is partitioned with Horizontal BT, and a 32x16 child node is not partitioned but uses Vertical BT partitioning, then all chroma CUs of the 32x16 chroma node can use CCLM.

[0049] For all other lumar and chromar coding tree partitioning conditions, CCLM is not allowed for chromar CU.

[0050] Multi-model LM (MMLM) The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, reconstructed neighbor samples are classified into two classes using a threshold that is the average of the reconstructed neighbor samples of the luma. The linear model for each class is derived using the least squares mean (LMS) method. The LMS method is also used to derive the linear model in the CCLM mode. Figure 2 shows two luma-chroma models obtained when the threshold for the luma (Y) is set to 17. Figure 2 shows the luma samples to two classes used in the derivation of two sets of α and β, the sample region (top), and the spatial region (bottom). Each luma-chroma model has its own linear model parameters α and β. As can be seen from the figure below, each luma-chroma model corresponds to the spatial segmentation of the content (i.e., to different objects or textures in the scene).

[0051] Convolutional Cross-Component Model (CCCM) Figure 3 shows the sample positions used for deriving the CCCM filter when six reference lines are used.

[0052] An improved version of cross-component prediction, known as CCCM, derives a lumer-chroma model using a 2D filter kernel. Filter coefficients are derived on the decoder side using a set of reconstructed input data and chroma samples. In deriving the filter coefficients, a reference sample area (consisting of reconstructed lumer and chroma samples) is defined for both lumer and chroma at the same location, as shown in Figure 3, where the commonly used 4:2:0 chroma downsampling is applied. The reference sample area for a given block can be six rows above and to the left, as shown in Figure 3, but any number of reference lines (achievable on both the encoder and decoder) can be used. Generally, the reference samples can include any chroma and lumer samples reconstructed by both the encoder and decoder. Once the reference samples are determined, the filter coefficients can be derived using various types of linear regression tools, such as standard least squares estimation, orthogonal matching tracking, optimized orthogonal matching tracking, ridge regression, or the least absolute condensation selection operator.

[0053] The dimensions of the filter kernel can be arbitrary, such as 1x3 (one-dimensional vertical), 3x1 (one-dimensional horizontal), 3x3, 7x7, etc., and its shape can also be a cross or a diamond (as shown in Figure 4) or any other shape (by selecting only a subset from all possible kernel positions). Figure 4 shows, from left to right, three taps arranged vertically, three taps arranged horizontally, five taps arranged in a cross shape, and 25 taps arranged in a diamond shape. The following notation is used to refer to samples within the filter kernel, as indicated by the letters N, E, S, W, and C in Figure 4: North (top), East (right), South (bottom), West (left), and Center.

[0054] The overall method for reconstructing chroma samples using a convolution operation between a filter kernel acquired on the decoder side and a set of input data is referred to herein as the Convolutional Cross-Component Model (CCCM). The following steps can be applied to perform CCCM operation. 1. Define a reference area that is in the same position for the lumen component and the chroma component. 2. Downsample lumens to align with the chroma grid (optional). 3. Scan the lumens and chroma samples of the reference area and collect available statistics (such as autocorrelation matrices and cross-correlation vectors) based on the filter shape. 4. Solve the filter coefficients by minimizing the squared error (or any other indicator) based on available parameters (such as the autocorrelation matrix and cross-correlation vector). 5. The predicted chroma block is computed by convolving the downsampled chroma samples with a filter kernel.

[0055] The chroma sample (which may be downsampled) is defined as a 2D array Y(x, y) indexed using the horizontal x-coordinate and vertical y-coordinate. The chroma sample at the same position is defined as a 2D array C(x, y), and the filter kernel (i.e., coefficients) is defined as a 3x3 array F(i, j). At the sample level, the convolution operation between Y and F is defined as follows:

[0056]

number

number

number

number

[0057] Multiple Reference Lines (MRL) Intra-Prediction Multiple Reference Line (MRL) intra-prediction uses more reference lines for intra-prediction. Figure 5 shows an example of four reference lines adjacent to a prediction block. Figure 5 shows an example of four reference lines, indicating that the sample values ​​for segments A and F are not fetched from reconstructed adjacent samples, but rather are interpolated with the closest samples from segments B and E, respectively. HEVC intra-picture prediction uses the closest reference line (i.e., reference line 0). MRL uses two auxiliary lines (reference line 1 and reference line 3).

[0058] The index of the selected reference line (mrl_idx) can be signaled and used to generate an intra predictor. For reference line idx greater than 0, only the additional reference line modes can be included in the MPM list, and only the MPM index can be signaled without including the remaining modes. The reference line index can be signaled before the intra predictor mode, and if a non-zero reference line index is signaled, the Planar mode can be excluded from the intra predictor mode.

[0059] MRL can be disabled for the first line of the inner block of the CTU, preventing the use of extended reference samples outside the current CTU line. Furthermore, PDPC can be disabled when auxiliary lines are used. For MRL mode, the derivation of the DC value for DC intra-prediction mode for non-zero reference line indices is aligned with that for reference line index 0. MRL requires storage of three adjacent lumen reference lines with the CTU to generate predictions. The Cross-Component Linear Model (CCLM) tool also requires three adjacent lumen reference lines for its downsampling filter. To reduce the decoder's storage requirements, the definition of MRL, which uses the same three lines, is aligned with CLM.

[0060] Intra-subpartition (ISP) Intra-subpartitions (ISPs) divide intra-predicted rumor blocks vertically or horizontally into two or four subpartitions depending on the block size. For example, the minimum block size for an ISP is 4x8 (or 8x4). If the block size is larger than 4x8 (or 8x4), the corresponding block is divided into four subpartitions. It has been found that Mx128 (M≦64) and 128xN (N≦64) ISP blocks can generate potential problems, including 64x64 VDPUs. For example, Mx128 CUs in the case of a single tree would result in Mx128 rumor TB and two corresponding

[0061]

number

[0062] Matrix-weighted intra-prediction (MIP) Matrix-weighted intra-prediction (MIP) is a new intra-prediction technique added to VVC. To predict samples of a rectangular block of width W and height H, Matrix-weighted intra-prediction (MIP) uses as input a line of reconstructed adjacent boundary samples of height H located on the left side of the block and a line of reconstructed adjacent boundary samples of width W located at the top of the block. If reconstructed samples are unavailable, they are generated in the same way as conventional intra-prediction. Figure 6 shows the Matrix-weighted intra-prediction process. The generation of the prediction signal is based on three steps: averaging, matrix-vector multiplication, and linear interpolation, as shown in Figure 6.

[0063] Decoder-side intra-mode derivation (DIMD) Applying Decoder-Side Intra-Mode Derivation (DIMD) derives two intra-modes from the reconstructed adjacent samples, and these two predictors are combined with the planar mode predictor using weights derived from the gradient, as described in JVET-O0449. The division operation in weight derivation is performed using the same lookup table (LUT)-based integerization scheme used by CCLM. For example, division operation in orientation calculation. Orient=G y / G x This is calculated using the following LUT-based method: x = Floor(Log2(Gx)) normDiff=((Gx<<4)>>x)&15 x + = (3 + (normDiff != 0) ? 1 : 0) Orient=(Gy * (DivSigTable[normDiff]|8)+(1<<(x-1)))>>x In the above equation, DivSigTable

[16] ={0,7,6,5,5,4,4,3,3,2,2,1,1,1,1,0} That is the case.

[0064] The derived intra-mode is included in the primary list of most probable modes (MPMs), so the DIMD process is executed before the MPM list is constructed. The primary derived intra-mode of a DIMD block is stored with the block and used to construct the MPM lists of adjacent blocks.

[0065] Figure 7 shows the HoG calculation from pixels of a template with a width of 3.

[0066] Fusion technology for template-based intra-mode derivation (TIMD) In each intra-prediction mode in MPM, the SATD between the template's predicted sample and the reconstructed sample is calculated. The first two intra-prediction modes with the minimum SATD are selected as TIMD modes. These two TIMD modes are fused with weights after applying PDPC processing, and such weighted intra-predictions are used for encoding the current CU. Position-dependent intra-prediction combinations (PDPCs) are included in the derivation of the TIMD modes.

[0067] The costs of the two selected modes are compared to a threshold, and in the test, a cost factor of 2 is applied as follows: costMode2<2 * costMode1

[0068] If this condition is true, fusion is applied; otherwise, only mode 1 is used.

[0069] The mode weights are calculated from the SATD costs as follows: weight1=costMode2 / (costMode1+costMode2); weight2 = 1 - weight1

[0070] Division operations are performed using the same lookup table (LUT)-based integerization scheme used by CCLM.

[0071] Low-frequency non-separated conversion (LFNST) In VVC, as shown in Figure 8, LFNST is applied between the forward linear transformation and quantization (on the encoder side) and between the inverse quantization and inverse linear transformation (on the decoder side). Figure 8 illustrates the Low Frequency Non-Separated Transform (LFNST) process. In LFNST, either a 4x4 non-separated transform or an 8x8 non-separated transform is applied depending on the block size. For example, a 4x4 LFNST is applied to smaller blocks (i.e., when the minimum (width, height) < 8), and an 8x8 LFNST is applied to larger blocks (i.e., when the minimum (width, height) > 4).

[0072] The application of the non-separated transform used in LFNST is explained below using an input example. To apply 4x4 LFNST, use a 4x4 input block X:

number

number

number

[0073] Inseparable transformation is

number

number

number

[0074] Reduction of non-separable transformations LFNST (Low Frequency Non-Separable Transform) is based on a direct matrix multiplication method to apply non-separable transform and is carried out in a single process without requiring multiple iterations. However, in order to minimize the computational complexity and the memory capacity for storing transform coefficients, the dimension of the non-separable transform matrix needs to be reduced. Therefore, in LFNST, a reduction method of non-separable transform (i.e., RST) method is used. The main idea of reducing non-separable transform is to map an N-dimensional vector (where N is usually equal to 64 in 8×8 NSST) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction coefficient. Therefore, instead of an N×N matrix, the RST matrix becomes an R×N matrix as follows.

[0075]

Number

[0076] LFNST conversion selection In LFNST, there are a total of four transformation sets, and two inseparable transformation matrices (kernels) are used for each transformation set. As shown in Figure 9A, the mapping from intra-prediction modes to transformation sets is predefined. Figure 9A shows the transformation method selection table. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_l_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transformation set 0 is selected for the current chroma block. In each transformation set, the selected inseparable quadratic transformation candidate is further specified by an explicitly signaled LFNST index. The index is signaled in the bitstream once per intra-CU after the transformation coefficients.

[0077] Figure 9A shows the conversion method selection table.

[0078] LFNST Index Signaling and Interaction with Other Tools Because LFNST is limited to being applicable only when all coefficients except the first coefficient subgroup are non-significant, LFNST index coding depends on the position of the last significant coefficient. Furthermore, LFNST indices are context-coded but independent of the intra-prediction mode, and only the first binary is context-coded. In addition, LFNST applies to intra-CU in both intra-slice and inter-slice, and to both lumar and chroma. When dual-tree is enabled, LFNST indices for lumar and chroma are signaled separately. In inter-slice (where dual-tree is disabled), only one LFNST index is signaled and used for both lumar and chroma components.

[0079] Given the existing maximum conversion size limit (64x64), and considering that large CUs exceeding 64x64 are implicitly partitioned (TU tiling), LFNST index lookup can quadruple data buffering for a given number of decoding pipeline stages. Therefore, the maximum size allowed for LFNST is limited to 64x64. Note that LFNST is only enabled in DCT2. LFNST index signaling is placed before the MTS index signal.

[0080] The use of scaling matrices in perceptual quantization is not clearly applicable to LFNST coefficients, as it is not apparent that scaling matrices specified for linear matrices may be useful for LFNST coefficients. Therefore, the use of scaling matrices for LFNST coefficients is not permitted. Chroma LFNST is not applied in single tree partitioning mode.

[0081] Intra-encoding extended multiple transform selection (MTS) In the current VVC design, only the DST7 and DCT8 translation kernels used for intra-coding and inter-coding are utilized in the MTS.

[0082] Additional linear transformations employed include DCT5, DST4, DST1, and the identity transformation (IDT). The MTS set is also configured depending on the TU size and intra-mode information. Sixteen different TU sizes are considered, and for each TU size, five classes are considered depending on the intra-mode information. For each class, one, four, or six different transformation pairs are considered. The number of intra-MTS candidates is adaptively selected (from one, four, or six MTS candidates) based on the sum of the absolute values ​​of the transformation coefficients. To determine the total number of allowed MTS candidates, the sum is compared against two fixed thresholds. One candidate: Total <= th0 4 candidates: th0 < total <= th1 6 candidates: Total > th1

[0083] While a total of 80 different classes are considered, it should be noted that many of these different classes share the exact same set of transformations. Therefore, the resulting LUT contains 58 (less than 80) unique entries.

[0084] In the case of angular modes, the co-symmetry between the TU shape and intra-prediction is considered. Therefore, mode i (i>34) with TU shape A×B is mapped to the same class as mode j=(68-i) with TU shape B×A. However, for each transformation pair, the order of the horizontal and vertical transformation kernels is swapped. For example, a 16×4 block with mode 18 (horizontal prediction) and a 4×16 block with mode 50 (vertical prediction) are mapped to the same class, however, the vertical and horizontal transformation kernels are swapped. In the case of wide-angle modes, the conventional angular mode closest to determining the transformation set is used. For example, mode 2 is used for all modes from -2 to -14. Similarly, mode 66 is used for modes 67 to 80.

[0085] Inter-Multiplexing Selection (MTS) Optimization For inter-encoded CUs, four candidates are used for each CU: {(DST7, DST7), (DST7, DCT8), (DCT8, DST7), (DCT8, DCT8)}. For higher resolution sequences (width > 1080), the maximum CU size when using Inter-MTS is set to 32 (i.e., Inter-MTS is used for CUs with width <= 32 and height <= 32), and for the remaining sequences (lower resolution), the maximum CU size is set to 16. For 4pt, 8pt, and 16pt conversions, the current AMT conversion cores, namely DST-7 and DCT-8, are replaced with separable KLTs as proposed in JVET-J0021.

[0086] Related intrablock copy and intrablock copy method based on template matching Intra template matching Intra-Template Matching Prediction (Intra TMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame where the L-shaped template matches the current template. Within a given search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the decoder performs a similar prediction operation.

[0087] The prediction signal is generated by matching the L-shaped causal neighbor of the current block with another block within a predetermined search area, as shown in Figure 9B. ●R1: Current CTU, ●R2: Upper left CTU, ●R3: Upper CTU, ●R4: Left CTU.

[0088] The sum of absolute differences (SAD) is used as the cost function.

[0089] Within each region, the decoder searches for the template with the minimum SAD (Same Address Distance) with the current block, and uses the block corresponding to that template as the predicted block.

[0090] The dimensions of all regions (SearchRange_w, SearchRange_h) are proportional to the block dimensions (BlkW, BlkH), and are set so that the number of SAD comparisons per pixel remains constant.

[0091] In other words, ●SearchRange_w=a*BlkW, ●SearchRange_h=a*BlkH. In the above equation, "α" is a constant used to control the balance between gain and complexity. In practice, "α" is equal to 5.

[0092] The intra-template matching tool is effective for CUs (Card Units) with a width and height of 64 or less. This maximum CU size for intra-template matching is configurable.

[0093] If DIMD is not used for the current CU, the intra-template matching prediction mode is signaled at the CU level through a dedicated flag.

[0094] Intrablock copying (IBC) using template matching In IBC, template matching is used in both IBC merge mode and IBC AMVP mode.

[0095] The IBC-TM merge list is modified according to the reduction method, similar to the normal TM merge mode, so that candidates are selected based on the distance of movement between candidates, compared to the one used by the normal IBC merge mode. The realization of the final zero movement is replaced with movement vectors to the left (-W, 0), up (0, -H), and upper left (-W, -H), where W is the width of the current CU and H is the height.

[0096] In IBC-TM merge mode, selected candidates are scrutinized using a template matching method before RDO or decoding. IBC-TM merge mode is in conflict with normal IBC merge mode, and the TM merge flag is signaled.

[0097] In IBC-TM AMVP mode, up to three candidates are selected from the IBC-TM merge list. Each of these three selected candidates is refined using a template matching method and sorted according to the resulting template matching cost. Then, in the motion estimation process, only the first two are considered as usual.

[0098] The refinement of template matching in both IBC-TM merge mode and AMVP mode is quite simple because the IBC motion vectors are constrained to (i) be integers and (ii) fit within the reference region. Therefore, in IBC-TM merge mode, all refinement processes are performed with integer precision, and in IBC-TM AMVP mode, they are performed with either integer precision or 4-pixel precision depending on the AMVR value. Such refinement accesses only samples without interpolation. In either case, the refined motion vectors and templates used in each refinement step must adhere to the reference region constraints.

[0099] Figure 10 shows the IBC reference region corresponding to the current CU position.

[0100] IBC Reference Area Figure 11 shows the IBC reference area when CTU(m,n) is encoded. In Figure 11, blocks with the symbol "++" at the top represent the current CTU, blocks with an asterisk * represent a reference area, and the remaining blocks in Figure 11 indicate that they are not valid reference area blocks.

[0101] The IBC reference area extends up to the CTU two rows above. Figure 11 shows the reference area for encoding CTU(m, n). Specifically, when encoding CTU(m, n), the reference area includes CTUs with indices (m-2, n-2)...(W, n-2), (0, n-1)...(W, n-1), (0, n)...(m, n), where W represents the maximum horizontal index value in the current tile, slice, or picture. If the CTU size is 256, the reference area is limited to the CTU one row above. This setting ensures that the IBC does not require additional memory on the current ETM platform when the CTU size is 128 or 256. The range of the sample-by-sample block vector search (also called local search) is limited horizontally to [-(C<<1), C>>2] and vertically to [-C, C>>2] to accommodate the expansion of the reference area, where C represents the CTU size.

[0102] Reconstruction-rearrangement IBC (RR-IBC) The Reconstruct-Rearrange IBC (RR-IBC) mode is allowed for IBC encoded blocks. When RR-IBC is applied, the samples within the reconstructed block are inverted according to the inversion type of the current block. On the encoder side, the original block is inverted before motion search and residual calculation, while the predicted block is derived without inversion. On the decoder side, the reconstructed block is inverted and the original block is restored.

[0103] For RR-IBC encoded blocks, two inversion schemes are supported: horizontal inversion and vertical inversion. Syntax flags are initially signaled to the IBC AMVP encoded block to indicate whether the reconstruction is inverted; if so, another flag is further signaled to identify the inversion type. In IBC merges, the inversion type is inherited from adjacent blocks, and no syntax signaling occurs. Given horizontal or vertical symmetry, the current block and reference block are typically aligned horizontally or vertically. Therefore, when horizontal inversion is applied, the vertical component of the BV is not signaled and is inferred to be equal to 0. Similarly, when vertical inversion is applied, the horizontal component of the BV is not signaled and is inferred to be equal to 0.

[0104] Figures 12A and 12B show examples of BV adjustment in the case of (a) horizontal inversion shown in Figure 12A and (b) vertical inversion shown in Figure 12B.

[0105] To more effectively utilize the symmetry property and refine the block vector candidates, a flip-aware BV adjustment method that takes inversion into account is applied. For example, as shown in Figures 12A and 12B, (xnbr, ynbr) represent the center sample coordinates of the adjacent block, (xcur, ycur) represent the center sample coordinates of the current block, BVnbr represents the BV of the adjacent block, and BVcur represents the BV of the current block. Instead of directly inheriting the BV from the adjacent block, if the adjacent block is encoded with a horizontal inversion, the horizontal component of BVcur is calculated by adding a motion shift to the horizontal component of BVnbr (represented as BVnbrh), i.e., BVcurh = 2(xnbr - xcur) + BVnbrh. Similarly, if adjacent blocks are encoded with a vertical inversion, the vertical component of BVcur is calculated by adding a motion shift to the vertical component of BVnbrv (represented as BVnbrv), i.e., BVcurv = 2(ynbr - ycur) + BVnbrv.

[0106] IBC merge mode using block vector difference (IBC-MBVD) Affine MMVD and GPM-MMVD were adopted in ECM as extensions to the standard MMVD mode. Extending the MMVD mode to IBC merge mode was a natural progression.

[0107] In IBC-MBVD, the distance settings are {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 ​​pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD direction is two horizontal directions and two vertical directions.

[0108] The base candidates are selected from the first five in the rearranged IBC merge list. Then, all possible MBVD refinement positions (20x4) for each base candidate are rearranged based on the SAD cost between the template of each refinement position (the block one row above and one column to the left of the current block) and its reference image. Finally, the top eight refinement positions with the lowest template SAD costs are retained as available positions and used for MBVD index coding. The MBVD index is then binaryized using a rice code scheme where the parameter is equal to 1.

[0109] Blocks encoded using the IBC-MBVD scheme do not inherit the inversion type from adjacent blocks encoded using the RR-IBC scheme.

[0110] As similarly mentioned above, block vector-guided cross-component prediction uses non-local regions to improve the prediction performance of cross-component models when intra-block copies are used. When identical blocks in a reference channel are encoded using an intra-block copy (IBC) scheme, the associated block vector (BV) is used to indicate the reference region for calculating the parameters of the cross-component prediction model.

[0111] When multiple blocks with associated block vectors are located in the same region, the selection of a reference sample can be performed in several ways. The reference sample must also reflect the intensity distribution of samples within the same block. Inappropriately selected reference samples directly lead to decreased performance in cross-component prediction.

[0112] Exemplary embodiments of the present invention provide several improvements to block vector-guided cross-component prediction. These improvements address the selection of reference samples, the processing of multiple block vectors within the same location area, the processing of overlapping reference samples, and robust reference sample selection based on identical luminance sample values.

[0113] Before describing in detail the exemplary embodiments disclosed herein, Figure 16 shows a simplified block diagram of various electronic devices suitable for use in carrying out exemplary embodiments of the present invention.

[0114] Figure 16 is a block diagram of one possible, non-limiting, exemplary system in which exemplary embodiments may be implemented. In Figure 16, as shown in Figure 16, user equipment (UE) 10 is communicating wirelessly with wireless network 1 or network 1. Wireless network 1 or network 1 in Figure 16 may comprise a communication network such as a mobile network, for example, mobile network 1 or first mobile network disclosed herein. References to wireless network 1 in Figure 16 herein can be considered as references to any wireless network disclosed herein. Furthermore, wireless network 1 in Figure 16 may also comprise hardwired functionality if required by the communication network. The UE is a wireless, typically mobile device, capable of accessing the wireless network. For example, the UE may be a mobile phone (or "cellular" phone) and / or a computer with mobile terminal capabilities. For example, the UE or mobile terminal may be a portable, pocket-sized, handheld, computer-integrated, or vehicle-mounted mobile device that performs language signaling and / or data exchange with the RAN.

[0115] UE10 includes one or more processors DP10A, one or more memory MEM10B, and one or more transceivers TRANS10D, which are interconnected via one or more buses. Each of the one or more transceivers TRANS10D includes a receiver and a transmitter. The one or more buses may be an address bus, a data bus, or a control bus, and may include a series of lines on a motherboard, or any interconnection mechanism such as an integrated circuit, optical fiber, or other optical communication equipment. Each of the one or more transceivers TRANS10D can optionally be connected to one or more antennas for communication with NN12 and NN13, respectively. One or more memory MEM10B contains computer program code PROG10C. UE10 communicates with NN12 and / or NN13 via wireless links 11 or 16.

[0116] NN12 (NR / 5G Node B, Evolutionary NB, or LTE device) is a network node, such as a master or secondary node base station (for NR or LTE Long Term Evolution), that communicates with devices such as NN13 and UE10 in Figure 16. NN12 enables wireless devices such as UE10 to access the wireless network 1. NN12 includes one or more processors DP12A, one or more memory MEM12B, and one or more transceivers TRANS12D, which are interconnected via one or more buses. According to an exemplary embodiment, these TRANS12D may include X2 and / or Xn interfaces used to perform the exemplary embodiment. Each of the one or more transceivers TRANS12D includes a receiver and a transmitter. One or more transceivers TRANS12D can optionally be connected to one or more antennas to communicate with UE10 via at least link 11. One or more memory MEM12B and computer program code PROG12C, together with one or more processors DP12A, are configured to cause NN12 to perform one or more of the operations described herein. NN12 can communicate with other gNB or eNB devices, or devices such as NN13, for example, via link 16. Furthermore, links 11, 16, and / or other links can be wired or wireless, or both, and can implement, for example, X2 or Xn interfaces. Additionally, links 11 and / or 16 may be configured through other network devices, such as NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF14 devices, as shown in Figure 16, but are not limited to these. NN12 can perform functions of an MME (Mobility Management Entity) or SGW (Service Gateway), such as user plane functions and / or access management functions for LTE, and similar functions for 5G.

[0117] NN13 may be for WiFi or Bluetooth, or other wireless devices associated with mobility function devices such as AMF or SMF, and further, NN13 may include base stations such as master or secondary node base stations (e.g., for NR or LTE Long Term Evolution) that communicate with devices such as NR / 5G node B, or potentially evolved NB, and NN12 and / or UE10 and / or wireless network 1. NN13 includes one or more processors DP13A, one or more memory MEM13B, one or more network interfaces, and one or more transceivers TRANS13D, which are interconnected via one or more buses. According to an exemplary embodiment, these network interfaces of NN13 may include X2 and / or Xn interfaces used to perform the exemplary embodiment. Each of the one or more transceivers TRANS13D includes a receiver and a transmitter, which are optionally connectable to one or more antennas. One or more memory MEM13B includes computer program code PROG13C. For example, one or more memory MEM13B and computer program code PROG13C, together with one or more processors DP13A, are configured to cause NN13 to perform one or more of the operations described herein. NN13 can communicate with other mobility function devices and / or eNBs, such as NN12 and UE10, or any other devices, using, for example, link 11 or link 16, or other links. Link 16, shown in Figure 16, can be used for communication with NN12. These links can be wired or wireless, or both, and can implement, for example, X2 or Xn interfaces. Furthermore, as stated above, links 11 and / or link 16 may be configured through other network devices, such as NCE / MME / SGW devices, such as NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF14, shown in Figure 16, but not limited to these.

[0118] One or more buses in the device shown in Figure 16 may be address buses, data buses, or control buses, and may include a series of lines on a motherboard, or any interconnection mechanism such as integrated circuits, optical fibers, other optical communication equipment, or wireless channels. For example, one or more transceivers TRANS12D, TRANS13D, and / or TRANS10D may be implemented as a remote radio head (RRH), with the other elements of NN12 located in a physically separate location from the RRH, and these devices may include one or more buses, some of which may be implemented as optical fiber cables, to connect the other elements of NN12 to the RRH.

[0119] Figure 16 shows network nodes such as NN12 and NN13, but it should be noted that any of these nodes can incorporate or be incorporated into eNodeB, eNB, or gNB for LTE, NR, etc., and are still configurable to perform the exemplary embodiment.

[0120] While the descriptions herein indicate that a “cell” performs a function, it should be clear that the gNB and / or user equipment and / or mobility management device forming the cell perform the function. Furthermore, a cell constitutes part of a gNB, and there may be multiple cells in a single gNB.

[0121] Wireless network 1 or any network it may represent may include or may not include NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF14, which may include (NCE) Network Control Element functions, MME (Mobility Management Entity) / SGW (Service Gateway) functions, and / or Service Gateway (SGW), as well as / or MME (Mobility Management Entity) and / or SGW (Service Gateway) functions, as well as / or User Data Management Function (UDM), as well as / or PCF (Policy Control) functions, as well as / or Access and Mobility Management Function (AMF) functions, as well as / or Session Management (SMF) functions, as well as / or Location Management Function (LMF), as well as / or Authentication Server (AUSF) functions, and provide connectivity to further networks such as telephone networks and / or data communication networks (e.g., the Internet), and are configured to perform 5G and / or NR operations in addition to or instead of other standard operations as of the present filing date. NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF14 can be configured to perform the operations described in the exemplary embodiments in any communication technology, including communication technologies based on LTE, NR, 5G, and / or any standards that are being implemented or discussed at the time of this filing. Furthermore, it should be noted that the operations described in the exemplary embodiments performed by NN12 and / or NN13 are also possible in NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF14.

[0122] The NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF14 includes one or more processors DP14A, one or more memory MEM14B, and one or more network interfaces (N / WI / F), which are interconnected via one or more buses and coupled to link 13 and / or link 16. According to an exemplary embodiment, these network interfaces may include X2 and / or Xn interfaces used to perform the exemplary embodiment. One or more memory MEM14B includes computer program code PROG14C. One or more memory MEM14B and computer program code PROG14C, together with one or more processors DP14A, are configured to cause the NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF14 to perform one or more operations that may be necessary to support the operation according to the exemplary embodiment.

[0123] It should be noted that NN12 and / or NN13 and / or UE10 can be configured (for example, based on standard implementations) to perform location management function (LMF) functions. The LMF function may be embodied in any of these network devices or any other device associated with these devices. Furthermore, at least the LMFs described later, such as the LMF of MME / SGW / UDM / PCF / AMF / SMF / LMF14 shown in Figure 16, may be located in the same position as UE10, separated from NN12 and / or NN13 in Figure 16, in order to perform the operations according to the exemplary embodiments disclosed herein.

[0124] Wireless Network 1 can implement network virtualization, which is the process of combining hardware and software network resources and network functions into a virtual network, a single software-based management entity. Network virtualization includes platform virtualization and is often combined with resource virtualization. Network virtualization is classified into external types, which consolidate many networks, or parts of networks, into a single virtual unit, and internal types, which provide network-like functionality to software containers on a single system. It should be noted that the virtualized entities resulting from network virtualization still have technical effects because they are implemented to some extent using hardware such as processors DP10, DP12A, DP13A, and / or DP14A, and memory MEM10B, MEM12B, MEM13B, and / or MEM14B.

[0125] The computer-readable memories MEM12B, MEM13B, and MEM14B may be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The computer-readable memories MEM12B, MEM13B, and MEM14B may also be means for performing storage functions. The processors DP10, DP12A, DP13A, and DP14A may be of any type suitable for the local technical environment and, as non-limiting examples, may include one or more general-purpose computers, dedicated computers, microprocessors, digital signal processors (DSPs), and processors based on multicore processor architectures. The processors DP10, DP12A, DP13A, and DP14A may be means for performing functions such as controlling UE10, NN12, NN13, and other functions described herein.

[0126] Generally, various embodiments of any of these devices, though not limited to them, may include cellular phones such as smartphones, tablets, personal digital assistants (PDAs) with wireless communication capabilities, portable computers with wireless communication capabilities, image capture devices such as digital cameras with wireless communication capabilities, gaming devices with wireless communication capabilities, music storage and playback devices with wireless communication capabilities, internet devices enabling wireless internet access and browsing, tablets with wireless communication capabilities, and portable units or terminals incorporating combinations of such capabilities.

[0127] Furthermore, various embodiments of any of these devices can be used in UE vehicles associated with a ground network, high-altitude platform stations, or any other such type of node, or in any drone-type radio, aircraft or other aircraft-mounted radio, or on waterborne vehicles such as ships.

[0128] As similarly described above, exemplary embodiments of the present invention offer several improvements to block vector-guided cross-component prediction. These improvements include the selection of reference samples, the processing of multiple block vectors within the same location area, the processing of overlapping reference samples, and the selection of robust reference samples based on identical luminance sample values.

[0129] In one embodiment, the reference area used to derive the cross-component model can be determined using all available block vectors present in the lumar coding unit or prediction unit at the same location (an example is shown in Figure 13). Figure 13 shows the chroma PU and the lumar PU at the same location.

[0130] In one embodiment, the reference area used to derive the cross-component model can be determined using only a subset of available block vectors within the same lumen PU (as shown in Figure 13). For example, only C, TL, TR, BL, and BR may be considered. The selection of the subset can also be inferred based on local characteristics such as the size and shape of the PU, or it can be signaled to the decoder by the encoder.

[0131] In one embodiment, if multiple block vectors are available in the same lumen PU, the reference sample area can be derived in such a way that redundant overlapping areas are discarded. Figure 14A shows two block vectors from coding units C and TL at the same location, pointing to different reference sample areas. The overlapping area (shaded portion) is considered only once. For example, as shown in Figure 14A, if two or more block vectors point to the same reference sample, the overlapping area (shaded portion) is considered only once. In addition to deriving a more accurate cross-component prediction model, removing redundant overlapping samples may also reduce computational complexity, particularly in the calculation of decoder-side parameters.

[0132] In one embodiment, if multiple block vectors are available in the same PU, the average of those block vectors can be used as the block vector pointing to the reference sample.

[0133] In one embodiment, based on the previously described embodiment, the mean value is applied only when the difference between block vectors is considered to be sufficiently small.

[0134] In one embodiment, based on the previously described embodiment, the average value can be calculated as a weighted average based on the size of PUs at the same location, thereby assigning a greater weight to the block vectors of larger PUs. Alternatively, the average value can be calculated based on a fixed grid of block vectors. For example, each 4x4 block within a block at the same location can be identified by its block vector, and the average of those block vectors can be determined to be the average block vector used to determine the location of the reference sample.

[0135] In one embodiment, if multiple block vectors are available in a PU at the same location, a linear or nonlinear estimation method can be used to derive the block that points to the reference sample. For example, the median, minimum, or maximum value of the block vector can be considered. The choice of operators (median, minimum, maximum, etc.) may be the same or different for the horizontal and vertical components of the block vector.

[0136] Figure 14B shows two block vectors from coding units C and TL at the same location, indicating that the maximum and minimum coordinate values ​​of a composite reference area defined by a block vector at the same location and a PU area at the same location are used to determine the reference area.

[0137] In one embodiment shown in Figure 14B, when multiple block vectors are available in a PU at the same location, the maximum and minimum coordinate values ​​of a composite reference area defined by the block vectors and PU areas at the same location are used to determine the reference area. For example, the left boundary of the reference area can be determined to have the minimum horizontal coordinate of the composite area, and the right boundary of the reference area can have the maximum horizontal coordinate of the composite area. Similarly, the top and bottom boundaries of the reference area can be determined using the minimum and maximum vertical coordinates of the composite area. If the reference area thus created is larger or smaller than the PU itself, the dimensions of the reference area can be adjusted to match the dimensions of the PU.

[0138] In one embodiment, if no block vector is associated with all identically located PUs, a temporary block vector can be determined for identically located PUs that do not have their own block vector. Such a temporary block vector can be determined, for example, by duplicating the block vector of a selected PU in a reference channel, or by interpolating a temporary block vector from the block vector of a selected PU in a reference channel. The temporary block vector can then be used to determine the reference area, just as a block vector obtained directly from an identically located PU.

[0139] In one embodiment, block vectors within the same PU and their spatial positions within the same-location area can be used to interpolate or estimate refined block vectors. For example, a planar or quadratic model can be used to derive additional block vectors at the center of the same-location lumen area. Such estimations can be performed, for example, using a linear regression solver. The derived block vectors are then used as pointers to reference samples for deriving the cross-component model.

[0140] In one embodiment, when multiple block vectors are available in the same PU, a template matching-based refinement mechanism can be used to find the best matching reference block and its corresponding block vector. Template matching-based refinement may use some or all of the samples in the same reference block and / or some or all of the adjacent reference samples in the current block.

[0141] According to the embodiments described above, template matching-based refinement can use one or more block vectors available from the same PU as the initial block vector in the refinement search. Alternatively, or additionally, the initial block vector or multiple vectors can be derived using other embodiments described herein.

[0142] In one embodiment, if multiple block vectors are available in a PU at the same location, the chroma PU can be divided into multiple cross-component models based on the spatial position of the PU at the same location.

[0143] Figure 15 shows two block vectors from prediction units C and TL at the same location, pointing to different reference sample areas. The chroma PU can be divided into two cross-component models (0 and 1) based on the spatial location of the PU to which the block vectors (bv0 and bv1) belong.

[0144] For example, as shown in Figure 15, block vector 0 is used to derive the cross-component model of the upper half of the chroma PU (the part marked with 0), and similarly, block vector 1 is used to derive the cross-component model of the lower half of the chroma PU (the part marked with 1).

[0145] In one embodiment, if multiple block vectors are available in the same PU, multiple BVs can be used to compute multiple cross-component models. The final prediction can then be obtained by combining the multiple prediction results. The weights for combining the multiple predictions may be defined in the codec specification, or the weights or weight identifier indices may be signaled in the bitstream, or they may be calculated on the decoder side based on block information such as reconstructed sample values ​​and block size.

[0146] In one embodiment, statistics of samples from blocks at the same location can be used to reduce the number of training samples in a reference block pointed to by a block vector. For example, the minimum and maximum intensity values ​​of samples in a block at the same location can be determined and used to reduce the number of samples so that those within the minimum and maximum intensity ranges are considered for parameter calculation or training of a cross-component predictive model. To provide a wider intensity range for training, the calculated minimum and maximum intensity values ​​can be extended by delta values ​​from the upper and lower limits of the range, which may be fixed or determined by the calculated minimum and maximum values. For example, the delta values ​​may be a constant percentage of the minimum and / or maximum values. In another example, the delta values ​​used to extend the upper and lower limits may be the difference between the minimum and maximum intensity values, or a modified version of the difference between the minimum and maximum intensity values ​​of block samples at the same location.

[0147] In one embodiment, the type of cross-component model can be determined based on the distribution of samples from blocks and / or reference blocks at the same location indicated by block vectors. For example, the decision of whether to use prediction by a single model or by multiple models can be made based on the distribution of samples with respect to the classification parameters of the cross-component model. Cross-component models typically use the mean of the training samples as classification parameters for computing multiple models. If the proportions of the models' samples are not evenly distributed based on the classification parameters, a variation of a single model may be determined. Alternatively, if the proportions of the models' samples are evenly distributed based on the classification parameters, a variation of multiple models may be determined for cross-component prediction. The sample proportions may be predetermined and may be signaled in the bitstream.

[0148] According to the embodiments described above, the model type may be determined by other parameters in addition to, or instead of, minimum and maximum values. For example, the mean value of a sample may be used. Samples that fall within a certain intensity distance of the mean value from upper and lower limits may be used to calculate the parameters. The intensity distances for determining the upper and lower limits may be the same or different. The intensity distance may be defined as a constant percentage of the mean value. The intensity distance value, or an indicator thereof, may be signaled in the bitstream.

[0149] In one embodiment, if multiple block vectors are available in the same PU, the reference block for model derivation is determined by finding the area that best matches the block at the same location within the reference channel. The search process can use one or more of the available block vectors. The search process for the best matching area may use cost calculation metrics such as the sum of absolute differences (SAD), sum of squared errors (SSE), sum of transformation errors (SATD), or any other metric.

[0150] In another embodiment, one or more correlation indices may be used to select a sample from a reference area pointed to by a block vector of PUs at the same location. For example, the correlation coefficient can be used as a distortion index, as previously proposed.

[0151] The following details how to use the correlation coefficient to select the sample with the highest correlation to the block at the same location within the reference channel.

[0152] Both consider two signals X and Y of length n. Here, we list five sums based on X and Y. S x =ΣX S y =ΣY S xx =ΣXX S yy =ΣYY S xy =ΣXY Furthermore, Pearson's correlation coefficient can be expressed as follows:

[0153]

number

[0154] The denominator in the above formula (the product of the standard deviations of X and Y) normalizes the covariance of X and Y to the range [-1, 1], so that P is useful for evaluating the correlation between X and multiple Y values. The following integer operations can be used:

number

[0155] Correlation cost is defined as follows: C(X, Y) = abs(P(X, Y)) In the above equation, a larger value of C indicates better alignment. Many existing coding tools, such as VVC, perform motion compensation and template matching by minimizing SSE or SAD using integer arithmetic. Correlation costs can be easily incorporated into the above procedure by considering the following:

[0156]

number

number

number

[0157] The described process iterates through different areas pointed to by block vectors from the same location PU and is used to compute a cross-component prediction model using the area showing the highest correlation.

[0158] Instead of using different block vectors, or in addition to them, refinement can be applied to one or more block vectors at the same location from the same PU. Refinement involves adding a delta to one or both directions of the BV relative to the initial block vectors and calculating the correlation cost of the areas associated with the refined BV. The optimal areas and BVs are then selected and used to calculate the parameters for cross-component prediction.

[0159] Figure 17 illustrates, but is not limited to, operations that may be performed by a device, such as UE10 in Figure 16. As shown in step 1710 of Figure 17, the video decoder may determine a reference sample area of ​​video clip samples for deriving the cross-component model. As shown in step 1720 of Figure 17, the determination is to use an identified subset of available block vectors in one of the colocation lumar coding units or colocation lumar prediction units. As shown in step 1730 of Figure 17, the subset is identified based on local properties. As shown in step 1740 of Figure 17, a cross-component prediction model may be obtained based on the determination. As shown in step 1750 of Figure 17, the cross-component prediction model may be used to decode the video clip samples.

[0160] According to the exemplary embodiments described in the paragraph above, the local property comprises at least one of the size of the prediction unit or the shape of the prediction unit.

[0161] According to the exemplary embodiments described in the paragraph above, at least one of the size or shape of the prediction unit is predetermined by the video decoder or received by the video encoder.

[0162] According to the exemplary embodiment described in the paragraph above, the reference sample area is an intrablock copy reference area.

[0163] According to the exemplary embodiments described in the paragraph above, the reference sample area is determined based on a plurality of block areas available in at least one prediction unit of the same position lumer.

[0164] According to the exemplary embodiments described in the paragraph above, the determination comprises deriving a reference sample area such that any overlapping or redundant areas are discarded.

[0165] According to the exemplary embodiment described in the paragraph above, the average of the available block vectors is used as a block vector that points to a reference sample area.

[0166] According to the exemplary embodiment described in the paragraph above, if the size difference of the block vectors is less than or equal to a threshold, the block vectors point to a reference sample area.

[0167] According to the exemplary embodiments described in the paragraph above, the mean value may be calculated as a weighted average based on the size of one of the colocation lumar coding units or colocation lumar prediction units.

[0168] According to the exemplary embodiment described in the paragraph above, the block vector of at least one prediction unit is assigned a weight based on the size of the prediction unit of at least one prediction unit.

[0169] According to the exemplary embodiment described in the paragraph above, the average value is calculated based on a fixed grid of block vectors.

[0170] According to the exemplary embodiments described in the paragraph above, at least one of a linear or nonlinear estimation method is used to derive a block vector that points to a reference sample area.

[0171] According to the exemplary embodiment described in the paragraph above, based on a plurality of block areas, the maximum and minimum coordinate values ​​of a composite reference area defined by identical position vectors and identical position prediction unit areas are used to determine a reference sample area.

[0172] According to the exemplary embodiments described in the paragraph above, based on the fact that the dimensions of the reference sample area differ from the dimensions of at least one prediction unit, the apparatus adjusts the dimensions of the reference sample area to match the dimensions of at least one prediction unit.

[0173] According to the exemplary embodiments described in the paragraph above, temporary block vectors of the same position prediction unit that do not have associated block vectors may be used to identify temporary block vectors of the same position prediction unit and to determine at least one of the block vectors obtained directly from a reference sample area or the same position prediction unit.

[0174] According to the exemplary embodiments described in the paragraph above, the spatial position of a block vector in one of the co-location lumens coding units or co-location lumens prediction units is used to interpolate or estimate a block vector pointing to a reference sample in a reference sample area for the derivation of the cross-component model.

[0175] According to the exemplary embodiment described in the paragraph above, a template matching-based refinement mechanism is used to find at least one matching reference area block or corresponding block vector from an identified subset of available block vectors, based on the fact that multiple block vectors are available in the same position prediction unit area.

[0176] According to the exemplary embodiments described in the paragraph above, the template matching-based refinement mechanism uses at least one of the at least one adjacent reference samples, or a co-located reference area block within the current block.

[0177] According to the exemplary embodiment described in the paragraph above, based on the availability of multiple block vectors within the same position prediction unit area, the chroma prediction unit is divided into multiple cross-component models based on the spatial position of the same position prediction unit.

[0178] According to the exemplary embodiment described in the paragraph above, block vector 0 is used to derive the cross-component model of the upper half of the chroma prediction unit marked with 0, and block vector 1 is used to derive the cross-component model of the lower half of the chroma prediction unit marked with 1.

[0179] According to the exemplary embodiment described in the paragraph above, based on the availability of multiple block vectors within the same position prediction unit area, multiple predictions are made using multiple block vectors to perform multiple cross-component models, and the final prediction is obtained by combining the multiple predictions.

[0180] According to the exemplary embodiments described in the paragraph above, the weights for combining multiple predictions can be those defined in the codec specification, weight identifier indices signaled to the video decoder, or weights calculated on the decoder side based on block information with reconstructed sample and block size.

[0181] According to the exemplary embodiment described in the paragraph above, statistics of blocks at the same location in multiple block areas are used to reduce the number of training samples in the reference area block of the reference sample area pointed to by at least one block vector.

[0182] According to the exemplary embodiment described in the paragraph above, the minimum and maximum intensity values ​​of samples within a block at the same location are determined, and samples within the minimum and maximum intensity ranges are used to reduce them so that they are considered suitable for at least one of the following: parameter calculation or training of a cross-component prediction model.

[0183] According to the exemplary embodiment described in the paragraph above, in order to provide a wider range for training, the minimum and maximum intensity values ​​are extended by delta values ​​from the upper and lower limits of the maximum intensity range, where the delta values ​​are either fixed or determined by the minimum and maximum intensity values.

[0184] According to the exemplary embodiment described in the paragraph above, the type of cross-component model is determined based on the distribution of samples from blocks and / or reference blocks at the same location indicated by the block vectors.

[0185] According to the exemplary embodiment described in the paragraph above, if the sample ratios of the model are distributed based on classification parameters, then multiple model variations of the cross-component prediction model are determined, and the sample ratios are predetermined or one of those signaled to the decoder.

[0186] According to the exemplary embodiment described in the paragraph above, the type of model is determined by other parameters in addition to, or instead of, minimum and maximum values, wherein the other parameters comprise the mean value of the samples used, the mean value comprises upper and lower limits based on intensity distances used in parameter calculations, and the intensity distance is predetermined or one of those signaled to the decoder.

[0187] According to the exemplary embodiment described in the paragraph above, based on the fact that multiple block vectors are available in the same position prediction unit area, the reference block for model derivation is determined by finding an area that matches a block at the same position in the reference channel.

[0188] According to the exemplary embodiment described in the paragraph above, one or more correlation indices are used to select a sample from a reference sample area pointed to by the block vector of the same position prediction unit.

[0189] A non-temporary computer-readable medium (MEM12B shown in Figure 16) stores program code (PROG10C shown in Figure 16), which is executed by at least one processor (DP10A and / or DP10F shown in Figure 16) to perform at least the operations described in the above paragraph.

[0190] According to the exemplary embodiments of the present invention described above, the video decoder provides means for determining a reference sample area of ​​video clip samples for deriving a cross-component model (as shown in Figure 16, one or more transceivers 12D and / or 13D, MEM12B and / or MEM13B, PROG12C and / or PROG13C, and DP12A and / or DP13A), wherein the determination is to use an identified subset of available block vectors in one of the colocation lumar coding units or colocation lumar prediction units, and the subset (as shown in Figure 16, one or more transceivers 12D and / or 13D, MEM12B and / or MEM13B, PROG12C and / or The apparatus comprises means for identifying PROG13C, and DP12A and / or DP13A) based on local properties, means for obtaining a cross-component predictive model (as shown in Figure 16, one or more transceivers 12D and / or 13D, MEM12B and / or MEM13B, PROG12C and / or PROG13C, and DP12A and / or DP13A) based on determining, and means for using the cross-component predictive model (as shown in Figure 16, one or more transceivers 12D and / or 13D, MEM12B and / or MEM13B, PROG12C and / or PROG13C, and DP12A and / or DP13A) to decode video clip samples.

[0191] In exemplary embodiments of the present invention as described in the paragraphs above, the means for determining, identifying, and using comprises a non-temporary computer-readable medium [MEM12B and / or MEM13B shown in Figure 5] encoded using a computer program [PROG12C and / or PROG13C] executable by at least one processor [DP12A and / or DP13A shown in Figure 16].

[0192] Furthermore, according to exemplary embodiments of the present invention, a circuit is provided for performing the operation according to the exemplary embodiments of the present invention disclosed herein. This circuit may include any type of circuit, such as a content encoding circuit, a content decoding circuit, a processing circuit, an image generation circuit, or a data analysis circuit. Furthermore, this circuit may include discrete circuits, application-specific integrated circuits (ASICs), and / or field-programmable gate array circuits (FPGAs), as well as a dual-core processor with a processor specifically configured to perform its respective functions by software, or a digital signal processor corresponding to the software. Furthermore, necessary inputs to and outputs from the circuit, the functions performed by the circuit, and interconnections (possibly via inputs and outputs) between the circuit and other components, which may include other circuits, are provided for performing the exemplary embodiments of the present invention described herein.

[0193] According to exemplary embodiments of the present invention disclosed herein, the “circuit” provided may include at least one or more or all of the following hardware circuits and / or processors, such as a hardware-only circuit implementation (e.g., an implementation in analog and / or digital circuits only), a combination of hardware circuits and software (if applicable), such as a combination of analog and / or digital hardware circuits and software / firmware, a hardware processor having software (including a digital signal processor) that works in conjunction to cause a device such as a mobile phone or server to perform various functions, such as functions or operations according to exemplary embodiments of the present invention disclosed herein, and (c) a microprocessor or a part of a microprocessor that requires software (e.g., firmware) for operation, but may not be present if such software is not required for operation.

[0194] According to an exemplary embodiment of the present invention, there is circuitry sufficient to perform at least a novel operation according to an embodiment of the present invention disclosed in this application. As used herein, "circuitry" refers to at least the following: (a) Circuit implementations consisting only of hardware (such as implementations in only analog circuits and / or digital circuits), and (b) Combinations of circuitry and software (and / or firmware), for example (where applicable): (i) combinations of processors, or (ii) processors / software (including digital signal processors), software, and a portion of memory that operate in cooperation to cause a device such as a mobile phone or a server to perform various functions, and (c) Circuitry such as a microprocessor or a portion of a microprocessor that requires software or firmware for operation even if the software or firmware does not physically exist.

[0195] This definition of "circuitry" applies to all uses of this term in this application, including the claims. As a further example, when used in this application, the term "circuitry" also encompasses implementations of merely a processor (or processors) or a portion of a processor and its (or their) accompanying software and / or firmware. The term "circuitry" includes, for example, a baseband integrated circuit or an application processor integrated circuit for a mobile phone, or a similar integrated circuit in a server, a cellular network device, or other network devices when applied to a particular claim element.

[0196] Generally, various embodiments may be implemented by hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. Various aspects of the invention may be described and illustrated using block diagrams, flowcharts, or some other graphical representation, but these blocks, devices, systems, techniques, or methods described herein may, by way of non-limiting example, be implemented by hardware, software, firmware, dedicated circuitry or logic, general purpose hardware or controllers, or other computing devices, or any combination thereof.

[0197] Embodiments of the present invention may be implemented in various components such as integrated circuit modules. The design of integrated circuits is generally a highly automated process. Complex and powerful software tools are available for converting a logic level design into a semiconductor circuit design that can be etched and formed on a semiconductor substrate.

[0198] As used herein, the term "exemplary" means "serving as an example, instance, or illustration". Embodiments described herein as "exemplary" should not necessarily be construed as being more preferred or advantageous than other embodiments. All embodiments described in the mode for carrying out the invention are exemplary embodiments provided to enable those skilled in the art to implement or use the invention, and do not limit the scope of the invention defined by the claims.

[0199] The above description provides a complete and useful explanation, using illustrative and non-limiting examples, of the best methods and apparatus currently devised by the inventors for carrying out the present invention. However, reading the above description in conjunction with the accompanying drawings and claims will make it clear to those skilled in the art that various modifications and adaptations are possible. Nevertheless, such changes and similar modifications to the teachings of exemplary embodiments of the present invention remain within the scope of the present invention.

[0200] It should be noted that the terms “connected,” “joined,” or variations thereof, mean any direct or indirect connection or joining between two or more elements, and may also include the presence of one or more intermediate elements between two “connected” or “joined” elements. The joining or connection between elements may be physical, logical, or a combination thereof. In this specification, as some non-exclusive and exhaustive examples, two elements may be considered “connected” or “joined” by one or more conductors, cables, and / or printed wiring connections, as well as by the use of electromagnetic energy, such as electromagnetic energy having wavelengths in the radio frequency domain, microwave domain, and optical (including visible and invisible light) domains.

[0201] Furthermore, some features of preferred embodiments of the present invention can be advantageously used without corresponding to other features. Therefore, the above description is merely illustrative of the principles of the present invention and should not be considered as a limitation thereof.

Claims

1. At least one processor, When executed by the at least one processor, the device The video decoder determines the reference sample area of ​​the video clip samples for deriving the cross-component model, The determination described above involves using an identified subset of available block vectors in one of the colocation lumar coding units or colocation lumar prediction units. The subset is identified based on local properties. To make a decision, Based on the above decision, a cross-component prediction model is obtained, The cross-component prediction model is used to decode the aforementioned video clip samples. At least one non-temporary memory that stores instructions that cause at least one to execute and A device equipped with the following features.

2. The apparatus according to claim 1, wherein the local property comprises at least one of the size of the prediction unit or the shape of the prediction unit.

3. The apparatus according to claim 2, wherein at least one of the sizes or shapes of the prediction unit is predetermined by the video decoder or received by the video encoder.

4. The apparatus according to claim 1, wherein the reference sample area is an intrablock copy reference area.

5. The apparatus according to claim 1, wherein the reference sample area is determined based on a plurality of block areas available in at least one prediction unit of the same position lumer.

6. The apparatus according to claim 5, wherein the determination comprises deriving the reference sample area such that any overlapping or redundant areas are discarded.

7. The apparatus according to claim 5, wherein the average value of the available block vectors is used as a block vector indicating the reference sample area.

8. The apparatus according to claim 7, wherein the block vectors point to the reference sample area when the size difference of the block vectors is less than or equal to a threshold.

9. The apparatus according to claim 7, wherein the average value is calculated as a weighted average based on the size of one of the co-location lumar coding units or co-location lumar prediction units.

10. The apparatus according to claim 9, wherein weights are assigned to the block vectors of the at least one prediction unit based on the size of the prediction unit of the at least one prediction unit.

11. The apparatus according to claim 7, wherein the average value is calculated based on a fixed grid of block vectors.

12. The apparatus according to claim 7, wherein at least one of a linear estimation method or a nonlinear estimation method is used to derive the block vector that points to the reference sample area.

13. The apparatus according to claim 7, wherein the maximum and minimum coordinate values ​​of a composite reference area defined by the same position vector and the same position prediction unit area, based on the plurality of block areas, are used to determine the reference sample area.

14. The apparatus according to claim 13, wherein, based on the fact that the dimensions of the reference sample area differ from the dimensions of the at least one prediction unit, the apparatus adjusts the dimensions of the reference sample area to match the dimensions of the at least one prediction unit.

15. An instruction stored in the at least one non-temporary memory is transmitted to the device. Identifying temporary block vectors of the same position prediction unit that do not have associated block vectors, The temporary block vector is used to determine at least one of the block vectors obtained directly from a reference sample area or the same position prediction unit. The apparatus according to claim 1, which is executed by at least one processor in order to perform the following.

16. The apparatus according to claim 1, wherein the spatial position of a block vector in one of the co-location lumens coding units or the co-location lumens prediction units is used to interpolate or estimate a block vector pointing to a reference sample in the reference sample area for the purpose of deriving the cross-component model.

17. The apparatus according to claim 7, wherein, based on the fact that multiple block vectors are available in the same position prediction unit area, a template matching-based refinement mechanism is used to find at least one of a matching reference area block or a corresponding block vector from the identified subset of the available block vectors.

18. The apparatus according to claim 17, wherein the refinement mechanism based on template matching uses at least one sample from the same-position reference area blocks within the current block, or at least one adjacent reference sample.

19. The apparatus according to claim 7, wherein, if multiple block vectors are available in the same position prediction unit area, the chroma prediction unit is divided into multiple cross-component models based on the spatial position of the same position prediction unit.

20. The apparatus according to claim 19, wherein block vector 0 is used to derive the cross-component model of the upper half of the chroma prediction unit marked with 0, and block vector 1 is used to derive the cross-component model of the lower half of the chroma prediction unit marked with 1.

21. The apparatus according to claim 7, wherein, if multiple block vectors are available in the same position prediction unit area, multiple predictions of multiple cross-component models are made using the multiple block vectors, and a final prediction is obtained by combining the multiple predictions.

22. The apparatus according to claim 21, wherein the weights for combining the multiple predictions may be those defined in the codec specification, weight identifier indices signaled to the video decoder, or weights calculated on the decoder side based on block information comprising reconstructed samples and block sizes.

23. The apparatus according to claim 5, wherein statistics of blocks at the same location in the plurality of block areas are used to reduce the number of training samples in the reference area block of the reference sample area indicated by at least one block vector.

24. The apparatus according to claim 23, wherein the minimum and maximum intensity values ​​of samples in a block at the same location are determined, and samples within the minimum and maximum intensity ranges are used to reduce them so that they are considered to be for parameter calculation or training of the cross-component prediction model, or

25. The apparatus according to claim 24, wherein, in order to provide a wider range for training, the minimum intensity value and the maximum intensity value are extended by delta values ​​from the upper and lower limits of the maximum intensity range, and the delta values ​​are one of a fixed value or determined by the minimum intensity value and the maximum intensity value.

26. The apparatus according to claim 24, wherein the type of the cross-component model is determined based on the distribution of samples from the block and / or reference block at the same location indicated by the block vector.

27. The apparatus according to claim 24, wherein, when the sample ratios of the models are distributed based on classification parameters, the deformations of multiple models of the cross-component prediction model are determined, and the sample ratios are predetermined or one of those signaled to the decoder.

28. The apparatus according to claim 26, wherein the type of the cross-component prediction model is determined by other parameters in addition to or instead of minimum and maximum values, the other parameters comprising the mean values ​​of the samples used, the mean values ​​comprising upper and lower limits based on intensity distances used in the parameter calculation, and the intensity distances being predetermined or one of those signaled to the decoder.

29. The apparatus according to claim 13, wherein, if multiple block vectors are available in the same position prediction unit area, the reference block for model derivation is determined by finding an area that matches a block at the same position in the reference channel.

30. The apparatus according to claim 1, wherein one or more correlation indices are used to select a sample from the reference sample area indicated by the block vectors of the same position prediction unit.

31. The video decoder determines the reference sample area of ​​the video clip samples for deriving the cross-component model, The determination described above involves using an identified subset of available block vectors in one of the colocation lumar coding units or colocation lumar prediction units. The subset is identified based on local properties. To make a decision, Based on the above decision, a cross-component prediction model is obtained, The cross-component prediction model is used to decode the aforementioned video clip samples. A method that includes [a certain feature].