Reference sample selection in block vector guided cross-component prediction
By selecting appropriate reference sample regions and optimizing weight calculation in block vector-guided cross-component prediction, the performance issues caused by improper reference sample selection are resolved, thereby improving the efficiency and quality of video codecs.
Patent Information
- Application Number
- CN202480023783.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-07
- Filing Date
- 2024-03-11
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, improper selection of reference samples in block vector-guided cross-component prediction can lead to performance degradation, affecting the efficiency and quality of video codecs.
By determining the available block vector subsets in the same-position luminance coding unit or the same-position luminance prediction unit, a reference sample region is selected based on local attributes. The video segment sample is decoded using a cross-component prediction model, and the reference sample selection is optimized through template matching and weight calculation.
It improves the accuracy and efficiency of block vector-guided cross-component prediction, reduces redundant areas, and enhances the performance of the video encoding and decoding process.
Smart Images

Figure CN120883604A_ABST
Abstract
Description
Technical Field
[0001] The teachings of exemplary embodiments of the present invention generally relate to at least improved block vector-guided cross-component prediction, and more specifically, to improved block vector-guided cross-component prediction that works at least by utilizing reference sample selection, processing of multiple block vectors in co-location regions, processing of overlapping reference samples, robust reference sample selection based on co-location luminance sample values, etc. Background Technology
[0002] This section is intended to provide background or context for the invention set forth in the claims. The description herein may include concepts that may be pursued, but are not necessarily previously conceived or pursued. Therefore, unless otherwise stated herein, the content described in this section is not prior art to the description and claims of this application, nor should it be considered prior art by virtue of its inclusion in this section.
[0003] Certain abbreviations that may appear in the specification and / or figures are defined herein as follows: AMVR: Adaptive Motion Vector Resolution BV: Block Vector CC: Cross Component CCCM: Cross-component linear model CCLM: Intra-frame prediction using a cross-component linear model CTU: Coding Tree Unit CU: Central Unit IBC: Intra-Block Copy ISP: Intra-Frame Sub-Partition LM: Linear Model LMS: Least Mean Square MRL: Multiple Reference Lines MMLM: Multi-model LM MVD: Motion Vector Difference TM: Template Matching VVC: Multi-functional Video Codec
[0004] Overview of previous developments
[0005] When using intra-block copying, cross-component prediction guided by block vectors for the video codec uses non-local regions to improve the prediction performance of the cross-component model. If the co-occurring block in the reference channel is encoded using the intra-block copying (IBC) method, the associated block vector (BV) is used to indicate the reference region for computing the parameters of the cross-component prediction model.
[0006] When a co-location region contains several blocks with associated block vectors, reference sample selection can be performed in several ways. The reference sample should also reflect the intensity distribution of samples within the co-location block. Inappropriately selected reference samples directly lead to impaired performance in cross-component prediction.
[0007] The exemplary embodiments of the present invention propose at least improved operations for cross-component prediction guided by block vectors. Summary of the Invention
[0008] This section includes examples of possible implementations and is not intended to be restrictive.
[0009] In one example aspect of the invention, there is an apparatus, such as a user equipment side apparatus, comprising: at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: determine, by a video decoder, a reference sample region for video segment samples derived from a cross-component model, wherein the determination uses an identified subset of available block vectors in one of a co-occurrence luminance coding unit or a co-occurrence luminance prediction unit, and wherein the subset is identified based on local attributes; based on the determination, obtain a cross-component prediction model; and use the cross-component prediction model to decode the video segment samples.
[0010] In another example aspect of the invention, there is a method comprising: determining, by a video decoder, a reference sample region for a video segment sample derived from a cross-component model, wherein the determination uses an identified subset of available block vectors in one of a co-occurrence luminance coding unit or a co-occurrence luminance prediction unit, and wherein the subset is identified based on local attributes; obtaining a cross-component prediction model based on the determination; and using the cross-component prediction model to decode the video segment sample. Another example embodiment is an apparatus and method comprising the apparatus and method described in the preceding paragraph, wherein local attributes include at least one of prediction cell size or prediction cell shape, wherein at least one of the prediction cell size or prediction cell shape is predetermined by a video decoder or received from a video encoder, wherein a reference sample region is an intra-frame block copy reference region, wherein the reference sample region is determined based on the availability of more than one block region in at least one prediction cell with co-positional luminance, wherein the determination includes deriving the reference sample region such that overlapping or redundant regions are discarded, wherein the average of available block vectors is used as a block vector pointing to the reference sample region, wherein when the block vector size difference is less than a threshold The block vector points to a reference sample region, wherein the average can be calculated as a weighted average based on the size of one of the corresponding luminance coding units or corresponding luminance prediction units, wherein weights are assigned to the block vector of at least one prediction unit based on the size of the prediction unit in at least one prediction unit, wherein the average is calculated based on a fixed grid of the block vector, wherein at least one of linear or nonlinear estimation is used to derive the block vector pointing to the reference sample region, wherein the maximum and minimum coordinates of a composite reference region defined by the corresponding vector and the corresponding prediction unit region are used to determine the reference sample region based on more than one block region, wherein the size of the reference sample region is different from the size of at least one prediction unit. The apparatus is configured to scale the size of a reference sample region to match the size of at least one prediction unit, wherein there is a temporary block vector identifying a co-position prediction unit that does not have an associated block vector; and to use the temporary block vector to determine at least one of the reference sample region or block vector directly obtained from the co-position prediction unit, wherein the spatial location of the block vector in the co-position region of one of the co-position luminance coding units or co-position luminance prediction units is used to interpolate or estimate a block vector pointing to a reference sample for a reference sample region derived for a cross-component model, wherein based on the availability of multiple block vectors in the co-position prediction unit region, a template matching-based refinement mechanism is used to find at least one of the following: a matching reference region block, or an available block vector. The corresponding block vector in the identified subset of quantities, wherein the template matching-based refinement mechanism uses at least one sample of at least one of the co-location reference region blocks in the current block or at least one of the at least one of the adjacent reference samples, wherein multiple block vectors are available in the co-location prediction unit region, the chroma prediction unit is divided into multiple cross-component models based on the spatial location of the co-location prediction unit, wherein block vector 0 is used to derive the cross-component model of the upper half of the chroma prediction unit marked as 0, and block vector 1 is used to derive the cross-component model of the lower half of the chroma prediction unit marked as 1, wherein multiple block vectors are available in the co-location prediction unit region, and multiple predictions are made for multiple cross-component models using multiple block vectors;The final prediction is obtained by combining multiple predictions, wherein the weights used to combine the multiple predictions can be one of the following: weights defined in the codec specification, or identifier indices of weights signaled to the video decoder, or weights calculated on the decoder side based on block information including reconstructed samples and block size, wherein statistics of co-occurring blocks of more than one block region are used to prune training samples in reference region blocks of reference sample regions pointed to by at least one block vector, wherein the minimum and maximum intensity values of samples in co-occurring blocks are determined and used for pruning such that samples within the minimum and maximum intensity ranges are considered for at least one of the following: parameter calculation, or training a cross-component prediction model, wherein the minimum and maximum intensity values are expanded by incremental values from the lower and upper limits of the maximum intensity range to provide a larger range for training, and wherein the incremental values are one of the following: fixed, or determined by the minimum and maximum intensity values, wherein the cross-component prediction model is used for at least one of the following: parameter calculation, or training a cross-component prediction model, wherein the minimum and maximum intensity values are expanded by incremental values from the lower and upper limits of the maximum intensity range to provide a larger range for training, and wherein the incremental values are one of the following: fixed, or determined by the minimum and maximum intensity values, wherein the cross-component prediction model is used for at least one of the following: parameter calculation, or training a cross-component prediction model. The component model type is determined based on the distribution of samples from the co-occurrence block and / or reference block pointed to by the block vector, wherein the sample ratio is distributed for the classification parameter-based model, and the multi-model variant is determined for the cross-component prediction model, wherein the sample ratio is one of the following: predefined or signaled to the decoder, wherein in addition to or replacing the minimum and maximum values, the model type is determined by other parameters, wherein the other parameters include the mean of the samples used, wherein the mean includes a lower and upper bound based on the intensity distance used for parameter calculation, and wherein the intensity distance is one of the following: predefined or signaled to the decoder, wherein multiple block vectors are available in the co-occurrence prediction unit region, the reference block for model derivation is determined by finding a matching region with the co-occurrence block in the reference channel, and / or wherein one or more correlation metrics are used to select samples from the reference sample region pointed to by the block vector of the co-occurrence prediction unit.
[0011] A non-transitory computer-readable medium storing program code that is executed by at least one processor to perform at least the methods described in the preceding paragraphs.
[0012] In another example aspect of the invention, there is an apparatus comprising: means for determining, by a video decoder, a reference sample region for video segment samples derived from a cross-component model, wherein the determination uses an identified subset of available block vectors in one of a co-occurrence luminance coding unit or a co-occurrence luminance prediction unit, and wherein the subset is identified based on local attributes; means for obtaining a cross-component prediction model based on the determination; and means for decoding the video segment samples using the cross-component prediction model.
[0013] According to the example embodiments described in the preceding paragraphs, at least the components used to determine and obtain include: a network interface, and computer program code stored on a computer-readable medium and executed by at least one processor.
[0014] A communication system includes a network-side device and a user equipment-side device for performing the above operations. Attached Figure Description
[0015] The above and other aspects, features, and benefits of various embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings, wherein like reference numerals are used to denote like or equivalent elements. The accompanying drawings are provided to facilitate a better understanding of the embodiments of the present disclosure and are not necessarily drawn to scale. In the drawings:
[0016] Figure 1A The locations of the samples used to derive α and β are shown;
[0017] Figure 1B This demonstrates how to derive the chromaticity prediction mode from the luminance mode when cclm_ is enabled;
[0018] Figure 1C A unified binarization table for chromaticity prediction modes is shown;
[0019] Figure 2 The diagram shows the classification of luminance samples into two categories for deriving two sets of α and β, (top) the sample domain and (bottom) the spatial domain;
[0020] Figure 3 The location of the sample used to derive the CCCM filter is shown when six reference lines are used;
[0021] Figure 4 From left to right: 3 taps vertical, 3 taps horizontal, 5 taps intersecting, 25 taps diamond;
[0022] Figure 5 An example of four reference lines adjacent to the prediction block is shown;
[0023] Figure 6 The matrix-weighted intra-frame prediction process is illustrated.
[0024] Figure 7 The HoG calculation is shown from a template with a width of 3 pixels;
[0025] Figure 8 The low-frequency non-separable transform (LFNST) process is illustrated;
[0026] Figure 9A The transformation selection table is shown;
[0027] Figure 9BThe intra-frame template matching search area used is shown;
[0028] Figure 11 The reference region of IBC is shown when CTU(m, n) is encoded;
[0029] Figure 12A and Figure 12B The following are examples for (a) Figure 12A The horizontal flip shown and (b) as shown Figure 12B The diagram shows the vertical flipping BV adjustment;
[0030] Figure 13 It consists of chromaticity PU and iso-positional luminance PU;
[0031] Figure 14A This shows two block vectors from co-occurring coding units C and TL, pointing to different reference sample regions. Overlapping regions (marked with a pattern) will only be considered once; and
[0032] Figure 14B Two block vectors from the co-occurrence coding unit C and the co-occurrence coding unit TL are shown, pointing to the maximum and minimum coordinates of the composite reference region defined by the co-occurrence block vectors, and the co-occurrence PU region is used to determine the reference region.
[0033] Figure 15 Two block vectors from the same prediction unit C and the same prediction unit TL, pointing to different reference sample regions, are shown. Based on the spatial location of the PU to which the block vectors (bv 0 and bv 1) belong, the chromaticity PU can be divided into two cross-component models (0 and 1);
[0034] Figure 16 A block diagram of a possible, and non-limiting, exemplary system in which example embodiments can be practiced is shown; and
[0035] Figure 17 A method according to an example embodiment of the present invention is shown, which can be made by, for example Figure 16 The apparatus shown is used for execution. Detailed Implementation
[0036] In exemplary embodiments of the present invention, at least one method and apparatus for improving block vector-guided cross-component prediction are proposed. These improvements involve reference sample selection, processing of multiple block vectors in co-location regions, processing of overlapping reference samples, and robust reference sample selection based on co-location brightness sample values.
[0037] Similarly, as described above, hybrid video codecs (e.g., ITU-T H.263, H.264 / AVC, and HEVC) can encode video information in two stages. First, pixel values in a picture region (or “block”) are predicted, for example, through motion compensation (finding and indicating a region closely corresponding to the block being encoded in one of the previously encoded video frames) or through spatial means (using pixel values around the block to be encoded in a specified manner). In this first stage, predictive coding can be applied, for example, as so-called sample prediction and / or so-called syntax prediction.
[0038] In sample prediction, pixel or sample values are predicted for a given region or "block" of an image. For example, these pixel or sample values can be predicted using one or more of the following mechanisms: motion compensation or intra-frame prediction.
[0039] Motion compensation mechanisms (also known as inter-frame prediction, temporal prediction, motion-compensated temporal prediction, or motion-compensated prediction, or MCP) involve finding and indicating regions in a previously encoded video frame that closely correspond to the block being encoded. Inter-frame prediction can reduce temporal redundancy.
[0040] Intra-frame prediction (where pixel or sample values can be predicted using spatial mechanisms) involves finding and indicating spatial region relationships. Intra-frame prediction leverages the fact that neighboring pixels within the same image can be correlated. Intra-frame prediction can be performed in the spatial domain or the transform domain; that is, sample values or transform coefficients can be predicted. Intra-frame prediction is typically used in intra-frame coding, where inter-frame prediction is not applied.
[0041] In syntax prediction (also known as parametric prediction), syntax elements and / or syntax element values and / or variables derived from syntax elements are predicted from earlier (de-encoded) syntax elements and / or earlier derived variables. A non-restrictive example of syntax prediction is provided below.
[0042] In motion vector prediction, motion vectors, such as those used for inter-frame and / or inter-frame view prediction, can be differentially encoded relative to block-specific predicted motion vectors. In many video codecs, predicted motion vectors are created in a predefined manner, for example, by calculating the median of encoded or decoded motion vectors from adjacent blocks. Another method for creating motion vector predictions (sometimes called Advanced Motion Vector Prediction (AMVP)) is to generate a list of candidate predictions based on adjacent and / or co-located blocks in a time-referenced image and signal the selected candidate as the motion vector predictor. In addition to predicting motion vector values, reference indices from previously encoded / decoded images can also be predicted. Reference indices are typically predicted based on adjacent and / or co-located blocks in a time-referenced image. Differential encoding of motion vectors is typically disabled across slice boundaries.
[0043] Block partitioning (e.g., from CTU to CU, and then to PU) can be predicted.
[0044] In filter parameter prediction, for example, filter parameters used for sample adaptive offset can be predicted.
[0045] Prediction schemes that use image information from previously encoded images can also be called inter-frame prediction methods, or temporal prediction and motion compensation.
[0046] Prediction schemes that use image information within the same image can also be called intra-frame prediction methods.
[0047] Secondly, the prediction error (i.e., the difference between the predicted pixel block and the original pixel block) is encoded. This can be achieved by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant thereof), quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the precision of the pixel representation (image quality) and the size of the resulting encoded video representation (file size at the transmission bitrate).
[0048] In many video codecs (including H.264 / AVC and HEVC), motion information is indicated by motion vectors associated with each motion-compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be encoded (in the encoder) or decoded (at the decoder) relative to the predicted source block in one of the previously encoded or decoded images (or pictures). Like many other video compression standards, H.264 / AVC and HEVC divide the picture into a rectangular grid, and for each rectangular grid, a similar block from a reference image is indicated for inter-frame prediction. The position of the predicted block is encoded as a motion vector indicating the position of the predicted block relative to the encoded block.
[0049] The following new encoding tools exist in the general-purpose video codec (VVC) currently under development. (Further details will be added in the final draft patent as needed.) • Intra-frame prediction: - 67-frame mode, using wide-angle mode extension - Block size and mode-dependent 4-tap interpolation filter - Location-dependent intra-frame prediction combination (PDPC) - Cross-component linear model intra-frame prediction (CCLM) - Multi-reference line intra-frame prediction - Intra-frame sub-partition - Weighted intra-frame prediction, using matrix multiplication; • Image prediction: - Block-based motion replication, using spatial, temporal, history-based, and pairwise average merge candidates. - Affine motion inter-frame prediction - Sub-block-based temporal motion vector prediction - Adaptive motion vector resolution - Motion compression based on 8×8 blocks for temporal motion prediction - High-precision (1 / 16 pixel) motion vector storage and motion compensation, with an 8-tap interpolation filter for the luminance component and a 4-tap interpolation filter for the chrominance component. - Triangular partition - Combined intra-frame and inter-frame prediction - Merging with MVD (MMVD) - Symmetric MVD encoding - Bidirectional optical flow - Decoder-side motion vector refinement - Bidirectional prediction using CU-level weights; • Transformation, quantization, and coefficient encoding: - Utilizing multiple principal transform selection of DCT2, DST7, and DCT8 - Auxiliary transform for low-frequency region - Sub-block transformation for inter-frame prediction residuals - Related quantification, where the maximum QP increased from 51 to 63 - Transform coefficient encoding using symbolic data hiding - Transform skips residual coding; • Entropy encoding: - Arithmetic coding engine with adaptive dual-window probability updates • In-loop filter: - Loop Shaping - Deblocking filter with robust long filter - Sample adaptive offset - Adaptive loop filter; • Screen content encoding: - Current image reference with reference area limitations; • 360-degree video encoding: - Horizontal orbital motion compensation; • Advanced syntax and parallel processing: - Reference image management, using direct reference image list signaling. - A tile group, containing rectangular tile groups.
[0050] Partitioning in VVC
[0051] In VVC, each image is divided into coding tree units (CTUs) similar to those in HEVC. Images can also be divided into slices, tiles, bricks, and sub-images. CTUs can be further subdivided into smaller CUs using a quadtree structure. Each CU can be partitioned using quadtrees and nested multi-type trees, including ternary splits and binary splits.
[0052] There are specific rules for inferring partitions within image boundaries.
[0053] Redundant splitting patterns are not allowed in nested multi-type partitions.
[0054] Cross-component linear model prediction (CCLM)
[0055] To reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode is used in VVC, where chromaticity samples are predicted based on reconstructed luminance samples from the same CU using the following linear model: in This represents the predicted chromaticity samples in the CU, and This represents downsampled reconstructed luminance samples from the same CU.
[0056] The CCLM parameters (α and β) are derived from at most four adjacent chroma samples and their corresponding downsampled luminance samples. Assuming the current chroma block size is W×H, then W' and H' are set as follows: - When applying the LM pattern, W'=W, H'=H; - When applying the LM-A mode, W' = W + H; - When applying the LM-L pattern, H' = H + W;
[0057] The adjacent positions mentioned above are denoted as S[0, -1]……S[W'-1, -1], and the left adjacent positions are denoted as S[-1, 0]……S[-1, H'-1]. Then, the four samples are selected as follows: - When LM mode is applied and both the upper neighbor sample and the left neighbor sample are available, S[W' / 4, -1], S[3] W' / 4, -1], S[-1, H' / 4], S[-1, 3 H' / 4]; - When applying LM-A mode or when only the upper adjacent sample is available, S[W' / 8, -1], S[3] W' / 8, -1]、S[5 W' / 8, -1]、S[7 W' / 8, -1]; - When applying the LM-L mode or when only the left adjacent sample is available, S[-1, H' / 8], S[-1, 3] H' / 8]、S[-1, 5 H' / 8]、S[-1, 7 H' / 8]; - Four adjacent luminance samples at the selected location are downsampled and compared four times to find two smaller values: x0A and x1A, and two larger values: x0B and x1B. Their corresponding chromaticity sample values are represented as y0A, y1A, y0B, and y1B. Then, xA, xB, yA, and yB are derived as follows: - ; ; ; Finally, the linear model parameters α and β are obtained according to the following equation. • ○
[0058] Figure 1 shows an example of the positions of the left and top samples, as well as the sample in the current block, involved in CCLM mode. The division operation used to calculate the parameter α is implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) is used. - And the parameter α is represented using the exponential notation. For example, diff is approximated using 4 significant bits and the exponent. Therefore, the table for 1 / diff is reduced to 16 elements for 16 significant values, as follows: - This will help reduce both the complexity of the calculations and the memory size required to store the necessary tables.
[0059] In addition to the upper and left templates being used together to calculate linear model coefficients, they can also be used alternately for the other two LM modes, known as the LM_A mode and the LM_L mode.
[0060] In LM_A mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to (W+H). In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to (H+W).
[0061] For non-square blocks, the top template is expanded to W+W, and the left template is expanded to H+H.
[0062] To match the chroma sample positions for a 4:2:0 video sequence, two types of downsampling filters are applied to the luminance samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filters is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "Type 0" and "Type 2" content, respectively.
[0063] Note that when the upper reference line is at the CTU boundary, only one luminance line (the universal line buffer in intra-frame prediction) is used to create the downsampled luminance sample.
[0064] This parameter calculation is performed as part of the decoding process, not just as part of the encoder's search operation. Therefore, no syntax is used to convey the α and β values to the decoder.
[0065] For chroma intra-mode coding, a total of eight intra-modes are allowed. These modes include five traditional intra-modes and three cross-component linear model modes (CCLM, LM_A, and LM_L). The chroma mode signaling and derivation process are shown in Table 1. Chroma mode coding directly depends on the intra-prediction mode of the corresponding luma block. Due to the separate block partitioning structure for luma and chroma components enabled in the I-slice, one chroma block can correspond to multiple luma blocks. Therefore, for the chroma DM mode, the intra-prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited.
[0066] Figure 1C A unified binarization table for chromaticity prediction modes is shown. Figure 1CIn the binary table, the first binary bit indicates whether it is normal (0) or LM mode (1). If it is LM mode, the next binary bit indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next 1 binary bit indicates whether it is LM_L (0) or LM_A (1). In this case, when sps_cclm_enabled_flag is 0, the first binary bit of the binarization table used for the corresponding intra_chroma_pred_mode can be discarded before entropy encoding. Or, in other words, the first binary bit is inferred to be 0 and therefore not encoded. This single binarization table is used for both the case where sps_cclm_enabled_flag equals 0 and the case where sps_cclm_enabled_flag equals 1. The first two binary bits in the table are context-encoded using their own context model, and the remaining binary bits are bypassed.
[0067] Furthermore, to reduce luma and chroma latency in dual-tree systems, when partitioning 64×64 luma coding tree nodes using Not Split (and ISP not used for 64×64 CUs) or QT, chroma CUs in 32×32 / 32×16 chroma coding tree nodes are allowed to use CCLM in the following manner: - If the 32×32 chroma node is not split or partitioned by QT, then all chroma CUs in the 32×32 node can use CCLM; - If a 32×32 chroma node is partitioned using horizontal BT, and the 32×16 child nodes are not split or are split using vertical BT, then all chroma CUs in the 32×16 chroma node can use CCLM.
[0068] CCLM is not allowed for chroma CU under all other luma and chroma coding tree splitting conditions.
[0069] Multi-model LM (MMLM)
[0070] The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, reconstructed neighboring samples are classified into two classes using a threshold that is the average of the brightness of the reconstructed neighboring samples. The linear model for each class is derived using the least mean square (LMS) method. The LMS method is also used to derive the linear model for the CCLM mode. Figure 2 The illustration shows two luminance-to-chrominance models obtained for a luminance (Y) threshold of 17. Figure 2The diagram illustrates classifying luminance samples into two categories for deriving two sets of α and β: (top) the sample domain and (bottom) the spatial domain. Each luminance-to-chrominance model has its own linear model parameters α and β. As can be seen from the figure below, each luminance-to-chrominance model corresponds to a spatial segmentation of the content (i.e., they correspond to different objects or textures in the scene).
[0071] Convolutional Cross-Component Model (CCCM)
[0072] Figure 3 The location of the sample used to derive the CCCM filter is shown when using six reference lines.
[0073] An improved version of cross-component prediction (called CCCM) uses a 2D filter kernel to derive the luma-to-chroma model. The filter coefficients are derived on the decoder side using the reconstructed input data and a set of chroma samples. For the filter coefficient derivation, as follows... Figure 3 As shown, a co-located reference sample region (composed of reconstructed luminance and chrominance samples) is defined for both luminance and chrominance, employing the commonly used 4:2:0 chrominance downsampling. For example, as... Figure 3 As shown, the reference sample region for a given block can be the six lines above and to the left, but any number of reference lines can be used (which can be implemented by the encoder and decoder). In general, the reference samples can contain any chroma and luminance samples reconstructed by both the encoder and decoder. After the reference samples are determined, the filter coefficients can be derived, for example, using different types of linear regression tools such as ordinary least squares estimation, orthogonal matching pursuit, optimized orthogonal matching pursuit, ridge regression, or minimum absolute shrinkage and selection operators.
[0074] The size of the filter kernel can be, for example, 1. x 3 (1D vertical), 3 x 1 (Level 1D), 3 x 3, 7 x 7 or any size, and can be shaped into a cross or rhombus (e.g., by selecting only a subset of all possible core locations) Figure 4 (as shown) or any given shape. Figure 4 From left to right: 3-tap vertical, 3-tap horizontal, 5-tap cross, 25-tap diamond. When referencing samples within the filter kernel, the following symbols are used: North (top), East (right), South (bottom), West (left), and Center, as shown below. Figure 4 The letters N, E, S, W, and C are used to represent this.
[0075] The overall approach to reconstructing chroma samples using convolutions between the filter kernels obtained from the decoder side and the input dataset is referred to here as the Convolutional Cross-Component Model (CCCM). The following steps can be applied to perform CCCM operations: 1. Define co-location reference regions for the luminance and chrominance components; 2. Downsample the luminance samples to match the chromaticity grid (optional); 3. Scan the luminance and chrominance samples of the reference region and collect available statistics (such as autocorrelation matrix and cross-correlation vector) based on the filter shape. 4. Solve for the filter coefficients by minimizing the squared error (or any other metric) based on available statistics (such as the autocorrelation matrix and cross-correlation vector); 5. The predicted chromaticity patch is calculated by convolving the downsampled luminance samples with the filter kernel.
[0076] Let's define (potentially downsampled) luminance samples as those using level x Coordinates and vertical y 2D array of coordinate indices Let's also define isotopic chromaticity samples as 2D arrays. And the filter kernel (i.e., coefficients) is defined as 3. x 3 arrays At the sample level, we will Y and F The convolution between them is defined as follows: . When using other data items (such as non-linear square root terms), the additional convolution becomes... , in These are filter coefficients located outside the 2D filter kernel, but have already been obtained as part of the system of linear equations used to solve for the 2D filter coefficients in step 4 above. Similarly, we can add a bias term to the convolution. .
[0077] Multi-reference line (MRL) intra-frame prediction
[0078] Multi-reference line (MRL) intra-prediction uses more reference lines for intra-prediction. Figure 5 An example of four reference lines adjacent to the prediction block is shown. Figure 5 The example depicts four reference lines, where the samples for segments A and F are not extracted from reconstructed neighboring samples, but are instead filled using the nearest samples from segments B and E, respectively. Intra-image prediction in HEVC uses the nearest reference line (i.e., reference line 0). In MRL, two additional lines (reference line 1 and reference line 3) are used.
[0079] The index of the selected reference line (mrl_idx) is signaled and used to generate the intra-predictor. For reference line idx greater than 0, only the additional reference line mode is included in the MPM list, and only the mpm index is signaled, not the remaining modes. The reference line index is signaled before the intra-predictor mode, and if a non-zero reference line index is signaled, the planar mode is excluded from the intra-predictor mode.
[0080] MRL is disabled for the first line block within the CTU to prevent the use of extended reference samples outside the current CTU line. Additionally, PDPC is disabled when using additional lines. For MRL mode, the derivation of the DC value in the intra-frame prediction mode for non-zero reference line indices is consistent with the derivation for reference line index 0. MRL requires the CTU to store three adjacent luma reference lines to generate predictions. The Cross-Component Linear Model (CCLM) tool also requires three adjacent luma reference lines for its own downsampling filter. The definition of MLR using the same three lines is consistent with CCLM to reduce the decoder's storage requirements.
[0081] Intra-frame sub-partition (ISP)
[0082] Intra-frame sub-partitioning (ISP) depends on the block size, dividing the luma intra-frame prediction block vertically or horizontally into 2 or 4 sub-partitions. For example, the minimum block size for ISP is 4×8 (or 8×4). If the block size is larger than 4×8 (or 8×4), the corresponding block is divided into 4 sub-partitions. It has already been noted that... M ×128 (of which M ≤ 64) and 128× N (in N ≤ 64) ISP blocks may pose potential problems for 64×64 VDPUs. For example, in a single-tree case, M ×128 CU has M ×128 TB brightness and two corresponding Chromaticity TB. If the CU uses an ISP, the luminance TB will be divided into four. M ×32 TB (horizontal partitioning is possible only), four M Each TB in a 32×32 TB is smaller than a 64×64 block. However, in the current ISP design, chroma blocks are not partitioned. Therefore, both chroma components have a size larger than a 32×32 block. Similarly, utilizing a 128× ISP... N A similar situation can also be created with the CU. Therefore, both situations pose a problem for a 64×64 decoder pipeline. For this reason, the CU size can be limited to a maximum of 64×64 using the ISP. All sub-partitions satisfy the condition of having at least 16 samples.
[0083] Matrix-weighted intra-frame prediction (MIP)
[0084] Matrix-weighted intra-prediction (MIP) is a new intra-prediction technique added to VVC. To predict a width of... W And the height is H For a rectangular block of samples, matrix-weighted intra-frame prediction (MIP) reconstructs H adjacent boundary samples from the left row of the block and the row above the block. W The reconstructed neighbor boundary samples are used as input. If the reconstructed sample is unavailable, it is generated in the manner of traditional intra-frame prediction. Figure 6 The matrix-weighted intra-frame prediction process is illustrated. For example... Figure 6 As shown, the generation of the predicted signal is based on the following three steps: averaging, matrix-vector multiplication, and linear interpolation.
[0085] Decoder-side Intra-Frame Mode Export (DIMD)
[0086] When applying DIMD, two intra-frame modes are derived from reconstructed neighboring samples, and these two predictors are combined with a planar mode predictor, where weights are derived from gradients, as described in JVET-O0449. The division operation in weight derivation is performed using the same lookup table (LUT)-based integerization scheme used by CCLM. For example, the partitioning operation in orientation computation... The following LUT-based scheme is used for calculation: in .
[0087] The exported intra-frame modes are included in the master list of most probable intra-frame modes (MPMs), so the DIMD process is performed before the MPM list is constructed. The master exported intra-frame modes of a DIMD block are stored with the block and used for the construction of the MPM list for adjacent blocks.
[0088] Figure 7 The HoG calculation is shown from a template with a width of 3 pixels.
[0089] Template-based intra-mode export (TIMD) fusion
[0090] For each intra-prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The top two intra-prediction modes with the minimum SATD are selected as TIMD modes. These two TIMD modes are weighted and fused after applying the PDPC procedure, and this weighted intra-prediction is used to encode the current CU. Location-dependent intra-prediction combination (PDPC) is included in the derivation of the TIMD modes.
[0091] The costs of the two selected modes are compared with the threshold. In the test where the cost coefficient is 2, the following application is made: .
[0092] If the condition is true, then apply fusion; otherwise, use only mode 1.
[0093] The weights of the patterns are calculated based on their SATD costs as follows: weight1 = costMode2 / (costMode1+costMode2); weight2 = 1 - weight1.
[0094] The division operation is performed using the same lookup table (LUT)-based integerization scheme used by CCLM.
[0095] Low-frequency non-separable transform (LFNST)
[0096] In VVC, such as Figure 8 As shown, LFNST is applied between the forward master transform and quantization (at the encoder) and between dequantization and inverse master transform (at the decoder). Figure 8 The Low Frequency Inseparable Transform (LFNST) process is illustrated. In LFNST, either a 4×4 or 8×8 inseparable transform is applied depending on the block size. For example, a 4×4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and an 8×8 LFNST is applied to larger blocks (i.e., min(width, height) > 4).
[0097] The following example illustrates the application of the inseparable transform used in LFNST. To apply 4×4 LFNST, the input block X is 4×4: First, it is represented as a vector. :
[0098] The inseparable transformation is calculated as ,in Indicates the transform coefficient vector, and T is a 16×16 transform matrix. A 16×1 coefficient vector Subsequently, it is reorganized into 4×4 blocks using the scan order (horizontal, vertical, or diagonal) of the block. Coefficients with smaller indices are placed in the 4×4 coefficient block together with smaller scan indices.
[0099] Reduced non-separable transform
[0100] LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on a direct matrix multiplication scheme, thus achieving it in a single pass without multiple iterations. However, it is necessary to reduce the dimension of the non-separable transform matrix to minimize the computational complexity and the storage space for storing the transform coefficients. Therefore, the reduced non-separable transform (or RST) method is used in LFNST. The main idea of the reduced non-separable transform is to map an N (for an 8×8 NSST, N is usually equal to 64) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, the RST matrix becomes an R×N matrix instead of an N×N matrix, as follows: The R rows of the transform form R bases in N-dimensional space. The inverse transform matrix for RT is the transpose of its forward transform. For 8×8 LFNST, a reduction factor of 4 is applied, and the 64×64 direct matrix (which is the size of a traditional 8×8 inseparable transform matrix) is reduced to a 16×48 direct matrix. Therefore, a 48×16 inverse RST matrix is used on the decoder side to generate the core (master) transform coefficients in the top-left region of the 8×8. When applying a 16×48 matrix instead of a 16×64 matrix with the same transform set configuration, each of the 16×48 matrices takes 48 input data from three 4×4 blocks (excluding the bottom-right 4×4 block) in the top-left 8×8 block. With the help of dimensionality reduction, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, with a reasonable performance degradation. To reduce complexity, LFNST is only applied if all coefficients outside the first coefficient subgroup are invalid. Therefore, when applying LFNST, all principal-only transform coefficients must be zero. This allows for adjustment of the LFNST index signaling at the last valid position, thus avoiding the extra coefficient scans in the current LFNST design, which are only used to check valid coefficients at specific locations. The worst-case handling of LFNST (in terms of multiplication per pixel) limits the non-separable transforms for 4×4 and 8×8 blocks to 8×16 and 8×48 transforms, respectively. In these cases, for other sizes smaller than 16, the last valid scan position must be less than 8 when applying LFNST. For blocks of shape 4×N and N×4 where N>8, the proposed constraints mean that LFNST is now applied only once, and only to the top-left 4×4 region. Since all principal-only coefficients are zero when applying LFNST, the number of operations required for the principal transform is reduced in this case. From the encoder's perspective, coefficient quantization is significantly simplified when testing LFNST transforms. For the first 16 coefficients (in scan order), rate-distortion optimized quantization must be performed to the maximum extent possible, while the remaining coefficients are forced to zero.
[0101] LFNST Transform Selection
[0102] LFNST has four transformation sets, and two inseparable transformation matrices (kernels) are used for each transformation set. For example... Figure 9A As shown, the mapping from intra-prediction mode to transform set is predefined. Figure 9AThe transform selection table is shown. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable secondary transform candidate is further specified by an LFNST index that is explicitly signaled. Following the transform coefficients, the index is signaled once per intra-frame CU in the bitstream.
[0103] Figure 9A The transformation selection table is shown.
[0104] LFNST index signaling and its interaction with other tools
[0105] Since LFNST only applies when all coefficients outside the first coefficient subgroup are invalid, LFNST index encoding depends on the position of the last significant coefficient. Furthermore, LFNST indexing is context-coded but does not depend on the intra-prediction mode, and only the first binary bit is context-coded. Additionally, LFNST applies to intra-CUs in both intra-slices and inter-slices, and to both luma and chroma. If dual-tree is enabled, the LFNST indexes for luma and chroma are signaled separately. For inter-slices (where dual-tree is disabled), a single LFNST index is signaled and used for both luma and chroma.
[0106] Considering that large CUs larger than 64×64 are implicitly split (TU blocks) due to the existing maximum transform size limit (64×64), LFNST index search can quadruple the data buffer size within a certain number of decoding pipeline stages. Therefore, the maximum allowed size of LFNST is limited to 64×64. Note that LFNST is only enabled under DCT2. LFNST index signaling is placed before MTS index signaling.
[0107] Using a scaling matrix for perceptual quantization is not very effective; a scaling matrix specified for the master matrix can be useful for LFNST coefficients. Therefore, scaling matrices are not allowed for LFNST coefficients. For single-tree partitioning mode, chroma LFNST is not applied.
[0108] Enhanced Multiple Transform Selection (MTS) for Intra-Frame Coding
[0109] In the current VVC design, only the DST7 transform core and DCT8 transform core are used for MTS, which are used for intra-frame coding and inter-frame coding.
[0110] Additional master transforms (including DCT5, DST4, and DST1) and identity transforms (IDTs) are employed. Furthermore, the MTS set depends on the TU size and intra-frame mode information. Sixteen different TU sizes are considered, and five different classes are considered for each TU size, depending on the intra-frame mode information. For each class, one, four, or six different transform pairs are considered. The number of intra-frame MTS candidates (between one, four, and six) is adaptively selected based on the sum of the absolute values of the transform coefficients. The sum is compared to two fixed thresholds to determine the total number of allowed MTS candidates. One candidate: sum <= th0 4 candidates: th0 <sum<= th1 6 candidates: sum > th1
[0111] Note that although there are a total of 80 different categories, some of these categories often share the exact same set of transformations. Therefore, the resulting LUT contains 58 unique entries (fewer than 80).
[0112] For angular modes, the joint symmetry of TU shape and intra-frame prediction is considered. Therefore, mode i (i>34) with TU shape A×B will be mapped to the same category as mode j=(68-i) with TU shape B×A. However, for each transform pair, the order of the horizontal and vertical transform kernels is interchanged. For example, a 16×4 block with mode 18 (horizontal prediction) and a 4×16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical and horizontal transform kernels are interchanged. For wide-angle modes, the most recent conventional angular mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80.
[0113] Inter-frame multiple transform selection (MTS) optimization
[0114] For the MTS of inter-frame coded CUs, four candidates are used for each CU: {(DST7, DST7), (DST7, DCT8), (DCT8, DST7), (DCT8, DCT8)}. For larger resolution sequences (width > 1080), the maximum CU size used for the inter-frame MTS is set to 32 (i.e., the inter-frame MTS uses CUs with width <= 32 and height <= 32), while for the remaining sequences (smaller resolution), the maximum CU size is set to 16. For 4-pt, 8-pt, and 16-pt transforms, the current AMT transform kernels (i.e., DST-7 and DCT-8) are replaced with the separable KLT proposed in JVET-J0021.
[0115] Related intra-block copying and template matching-based intra-block copying methods:
[0116] Intra-frame template matching
[0117] Intra-template matching prediction (intra-TMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side.
[0118] The prediction signal is generated by matching the L-shaped causal neighbors of the current block with another block in the predefined search area below, such as... Figure 9B As shown, the search area includes: R1: Current CTU; R2: Top left CTU; R3: Above CTU; R4: Left CTU.
[0119] The sum of absolute differences (SAD) is used as the cost function.
[0120] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.
[0121] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block size (BlkW, BlkH) so that each pixel has a fixed number of SAD comparisons.
[0122] Right now: · ; · . in' a ' is a constant used to control the gain / complexity tradeoff. In practice, ' a 'Equals 5.'
[0123] For CUs with a width and height less than or equal to 64, enable the intra-frame template matching tool. The maximum CU size for intra-frame template matching is configurable.
[0124] When DIMD is not used in the current CU, the intra-template matching prediction mode signals this at the CU level via a dedicated flag.
[0125] Intra-block copying (IBC) using template matching
[0126] Template matching is used in both the IBC merge mode and the IBC AMVP mode in IBC.
[0127] Compared to the list used in the regular IBC merging mode, the IBC-TM merging list is modified so that candidates are selected based on a pruning method, where the motion distance between candidates is the same as in the regular TM merging mode. The zero motion fulfillment is replaced by motion vectors to the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU.
[0128] In IBC-TM merging mode, the selected candidates are refined using template matching before the RDO or decoding process. IBC-TM merging mode is now competing with the regular IBC merging mode, and a TM merging flag is signaled.
[0129] In IBC-TM AMVP mode, up to three candidates are selected from the IBC-TM merge list. Each of these three candidates is refined using a template matching method and ranked according to the template matching cost it produces. Then, during motion estimation, typically only the top two candidates are considered.
[0130] Template matching refinement for both IBC-TM merging and AMVP modes is quite straightforward because the IBC motion vectors are constrained to be (i) integers and (ii) within the reference region. Therefore, in IBC-TM merging mode, all refinements are performed with integer precision, while in IBC-TM AMVP mode, they are performed with either integer or 4-pel precision depending on the AMVR value. This refinement only accesses the samples and does not perform interpolation. In both cases, the refined motion vectors and the template used in each refinement step must respect the constraints of the reference region.
[0131] Figure 10 The IBC reference area is shown, depending on the current CU location.
[0132] IBC Reference Area
[0133] Figure 11 This shows the reference region for IBC when CTU(m, n) is encoded. The double ++ symbol at the top... Figure 11 The block represents the current CTU; it is marked with an asterisk at the top. The block represents the reference region; Figure 11 The remaining blocks are invalid reference region blocks.
[0134] The reference area for IBC is extended to the two CTU rows above. Figure 11 The diagram illustrates the reference region used for encoding CTU(m, n). Specifically, for the CTU(m, n) to be encoded, the reference region includes CTUs with indices (m-2, n-2)……(W, n-2), (0, n-1)……(W, n-1), (0, n)……(m, n), where W represents the maximum horizontal index within the current tile, slice, or image. When the CTU size is 256, the reference region is limited to the top row of CTUs. This setting ensures that IBC does not require additional memory in the current ETM platform when the CTU size is 128 or 256. The vector search (or local search) range per sample block is limited to [-(C<<1), C>>2] horizontally and [-C, C>>2] vertically to accommodate the reference region expansion, where C represents the CTU size.
[0135] Reconstruct and reorder IBC (RR-IBC)
[0136] For IBC coded blocks, Reconstructed Reordered IBC (RR-IBC) mode is allowed. When RR-IBC is applied, samples in the reconstructed block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual computation, while the prediction block is derived without flipping. On the decoder side, the reconstructed block is flipped back to recover the original block.
[0137] For RR-IBC coded blocks, two flipping methods are supported: horizontal flipping and vertical flipping. First, a syntax flag is signaled for the IBC AMVP coded block to indicate whether the reconstruction is flipped. If the reconstruction is flipped, another flag is further signaled to specify the flipping type. For IBC merging, the flipping type is inherited from adjacent blocks without syntax signaling. Considering horizontal or vertical symmetry, the current block and the reference block are typically horizontally or vertically aligned. Therefore, when a horizontal flip is applied, the vertical component of the BV is not signaled and is inferred to be equal to 0. Similarly, when a vertical flip is applied, the horizontal component of the BV is not signaled and is inferred to be equal to 0.
[0138] Figure 12A and Figure 12B It shows (a) as Figure 12A The horizontal flip shown and (b) as shown Figure 12B The diagram shows the vertical flipping BV adjustment.
[0139] To better utilize symmetry, a flip-aware BV adjustment scheme is applied to refine the block vector candidates. For example, as... Figure 12A and Figure 12BAs shown, (xnbr, ynbr) and (xcur, ycur) represent the coordinates of the center samples of the neighboring block and the current block, respectively, and BVnbr and BVcur represent the BV of the neighboring block and the current block, respectively. When neighboring blocks are encoded using horizontal flipping, the BV is not directly inherited from the neighboring blocks. Instead, the horizontal component of BVcur is calculated by adding a motion offset to the horizontal component of BVnbr (denoted as BVnbrh), i.e., BVcurh = 2(xnbr-xcur) + BVnbrh. Similarly, when neighboring blocks are encoded using vertical flipping, the vertical component of BVcur is calculated by adding a motion offset to the vertical component of BVnbr (denoted as BVnbrv), i.e., BVcurv = 2(ynbr-ycur) + BVnbrv.
[0140] IBC Merging Mode Utilizing Block Vector Difference (IBC-MBVD)
[0141] As an extension of the standard MMVD mode, affine MMVD and GPM-MMVD are adopted for ECM. It is natural to extend the MMVD mode to the IBC merging mode.
[0142] In IBC-MBVD, the distance set is {1-pel, 2-pel, 4-pel, 8-pel, 12-pel, 16-pel, 24-pel, 32-pel, 40-pel, 48-pel, 56-pel, 64-pel, 72-pel, 80-pel, 88-pel, 96-pel, 104-pel, 112-pel, 120-pel, 128-pel}, and the BVD directions are two horizontal directions and two vertical directions.
[0143] The basic candidate is selected from the top five candidates in the reordered IBC merge list. Furthermore, all possible MBVD refinement positions (20×4) for each basic candidate are reordered based on the SAD cost between the template (the top row and left column of the current block) and its reference at each refinement position. Finally, the top 8 refinement positions with the lowest template SAD cost are reserved as available positions for MBVD index encoding. The MBVD index is binary-coded using rice code with a parameter equal to 1.
[0144] IBC-MBVD encoded blocks do not inherit the flip type of RR-IBC encoded neighbor blocks.
[0145] Similarly, as described above, when using intra-block replication, block vector-guided cross-component prediction uses non-local regions to improve the prediction performance of the cross-component model. If the intra-block replication (IBC) method is used to encode co-occurring blocks in the reference channel, the relevant block vector (BV) is used to indicate the reference region used to compute the parameters of the cross-component prediction model.
[0146] When a co-location region contains several blocks with associated block vectors, reference sample selection can be performed in several ways. The reference sample should also reflect the intensity distribution of samples within the co-location block. Inappropriately selected reference samples directly lead to impaired performance in cross-component prediction.
[0147] The exemplary embodiments of the present invention provide several improvements for cross-component prediction guided by block vectors. These improvements involve reference sample selection, handling of multiple block vectors in co-location regions, handling of overlapping reference samples, robust reference sample selection based on co-location brightness sample values, etc.
[0148] Before describing the exemplary embodiments disclosed herein in detail, please refer to Figure 16 The illustration shows simplified block diagrams of various electronic devices applicable to practicing exemplary embodiments of the present invention.
[0149] Figure 16 A block diagram of one possible, non-limiting, exemplary system in which example embodiments can be practiced is shown. Figure 16 In this context, User Equipment (UE) 10 communicates wirelessly with Wireless Network 1 or Network 1, such as... Figure 16 As shown. Figure 16 The wireless network 1 or network 1 shown may include a communication network, such as a mobile network, for example, mobile network 1 or the first mobile network disclosed herein. Figure 16 Any reference to Wireless Network 1 in this document can be considered a reference to any wireless network disclosed herein. Furthermore, as... Figure 16 The wireless network 1 shown may also include hardwired features that a communication network might require. A UE is a wireless device that can access the wireless network, typically a mobile device. For example, a UE may be a mobile phone (or "cellular") and / or a computer with mobile terminal capabilities. For example, a UE or mobile terminal may also be a portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted mobile device that performs voice signaling and / or data exchange with the RAN.
[0150] UE 10 includes one or more processors DP 10A, one or more memories MEM 10B, and one or more transceivers TRANS 10D interconnected via one or more buses. Each transceiver in the one or more transceivers TRANS 10D includes a receiver and a transmitter. The one or more buses may be address buses, data buses, or control buses, and may include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optic cables, or other optical communication devices. The one or more transceivers TRANS 10D may optionally be connected to one or more antennas to communicate with NN 12 and NN 13 respectively. The one or more memories MEM 10B include computer program code PROG 10C. UE 10 communicates with NN 12 and / or NN 13 via wireless link 11 or wireless link 16.
[0151] NN 12 (NR / 5G Node B, evolved NB, or LTE equipment) is related to, for example, NR / 5G Node B, evolved NB, or LTE equipment. Figure 16 Network nodes, such as primary or secondary base stations (e.g., for NR or LTE Long Term Evolution), communicate with devices like NN 13 and UE 10. NN 12 provides access to wireless devices (such as UE 10) of the wireless network 1. NN 12 includes one or more processors DP 12A, one or more memories MEM 12B, and one or more transceivers TRANS 12D interconnected via one or more buses. According to an example embodiment, these TRANS 12Ds may include X2 and / or Xn interfaces for performing the example embodiment. Each transceiver in the one or more transceivers TRANS 12D includes a receiver and a transmitter. The one or more transceivers TRANS 12D may optionally be connected to one or more antennas to communicate with UE 10 via at least link 11. One or more memories MEM 12B and computer program code PROG 12C are configured, together with one or more processors DP 12A, to cause NN 12 to perform one or more operations described herein. NN 12 can communicate with another gNB or eNB, or devices such as NN 13, for example, via link 16. Furthermore, link 11, link 16, and / or any other link can be wired, wireless, or both, and can implement, for example, an X2 interface or an Xn interface. Additionally, link 11 and / or link 16 can communicate via other network devices, such as, but not limited to, Figure 16 The NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF 14 devices are shown. NN 12 can perform the functions of MME (Mobility Management Entity) or SGW (Serving Gateway), such as user plane functions and / or access management functions for LTE and similar functions for 5G.
[0152] NN 13 can be used for WiFi or Bluetooth or other wireless devices associated with mobile function devices such as AMF or SMF. Furthermore, NN 13 may include an NR / 5G node B or a possible evolved NB, a base station communicating with devices such as NN 12 and / or UE 10 and / or Wireless Network 1, such as a primary or secondary node base station (e.g., for NR or LTE Long Term Evolution). NN 13 includes one or more processors DP 13A, one or more memories MEM 13B, one or more network interfaces, and one or more transceivers TRANS 13D interconnected via one or more buses. According to an example embodiment, these network interfaces of NN 13 may include an X2 interface and / or an Xn interface for performing the example embodiment. Each transceiver in the one or more transceivers TRANS 13D includes a receiver and a transmitter, which may optionally be connected to one or more antennas. The one or more memories MEM 13B include computer program code PROG 13C. For example, one or more memory MEM 13B and computer program code PROG 13C are configured, together with one or more processors DP 13A, to cause NN 13 to perform one or more operations described herein. NN 13 can communicate with another mobility function device and / or eNB (such as NN 12 and UE 10 or any other device) using, for example, link 11 or link 16 or another link. Figure 16 Link 16 shown can be used to communicate with NN 12. These links can be wired, wireless, or both, and can implement, for example, an X2 interface or an Xn interface. Furthermore, as mentioned above, links 11 and / or 16 can be connected via other network devices, such as, but not limited to, NCE / MME / SGW devices, such as... Figure 16 NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF 14.
[0153] Figure 16 One or more buses of the device can be address, data, or control buses, and can include any interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optic or other optical communication devices, wireless channels, etc. For example, one or more transceivers TRANS 12D, TRANS 13D, and / or TRANS 10D can be implemented as remote radio heads (RRHs), where other components of NN 12 are physically located separately from the RRH, and these devices can include one or more buses, which can be partially implemented as fiber optic cables to connect other components of NN 12 to the RRH.
[0154] It should be noted that, although Figure 16Network nodes such as NN 12 and NN 13 are shown, but any of these nodes can be merged into or incorporated into an eNodeB, eNB, or gNB, such as for LTE and NR, and can still be configured to perform the example embodiments.
[0155] It should also be noted that while the description in this document indicates that a “cell” performs functions, it should be clear that a gNB forms a cell and / or is a user equipment and / or mobility management function device that will perform these functions. Furthermore, a cell constitutes part of a gNB, and there can be multiple cells for each gNB.
[0156] Wireless Network 1 or any network that it may represent may include or may not include NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF 14, which may include (NCE) network control element functions, MME (Mobility Management Entity) / SGW (Serving Gateway) functions, and / or Serving Gateway (SGW), and / or MME (Mobility Management Entity) and / or SGW (Serving Gateway) functions, and / or User Data Management Function (UDM), and / or PCF (Policy Control) functions, and / or Access and Mobility Management Function (AMF) functions, and / or Session Management (SMF) functions, and / or Location Management Function (LMF), and / or Authentication Server (AUSF) functions, and provide connectivity to another network (such as a telephone network and / or a data communication network (e.g., the Internet)), and be configured to perform any 5G and / or NR operations as a supplement to or replacement of other standard operations at the time of this application. NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF 14 can be configured to perform operations according to the exemplary embodiments in any of LTE, NR, 5G and / or any standards-based communication technology performed or discussed in this application. Furthermore, it should be noted that operations performed by NN 12 and / or NN 13 according to the exemplary embodiments can also be performed at NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF 14.
[0157] NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF 14 includes one or more processor DP 14A, one or more memory MEM 14B, and one or more network interfaces (N / WI / F) interconnected via one or more buses coupled to link 13 and / or link 16. According to an example embodiment, these network interfaces may include X2 interfaces and / or Xn interfaces for performing the example embodiment. One or more memory MEM 14B includes computer program code PROG 14C. The one or more memory MEM 14B and computer program code PROG 14C are configured, together with one or more processor DP 14A, to cause NCE / MME / SGW / UDM / PCF / AMF / SMF / LMF 14 to perform one or more operations, which may be necessary to support the operations according to the example embodiment.
[0158] It should be noted that NN 12 and / or NN 13 and / or UE 10 can be configured (e.g., based on standard implementations, etc.) to perform Location Management Function (LMF) functionality. LMF functionality can be embodied in any of these network devices or other devices associated with them. Furthermore, such as Figure 16 The LMF of MME / SGW / UDM / PCF / AMF / SMF / LMF 14, etc., can be co-located with UE 10 at least as described below, so as to be compatible with... Figure 16 NN 12 and / or NN 13 are separated to perform operations according to the example embodiments disclosed herein.
[0159] Wireless Network 1 can implement network virtualization, which is a process of combining hardware and software network resources and network functions into a single software-based management entity (virtual network). Network virtualization involves platform virtualization, which is often combined with resource virtualization. Network virtualization is classified as either external, which combines many networks or parts of networks into virtual units, or internal, which provides network-like functionality to software containers on a single system. Note that the virtualized entities created by network virtualization are still implemented to some extent using hardware such as processors DP10, DP12A, DP13A and / or DP14A and memory MEM 10B, MEM 12B, MEM 13B and / or MEM 14B, and such virtualized entities also produce technical effects.
[0160] Computer-readable storage devices MEM 12B, MEM 13B, and MEM 14B can be of any type suitable for the local technical environment and can be implemented using any suitable data storage technology, such as semiconductor-based memory devices, flash memory, magnetic storage devices and systems, optical storage devices and systems, fixed memory, and removable memory. Computer-readable storage devices MEM 12B, MEM 13B, and MEM 14B can be components for performing storage functions. Processors DP10, DP12A, DP13A, and DP14A can be of any type suitable for the local technical environment and can include one or more of general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), and processors based on multi-core processor architectures, as non-limiting examples. Processors DP10, DP12A, DP13A, and DP14A can be components for performing functions, such as controlling UE 10, NN 12, NN 13, and other functions described herein.
[0161] In general, various embodiments of any of these devices may include, but are not limited to, cellular phones, such as smartphones, tablets, personal digital assistants (PDAs) with wireless communication capabilities, portable computers with wireless communication capabilities, image capture devices (such as digital cameras) with wireless communication capabilities, gaming devices with wireless communication capabilities, music storage and playback devices with wireless communication capabilities, internet devices that allow wireless internet access and browsing, tablets with wireless communication capabilities, and portable units or terminals that combine such functions.
[0162] Furthermore, various embodiments of any of these devices can be used in conjunction with: UE vehicles, high-altitude platform stations, or any other type of node associated with a ground network, or any type of drone radio, or radio in an aircraft or other airborne vehicle or a vessel (such as a ship) navigating on water.
[0163] Similarly, as described above, exemplary embodiments of the present invention provide several improvements to block vector-guided cross-component prediction. These improvements involve reference sample selection, handling of multiple block vectors in co-location regions, handling of overlapping reference samples, robust reference sample selection based on co-location brightness sample values, etc.
[0164] In an embodiment, the reference region for cross-component model derivation can be determined using all available block vectors in the same-position luminance coding unit or prediction unit. Figure 13 The image in the middle shows an example. Figure 13 It consists of chromaticity PU and iso-positional luminance PU.
[0165] In an embodiment, the reference region used for cross-component model derivation can be determined using only a subset of the available block vectors in the co-position brightness PU (e.g., Figure 13 (As shown). For example, only C, TL, TR, BL, and BR can be considered. The selection of a subset can be inferred based on local properties such as PU size and PU shape, and this selection can also be signaled to the decoder by the encoder.
[0166] In an embodiment, when multiple block vectors are available in the same brightness PU, a reference sample region can be derived so that redundant overlapping regions are discarded. Figure 14A This illustrates two block vectors from co-occurring coding units C and TL, pointing to different reference sample regions. Overlapping regions (marked with a pattern) will be considered only once. For example, as shown... Figure 14A As shown, when two or more block vectors point to the same reference sample, the overlapping region (marked by the pattern) will only be considered once. In addition to deriving a more accurate cross-component prediction model, removing redundant overlapping samples can also reduce the computational complexity of parameter calculations, especially on the decoder side.
[0167] In an embodiment, when multiple block vectors are available in the same PU, the average of the block vectors can be used as a block vector pointing to a reference sample.
[0168] In embodiments based on previous examples, the average can only be applied if the block vector difference is considered sufficiently small.
[0169] In embodiments based on previous examples, the average can be calculated as a weighted average based on the size of the co-located PUs, so that the block vectors of larger PUs are assigned larger weights. Alternatively, the average can be calculated based on a fixed grid of block vectors. For example, each 4×4 block in the co-located blocks with block vectors can be identified, and the average of these block vectors can be determined as the average block vector used to determine the position of the reference sample.
[0170] In an embodiment, when multiple block vectors are available in the same PU, a block pointing to a reference sample can be derived using linear or nonlinear estimation. For example, the median, minimum, or maximum value of the block vectors can be considered. The choice of operators (median, minimum, maximum, etc.) can be the same or different for the horizontal and vertical block vector components.
[0171] Figure 14B It shows two block vectors from the co-occurrence coding unit C and the co-occurrence coding unit TL, pointing to the maximum and minimum coordinates of the composite reference region defined by the co-occurrence block vectors, and the co-occurrence PU region is used to determine the reference region.
[0172] In such Figure 14BIn the illustrated embodiment, when multiple block vectors are available in the co-located PU, the maximum and minimum coordinates of the composite reference region defined by the co-located block vectors and the co-located PU region are used to determine the reference region. For example, the reference region can be determined with its left boundary at the minimum horizontal coordinate of the composite region and its right boundary at the maximum horizontal coordinate of the composite region. Similarly, the top and bottom boundaries of the reference region can be determined using the minimum and maximum vertical coordinates of the composite region. If the reference region generated in this way is larger or smaller than the PU itself, the size of the reference region can be scaled to match the size of the PU.
[0173] In an embodiment, when none of the co-located PUs have an associated block vector, temporary block vectors can be determined for those co-located PUs that do not have their own block vectors. For example, such temporary block vectors can be determined by copying the block vectors of the selected PUs in the reference channel, or by interpolating the temporary block vectors from the block vectors of the selected PUs in the reference channel. Then, similar to block vectors obtained directly from co-located PUs, the temporary block vectors can be used to determine the reference region.
[0174] In an embodiment, the block vectors in the co-position PU and their spatial locations within the co-position region can be used for interpolation or estimation of fine block vectors. For example, a planar model or a quadratic model can be used to derive additional block vectors at the center of the co-position brightness region. Such estimation can be performed, for example, using a linear regression solver. The derived block vectors are then used as pointers to reference samples derived from the cross-component model.
[0175] In an embodiment, when multiple block vectors are available in the co-located PU, a template-matching-based refinement mechanism can be used to find the best-matching reference block and its corresponding block vector. Template-matching-based refinement can use some or all samples from the co-located reference block and / or some or all neighboring reference samples from the current block.
[0176] According to the foregoing embodiments, template-matching-based refinement can use one or more available block vectors from the corresponding PU as initial block vectors in the refinement search. Alternatively or additionally, one or more initial block vectors can be derived using other embodiments described herein.
[0177] In an embodiment, when multiple block vectors are available in the same PU, the chromaticity PU can be divided into multiple cross-component models based on the spatial location of the same PU.
[0178] Figure 15 Two block vectors from the same prediction unit C and the same prediction unit TL, pointing to different reference sample regions, are shown. Based on the spatial location of the PU to which the block vectors (bv 0 and bv 1) belong, the chromaticity PU can be divided into two cross-component models (0 and 1).
[0179] For example, such as Figure 15 As shown, block vector 0 is used to derive the cross-component model of the upper half of the chroma PU (labeled 0), and conversely, block vector 1 is used to derive the cross-component model of the lower half of the chroma PU (labeled 1).
[0180] In an embodiment, when multiple block vectors are available in the same PU, multiple cross-component models can be computed using multiple BVs. A final prediction can then be obtained by combining the multiple predictions. The weights used to combine the multiple predictions can be defined in the codec specification, or the weights or their identifier indices can be signaled in the bitstream, or they can be computed on the decoder side based on block information such as sample reconstruction samples, block size, etc.
[0181] In an embodiment, the sample statistics of the co-location block can be used to prune training samples in the reference block to which the block(s) are pointed. For example, the minimum and maximum intensity values of the samples in the co-location block can be determined and used for pruning such that samples within the minimum and maximum intensity ranges are considered for parameter calculation or training a cross-component prediction model. The calculated minimum and maximum intensity values can be expanded by incremental values from the lower and upper bounds of the range to provide a larger intensity range for training. The incremental values can be fixed or determined by the calculated minimum and maximum values. For example, the incremental value can be a percentage of the minimum and / or maximum values. In another example, the incremental value used to expand the lower and upper bounds can be the difference between the minimum and maximum intensity values, or it can be a scaled version of the difference between the minimum and maximum intensity values of the co-location block samples.
[0182] In an embodiment, the cross-component model type can be determined based on the sample distribution from the co-occurring block and / or reference block pointed to by the block vector. For example, the decision to use a single model or multiple models for prediction can be based on the sample distribution with respect to the classification parameters of the cross-component model. Cross-component models typically use the mean of the training samples as classification parameters to compute multiple models. When the sample ratio distribution of the model based on the classification parameters is poor, a single-model variant can be determined. Alternatively, when the sample ratio distribution of the model based on the classification parameters is good, a multiple-model variant can be determined for cross-component prediction. The sample ratio can be predefined or signaled in the bitstream.
[0183] Based on the preceding embodiments, the model type can be determined by other parameters besides or instead of minimum and maximum values. For example, the mean of the samples can be used. Samples falling within a certain intensity distance of the mean and the lower and upper limits can be used to calculate parameters. The intensity distances used to determine the lower and upper limits can be the same or different. The intensity distance can be defined as a percentage of the mean. The intensity distance value or its indicator can be signaled in the bitstream.
[0184] In an embodiment, when multiple block vectors are available in the co-located PU, a reference block for model derivation is determined by searching for the best matching region with the co-located block in the reference channel. One or more available block vectors can be used during the search process. The best matching region during the search process can be calculated using cost metrics such as Sum of Absolute Differences (SAD), Sum of Squared Errors (SSE), Sum of Transform Differences (SATD), or any other metric.
[0185] In another embodiment, one or more correlation metrics can be used to select samples from the reference region pointed to by the block vector of the co-located PU. For example, as previously described, the correlation coefficient can be used as a distortion metric.
[0186] The following are the details of how to use correlation coefficients to select the samples most strongly correlated with the co-blocks in the reference channel. :
[0187] Let's consider two strings of length 1. n Given signals X and Y. We now list five sums based on X and Y. Furthermore, the Pearson correlation coefficient can be expressed as, .
[0188] The denominator in the above equation (the product of the standard deviations of X and Y) normalizes the covariances of X and Y to the range [-1, 1], thus making P usable for assessing the correlation between X and multiple Ys. For integer arithmetic, we can use... in bitshift Choose based on the required level of precision. compared to, The secondary behavior will maintain the order / rank of the correlation coefficient, but the size will decrease more quickly.
[0189] We define the relevant costs as follows: , A larger value of C indicates a better match. Many existing coding tools (such as VVC) perform motion compensation and template matching by minimizing SSE or SAD using integer arithmetic. By considering the following formula, we can easily insert the relevant costs into the process. , Furthermore, for situations where negative distortion is not allowed, we can use, ,or , It depends on the situation.
[0190] The described process can iterate over different regions pointed to by block vectors from co-position PUs and produce regions with the highest correlation for computing cross-component prediction models.
[0191] In addition to or instead of using different block vectors, a refinement process can be applied to one or more co-position block vectors from the co-position PU. The refinement process can add increments to the initial block vector in either or both directions of the BV and calculate the associated cost of the region by which the refined BV is applied. The optimal region and BV are then selected and used to calculate the parameters for cross-component prediction.
[0192] Figure 17 The illustration shows that it can be made by, but is not limited to, devices (e.g., such as...) Figure 16 Operations performed by devices such as UE 10. Figure 17 As shown in step 1710, the video decoder determines the reference sample region for the video segment samples derived from the cross-component model. For example... Figure 17 As shown in step 1720, the identified subset of available block vectors is determined using either the same-position luminance coding unit or the same-position luminance prediction unit. Figure 17 As shown in step 1730, the subset is identified based on local attributes. Figure 17 As shown in step 1740, a cross-component prediction model is obtained based on this determination. Then, as... Figure 17 As shown in step 1750, a cross-component prediction model is used to decode video segment samples.
[0193] According to the example embodiments described in the preceding paragraphs, the local attributes include at least one of prediction cell size or prediction cell shape.
[0194] According to the example embodiments described in the preceding paragraphs, at least one of the prediction unit size or prediction unit shape is predetermined by the video decoder or received from the video encoder.
[0195] According to the example embodiment described in the above paragraph, the reference sample region is an intra-block copy reference region.
[0196] According to the example embodiments described in the preceding paragraphs, the reference sample region is determined based on the availability of more than one block region in at least one prediction unit with the same brightness.
[0197] According to the example embodiment described in the preceding paragraph, the determination includes: deriving a reference sample region such that overlapping or redundant regions are discarded.
[0198] According to the example embodiment described in the above paragraph, the average of the available block vectors is used as the block vector pointing to the reference sample region.
[0199] According to the example embodiment described in the above paragraph, when the block vector size difference is less than a threshold, the block vector points to the reference sample region.
[0200] According to the example embodiments described in the preceding paragraphs, the average can be calculated as a weighted average based on the size of either the same-position luminance coding unit or the same-position luminance prediction unit.
[0201] According to the example embodiment described in the preceding paragraph, weights are assigned to the block vector of at least one prediction unit based on the size of the prediction unit in at least one prediction unit.
[0202] According to the example embodiment described in the above paragraph, the average is calculated based on a fixed grid of block vectors.
[0203] According to the example embodiments described in the preceding paragraphs, at least one of linear estimation or nonlinear estimation is used to derive a block vector pointing to a reference sample region.
[0204] According to the example embodiment described in the above paragraph, the maximum and minimum coordinates of a composite reference region defined by co-location vectors and co-location prediction unit regions, based on more than one block region, are used to determine the reference sample region.
[0205] According to the example embodiment described in the preceding paragraph, where the size of the reference sample region differs from the size of at least one prediction unit, the device is configured to scale the size of the reference sample region to match the size of at least one prediction unit.
[0206] According to the example embodiments described in the preceding paragraphs, a temporary block vector is identified for a co-position prediction unit that does not have a related block vector; and the temporary block vector is used to determine at least one of a reference sample region or a block vector obtained directly from the co-position prediction unit.
[0207] According to the example embodiments described in the preceding paragraphs, the spatial location of the block vector in the co-location region of one of the co-location luminance coding units or co-location luminance prediction units is used to interpolate or estimate the block vector pointing to the reference sample region for cross-component model derivation.
[0208] According to the example embodiment described in the preceding paragraph, where multiple block vectors are available in the co-position prediction unit region, a template matching-based refinement mechanism is used to find at least one of the following: a matching reference region block, or a corresponding block vector in an identified subset of available block vectors.
[0209] According to the example embodiments described in the preceding paragraphs, the template-matching-based refinement mechanism uses at least one sample from at least one of the co-located reference region blocks in the current block or at least one of the adjacent reference samples.
[0210] According to the example embodiment described in the above paragraph, where multiple block vectors are available in the co-position prediction unit region, the chroma prediction unit is divided into multiple cross-component models based on the spatial location of the co-position prediction unit.
[0211] According to the example embodiment described in the above paragraph, block vector 0 is used to derive the cross-component model of the upper half of the chromaticity prediction unit marked as 0, and block vector 1 is used to derive the cross-component model of the lower half of the chromaticity prediction unit marked as 1.
[0212] According to the example embodiment described in the above paragraph, multiple predictions are made using multiple block vectors for multiple cross-component models based on the availability of multiple block vectors in the same prediction unit region; and the final prediction is obtained by combining multiple predictions.
[0213] According to the example embodiments described in the preceding paragraphs, the weights used to combine multiple predictions can be one of the following: weights defined in the codec specification, or identifier indexes of weights signaled to the video decoder, or weights calculated on the decoder side based on block information including reconstructed samples and block sizes.
[0214] According to the example embodiment described in the above paragraph, the statistics of co-location blocks of more than one block region are used to prune training samples in the reference region block of at least one block vector pointing to the reference sample region.
[0215] According to the example embodiment described in the preceding paragraph, the minimum and maximum intensity values of samples in the co-location block are determined and used for pruning such that samples within the minimum and maximum intensity ranges are considered for at least one of the following: parameter calculation or training a cross-component prediction model.
[0216] According to the example embodiment described in the preceding paragraph, the minimum intensity value and the maximum intensity value are extended by incremental values from the lower and upper limits of the maximum intensity range to provide a larger range for training, and wherein the incremental value is one of the following: fixed, or determined by the minimum intensity value and the maximum intensity value.
[0217] According to the example embodiments described in the preceding paragraphs, the cross-component model type is determined based on the distribution of samples from the co-located block and / or reference block to which the block vector points.
[0218] According to the example embodiments described in the preceding paragraphs, where the sample ratios are distributed for a model based on classification parameters, a multi-model variant is determined for a cross-component prediction model, wherein the sample ratios are one of the following: predefined or signaled to the decoder.
[0219] According to the exemplary embodiments described in the preceding paragraphs, the model type is determined by other parameters in addition to or in place of minimum and maximum values, wherein the other parameters include the mean of the samples used, wherein the mean includes a lower limit and an upper limit based on the intensity distance used for parameter calculation, and wherein the intensity distance is one of the following: predefined or signaled to the decoder.
[0220] According to the example embodiment described in the above paragraph, where multiple block vectors are available in the co-position prediction unit region, the reference block for model derivation is determined by finding a matching region with the co-position block in the reference channel.
[0221] According to the example embodiment described in the preceding paragraph, one or more of the correlation metrics are used to select samples from the reference sample region pointed to by the block vector of the co-position prediction unit.
[0222] A stored procedure code (such as) Figure 16 The program code is executed by at least one processor (DP 10A and / or DP 10F in Figure 19) on a non-transitory computer-readable medium (MEM 12B in Figure 18) to perform at least the operations described in the above paragraphs.
[0223] According to the exemplary embodiments of the present invention described above, there exists an apparatus comprising: for determining (e.g., by a video decoder) Figure 16One or more transceivers 12D and / or 13D; MEM 12B and / or MEM 13B; PROG 12C and / or PROG 13C; and DP 12A and / or DP 13A) are used as components for determining a reference sample region for video clip samples derived across component models, wherein the determination uses an identified subset of available block vectors in one of the in-place luminance coding units or in-place luminance prediction units, and wherein the subset is identified based on local attributes (such as...). Figure 16 One or more transceivers 12D and / or 13D; MEM 12B and / or MEM 13B; PROG 12C and / or PROG 13C; and DP 12A and / or DP13A); used to obtain (e.g.) based on this determination. Figure 16 One or more transceivers 12D and / or 13D; MEM 12B and / or MEM 13B; PROG 12C and / or PROG 13C; and DP 12A and / or DP 13A) in the cross-component prediction model; and components for use (such as Figure 16 One or more transceivers (12D and / or 13D; MEM 12B and / or MEM 13B; PROG 12C and / or PROG 13C; and DP 12A and / or DP 13A) are used in the component prediction model to decode video clip samples.
[0224] In an example aspect of the invention according to the foregoing paragraphs, at least the components for identifying, identifying, and using include non-transitory computer-readable media [such as...]. Figure 5 [MEM 12B and / or MEM 13B in the medium], the medium is used by at least one processor [such as Figure 6 The executable computer program [PROG 12C and / or PROG 13C] of DP 12A and / or DP 13A is encoded.
[0225] Furthermore, according to exemplary embodiments of the present invention, there exists a circuit system for performing operations according to exemplary embodiments of the invention disclosed herein. This circuit system may include any type of circuit system, including content encoding circuit systems, content decoding circuit systems, processing circuit systems, image generation circuit systems, data analysis circuit systems, etc. Furthermore, the circuit system may include discrete circuit systems, application-specific integrated circuit systems (ASICs) and / or field-programmable gate array circuit systems (FPGAs), processors specifically configured by software to perform corresponding functions, or dual-core processors having software and corresponding digital signal processors, etc. In addition, the necessary inputs and outputs, the functions performed by the circuit system, and the interconnections (possibly via inputs and outputs) between the circuit system and other components that may include other circuit systems are provided to perform the exemplary embodiments of the invention described herein.
[0226] According to the exemplary embodiments of the invention disclosed in this application, the provided "circuit system" may include at least one or more, or all of the following: a hardware circuit implementation only (such as an implementation in an analog and / or digital circuit system only); a combination of hardware circuitry and software, such as (if applicable): a combination of (multiple) analog and / or digital hardware circuitry with software / firmware; and any portion of (multiple) hardware processors (including (multiple) digital signal processors), software, and (multiple) memories having software, which cooperate to enable a device such as a mobile phone or server to perform various functions, such as the functions or operations of the exemplary embodiments of the invention disclosed herein); and (c) (multiple) hardware circuitry and / or (multiple) processors, such as (multiple) microprocessors or portions thereof, which require software (e.g., firmware) to operate, but may be absent when not required to operate.
[0227] According to exemplary embodiments of the present invention, there are sufficient circuit systems for performing at least novel operations according to exemplary embodiments of the invention disclosed in this application, wherein "circuit system" as used herein refers to at least the following:
[0228] (a) Hardware circuit implementation only (such as implementation only in analog and / or digital circuit systems); and
[0229] (b) Combinations of circuitry and software (and / or firmware), such as (if applicable): (i) combinations of (multiple) processors, or (ii) portions of (multiple) processors / software (including (multiple) digital signal processors), software, and (multiple) memories, which work together to enable a device such as a mobile phone or server to perform various functions; and
[0230] (c) Circuits that require software or firmware to operate, such as (multiple) microprocessors or a portion thereof, even if the software or firmware does not exist physically.
[0231] This definition of "circuit system" applies to all uses of the term in this application, including in any claim. As another example, as used in this application, the term "circuit system" will also cover an implementation of only one processor (or multiple processors) or a portion of a processor and its accompanying software and / or firmware. For example, if applicable to a particular claim element, the term "circuit system" will also cover a baseband integrated circuit or application processor integrated circuit for a mobile phone, or a similar integrated circuit in a server, cellular network device, or other network device.
[0232] Generally, various embodiments can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while others may be implemented in firmware or software, which may be executed by a controller, microprocessor, or other computing device, but the invention is not limited thereto. While various aspects of the invention may be shown and described using block diagrams, flowcharts, or some other graphical representation, it is well understood that, by way of non-limiting example, the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0233] Embodiments of the present invention can be implemented in various components, such as integrated circuit modules. The design of integrated circuits is largely a highly automated process. Complex and powerful software tools can be used to convert logic-level designs into semiconductor circuit designs for etching and formation on semiconductor substrates.
[0234] As used herein, the term "exemplary" means "as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to other embodiments. All embodiments described in this detailed description are exemplary embodiments intended to enable those skilled in the art to make or use the invention, and are not intended to limit the scope of the invention as defined by the claims.
[0235] The foregoing description provides a complete and informative description of the best methods and apparatus currently contemplated by the inventors for carrying out the invention, through exemplary and non-limiting examples. However, various modifications and adaptations will become apparent to those skilled in the art when read in conjunction with the accompanying drawings and claims, in view of the foregoing description. Nevertheless, all such and similar modifications to the teachings of the exemplary embodiments of the invention will still fall within the scope of the invention.
[0236] It should be noted that the terms “connection,” “coupling,” or any variation thereof refer to any direct or indirect connection or coupling between two or more elements, and may encompass the presence of one or more intermediate elements between two elements that are “connected” or “coupled” together. The coupling or connection between elements can be physical, logical, or a combination thereof. As adopted herein, as several non-limiting and non-exhaustive examples, two elements may be considered “connected” or “coupled” together by means of one or more wires, cables, and / or printed electrical connections, and by means of electromagnetic energy (such as electromagnetic energy with wavelengths in the radio frequency region, microwave region, and optical (visible and invisible) region).
[0237] Furthermore, some features of the preferred embodiments of the invention can be used advantageously without the need for corresponding use of other features. Therefore, the above description should be considered merely as illustrating the principles of the invention, and not as limiting it.
Claims
1. An apparatus comprising: At least one processor; as well as At least one non-transitory memory stores instructions that, when executed by the at least one processor, cause the device to at least: The video decoder determines the reference sample region for the video segment samples derived from the cross-component model. The determination of the identified subset of available block vectors using either the same-position luminance coding unit or the same-position luminance prediction unit, and The subsets mentioned above are identified based on local attributes; Based on the determination, a cross-component prediction model is obtained; and The cross-component prediction model is used to decode the video segment samples.
2. The apparatus of claim 1, wherein the local attribute includes at least one of prediction cell size or prediction cell shape.
3. The apparatus of claim 2, wherein at least one of the prediction unit size or prediction unit shape is predetermined by the video decoder or received from the video encoder.
4. The apparatus of claim 1, wherein the reference sample region is an intra-frame block copy reference region.
5. The apparatus of claim 1, wherein the reference sample region is determined based on the availability of more than one block region in at least one prediction unit of the same brightness.
6. The apparatus of claim 5, wherein the determination comprises: The reference sample region is exported so that overlapping or redundant regions are discarded.
7. The apparatus of claim 5, wherein the average of the available block vectors is used as a block vector pointing to the reference sample region.
8. The apparatus of claim 7, wherein when the block vector size difference is less than a threshold, the block vector points to the reference sample region.
9. The apparatus of claim 7, wherein the average is calculated as a weighted average based on the magnitude of either the same-position luminance encoding unit or the same-position luminance prediction unit.
10. The apparatus of claim 9, wherein weights are assigned to the block vector of the at least one prediction unit based on the size of the prediction unit in the at least one prediction unit.
11. The apparatus of claim 7, wherein the average is calculated based on a fixed grid of block vectors.
12. The apparatus of claim 7, wherein at least one of linear estimation or nonlinear estimation is used to derive the block vector pointing to the reference sample region.
13. The apparatus of claim 7, wherein the maximum and minimum coordinates of a composite reference region defined by the co-location vector and the co-location prediction unit region are used to determine the reference sample region based on the more than one block region.
14. The apparatus of claim 13, wherein, based on the fact that the size of the reference sample region differs from the size of the at least one prediction unit, the apparatus is configured to scale the size of the reference sample region to match the size of the at least one prediction unit.
15. The apparatus of claim 1, wherein the at least one non-transitory memory stores instructions, the instructions being executed by the at least one processor, such that the apparatus: Identify temporary block vectors used for co-position prediction units that do not have associated block vectors; and The temporary block vector is used to determine at least one of the reference sample region or block vector obtained directly from the co-position prediction unit.
16. The apparatus of claim 1, wherein the spatial location of the block vector in the co-location region of one of the co-location brightness encoding units or the co-location brightness prediction units is used to interpolate or estimate the block vector pointing to the reference sample region derived for the cross-component model.
17. The apparatus of claim 7, wherein, based on the availability of multiple block vectors in the co-position prediction unit region, a template matching-based refinement mechanism is used to find at least one of the following: a matching reference region block, or a corresponding block vector in the identified subset of the available block vectors.
18. The apparatus of claim 17, wherein the template-matching-based refinement mechanism uses at least one sample of at least one of the co-located reference region blocks in the current block or at least one of the adjacent reference samples.
19. The apparatus of claim 7, wherein when multiple block vectors are available in the co-position prediction unit region, the chroma prediction unit is divided into multiple cross-component models based on the spatial location of the co-position prediction unit region.
20. The apparatus of claim 19, wherein block vector 0 is used to derive the cross-component model of the upper half of the chromaticity prediction unit marked as 0, and block vector 1 is used to derive the cross-component model of the lower half of the chromaticity prediction unit marked as 1.
21. The apparatus of claim 7, wherein when multiple block vectors are available in the co-position prediction unit region, multiple predictions are made using the multiple block vectors for multiple cross-component models; and a final prediction is obtained by combining the multiple predictions.
22. The apparatus of claim 21, wherein the weights for combining the plurality of predictions can be one of the following: weights defined in a codec specification, or an identifier index of weights signaled to the video decoder, or weights calculated on the decoder side based on block information including reconstructed samples and block size.
23. The apparatus of claim 5, wherein the statistics of co-location blocks of the more than one block region are used to prune training samples in the reference region blocks of the reference sample region pointed to by at least one block vector.
24. The apparatus of claim 23, wherein the minimum and maximum intensity values of samples in the co-location block are determined and used for pruning such that samples within the minimum and maximum intensity ranges are considered for at least one of: parameter calculation or training the cross-component prediction model.
25. The apparatus of claim 24, wherein the minimum intensity value and the maximum intensity value are extended by an incremental value from the lower limit and the upper limit of the maximum intensity range to provide a wider range for training, and wherein the incremental value is one of: fixed, or determined by the minimum intensity value and the maximum intensity value.
26. The apparatus of claim 24, wherein the cross-component model type is determined based on the distribution of samples derived from the co-located block and / or reference block to which the block vector points.
27. The apparatus of claim 24, wherein, for the distribution of sample ratios based on the classification parameters, a multi-model variant is determined for the cross-component prediction model, wherein the sample ratios are one of: predefined, or signaled to the decoder.
28. The apparatus of claim 26, wherein, in addition to or instead of minimum and maximum values, the cross-component prediction model type is determined by other parameters, wherein the other parameters include the mean of the samples used, wherein the mean includes a lower limit and an upper limit of the intensity distance used for calculation of the parameters, and wherein the intensity distance is one of the following: predefined, or signaled to the decoder.
29. The apparatus of claim 13, wherein when multiple block vectors are available in the co-position prediction unit region, the reference block for model derivation is determined by finding a matching region for the co-position block in the reference channel.
30. The apparatus of claim 1, wherein one or more of the correlation metrics are used to select samples from the reference sample region pointed to by the block vector of the co-position prediction unit.
31. A method comprising: The video decoder determines the reference sample region for the video segment samples derived from the cross-component model. The determination of the identified subset of available block vectors using either the same-position luminance coding unit or the same-position luminance prediction unit, and The subsets mentioned above are identified based on local attributes; Based on the determination, a cross-component prediction model is obtained; and The cross-component prediction model is used to decode the video segment samples.