Method and apparatus for chrominance component prediction based on reconstructed luminance block
By using inter-component prediction with luminance blocks to enhance chrominance block prediction, the method addresses inefficiencies in video encoding/decoding, improving data compression and restoration accuracy.
Patent Information
- Application Number
- PCT/KR2025/005128
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-15
- Filing Date
- 2025-04-15
- Publication Date
- 2025-10-23
AI Technical Summary
Existing video encoding/decoding technologies face inefficiencies in predicting chrominance components due to insufficient utilization of correlations with luminance components, leading to suboptimal data compression and accuracy in video signal restoration.
The method employs inter-component prediction using restored luminance blocks to predict chrominance blocks, utilizing various filters, scaling values, and nonlinear functions to enhance prediction accuracy and efficiency, including linear prediction methods and filter-based approaches.
Improves encoding/decoding efficiency and accuracy by leveraging correlations between luminance and chrominance components, resulting in more effective data compression and restoration of video signals.
Smart Images

Figure KR2025005128_23102025_PF_FP_ABST
Abstract
Description
Method and device for predicting chrominance components based on restored luminance blocks
[0001] The present disclosure relates to a method, device, and recording medium for encoding / decoding a video signal.
[0002] Video data is a massive data volume, and various compression techniques are being studied to reduce data size for efficient storage and transmission. Statistical analysis of redundant signals and various techniques utilizing these characteristics are being applied to video compression technologies. Furthermore, effective methods are being developed for transmitting the information required for the decoder to restore signals removed during the encoding process.
[0003] In video encoding / decoding technology, efficient data compression is performed by removing redundant information and restoring the removed redundant data using a method agreed upon between the encoding / decoding unit / decoding unit. In this case, in the case of chrominance components, there may be methods that utilize the similarity with the luminance component, and depending on the type and method of use of similar information, it may be possible to restore information on information that has been explicitly or implicitly removed.
[0004] The present disclosure aims to improve encoding / decoding efficiency in predicting a chrominance block using correlation between components.
[0005] The present disclosure seeks to improve the accuracy of component-by-component prediction.
[0006] The video encoding / decoding method, device, and recording medium of the present disclosure may include the steps of: determining whether to perform inter-component prediction using a corresponding luminance block for a current chrominance block; determining, in response to performing the inter-component prediction for the current chrominance block, one inter-component prediction mode among a plurality of inter-component prediction modes available to the current chrominance block; and predicting the current chrominance block based on the determined inter-component prediction mode to generate a prediction block of the current chrominance block.
[0007] In the video encoding / decoding method, device and recording medium of the present disclosure, prediction of the current color difference block can be performed by applying any one of a plurality of filters.
[0008] In the video encoding / decoding method, device, and recording medium of the present disclosure, the filter coefficients of the filter can be obtained using the surrounding reference sample area of the current chrominance block and the surrounding reference sample area of the luminance block.
[0009] In the video encoding / decoding method, device and recording medium of the present disclosure, the filter coefficient of the filter may differ depending on the type of the filter.
[0010] In the video encoding / decoding method, device and recording medium of the present disclosure, the function of the filter may include a nonlinear function using bit depth.
[0011] In the video encoding / decoding method, device and recording medium of the present disclosure, prediction of the current chrominance block can be performed based on a scaling value explicitly signaled from a bitstream and an offset calculated based on a surrounding reference sample value of the current chrominance block.
[0012] In the video encoding / decoding method, device and recording medium of the present disclosure, prediction of the current chrominance block can be performed using the average value or median value of all or part of the samples within the surrounding reference area of the luminance block.
[0013] In the video encoding / decoding method, device, and recording medium of the present disclosure, the scaling value can be determined by information indicating whether the scaling value explicitly signaled from the bitstream is 0, sign information of the scaling value, and size information of the scaling value.
[0014] In the video encoding / decoding method, device, and recording medium of the present disclosure, prediction of the current chrominance block can be performed using only one of the upper reference sample, the left reference sample, and the upper and left reference samples of the current chrominance block.
[0015] In the video encoding / decoding method, device, and recording medium of the present disclosure, prediction of the current chrominance block can be performed using the calculated DC value of the current chrominance block.
[0016] In the video encoding / decoding method, device, and recording medium of the present disclosure, the calculation of the DC value can be performed based on the sample bit depth depending on whether a surrounding reference sample exists.
[0017] According to the present disclosure, encoding / decoding efficiency can be improved in predicting a chrominance block using correlation between components.
[0018] The present disclosure can improve the accuracy of component-by-component prediction.
[0019] Figure 1 illustrates an embodiment in which prediction is performed through intra-screen prediction for a chrominance block.
[0020] Figure 2 illustrates an embodiment that represents directionality information of on-screen prediction through the main direction and the delta direction.
[0021] Figure 3 illustrates an example of a type of inter-component prediction mode through a luminance component.
[0022] Figure 4 illustrates an embodiment of prediction in CFL_EXPLICIT mode.
[0023] Figure 5 illustrates a list for an embodiment of signaling α of Cb and Cr through one index.
[0024] Figure 6 illustrates an embodiment of prediction in CFL_DERIVED mode.
[0025] Figure 7 illustrates an example of a surrounding reference sample area used to derive the α value.
[0026] Figure 8 illustrates an embodiment of a reference sample area in a corresponding luminance block of a current chrominance block to be encoded / decoded.
[0027] Figure 9 illustrates an embodiment of a downsampling filter.
[0028] Figure 10 illustrates reference samples used according to each mode.
[0029] Figure 11 illustrates an example of segmentation of a region according to a directional mode.
[0030] Figure 12 illustrates an example of segmentation of a region according to directional mode.
[0031] Figure 13 shows the filter coefficients of CFL_FITER mode.
[0032] Figure 14 illustrates an embodiment in which CFL_MERGE mode is performed.
[0033] FIG. 15 illustrates an encoding process according to one embodiment of the present disclosure in an encoder.
[0034] The present disclosure may be modified in various ways and may have various embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Similar reference numerals have been used to designate similar components throughout the description of each drawing.
[0035] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0036] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0037] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0038] Hereinafter, preferred embodiments of the present disclosure will be described in more detail with reference to the attached drawings. Hereinafter, identical components in the drawings will be designated by the same reference numerals, and redundant descriptions of identical components will be omitted.
[0039]
[0040] Intercomponent prediction using luminance components in chrominance component prediction (CFL: Chroma From Luma) method
[0041] In an embodiment of the present disclosure, when generating a prediction value of a chrominance block, prediction of the current chrominance block can be performed based on the restored luminance sample value of the corresponding luminance block. According to an embodiment, inter-component prediction can be performed through one or more of a linear prediction method based on a scale parameter and a filter-based prediction method derived from a reference sample identical to that of the encoder / decoder. Such a prediction method can be determined by flag or index information signaled from the bitstream.
[0042] In one embodiment, when linear prediction based on scale parameters is performed to generate a prediction value of a chrominance block, the AC value of the corresponding restored luminance block ( ) is multiplied by a scale parameter (α) that is explicitly transmitted and determined from the bitstream, and the chrominance DC value ( derived from the surrounding pre-decoded reference sample values of the current sub / decoding target chrominance block) is ) through the prediction value of the color difference block as follows ( ) can be generated. Unlike the above embodiment, the scale parameter can also be implicitly determined by information explicitly transmitted from the bitstream.
[0043]
[0044] At this time, the AC component of the corresponding restored luminance block ( ) can be calculated by differentiating each sample value in the corresponding restored luminance block and the average value, maximum value, minimum value, median value or weighted average value of all or some samples in the corresponding luminance block according to an embodiment.
[0045] In another embodiment, the scale parameter can be derived and used from the previously decoded chrominance samples surrounding the current sub / decoding target chrominance block and the corresponding reconstructed luma samples. Alternatively, the scale parameter can be derived and used from the chrominance samples of some or all of the previously decoded chrominance blocks surrounding the current sub / decoding target chrominance block and the luma samples of some or all of the reconstructed luma blocks corresponding thereto. Alternatively, the scale parameter can be derived and used from the chrominance samples of some or all of the previously decoded chrominance blocks surrounding the current sub / decoding target chrominance block and the luma samples of some or all of the previously decoded luma blocks surrounding the luma blocks corresponding to the current sub / decoding target chrominance block. The method applies the relationship between the chrominance samples and the luma samples surrounding the current sub / decoding target block to the current sub / decoding target block, and the scale parameter can be calculated by calculating a linear regression formula that takes as input the values of the chrominance samples and the luma samples surrounding the current sub / decoding target block.
[0046] Based on the calculated scale parameters, inter-component prediction can be performed using the following formula.
[0047]
[0048] At this time, in the case of an embodiment in which a scale parameter is derived, the AC component of the restored luminance block corresponding to the current sub / decoded chrominance block ( ) can be calculated by differentiating each sample value in the corresponding restored luminance block and the average value, maximum value, minimum value, median value or weighted average value of all or some samples in the corresponding luminance block according to an embodiment.
[0049]
[0050] Figure 1 illustrates an embodiment in which prediction is performed through intra-screen prediction for a chrominance block.
[0051] Chroma block prediction mode decision unit
[0052] The chrominance block prediction mode determination unit of the decoder of the present disclosure can determine the intra-picture prediction mode of the current chrominance block by parsing an intra-picture prediction mode index from a bitstream. At this time, according to an embodiment, a flag indicating whether the prediction mode of the chrominance block is a mode that performs inter-component prediction can be parsed to determine whether the inter-component prediction mode for the chrominance component has been applied, and if the mode is not applied, a mode index can be additionally received to determine the intra-picture prediction mode of the chrominance component of the current block. Here, the intra-picture prediction mode of the chrominance component can mean a general intra-picture prediction mode including a non-directional or directional prediction mode that generates a prediction block by using pre-decoded reference samples around the chrominance component.
[0053] A method for determining the intra-screen prediction mode of a chrominance component through the above mode index may include a method of determining dependently based on the intra-screen prediction mode of a luminance block corresponding to the chrominance block currently being decoded.
[0054] In one embodiment, the prediction mode for the chrominance component can be implicitly determined through information about the prediction mode for the luminance component, the surrounding modes of the pre-decoded chrominance component, etc., without the step of parsing the flag for the mode using inter-component prediction for the prediction mode for the chrominance component.
[0055] Alternatively, the application of inter-component prediction can be explicitly conveyed by indicating a mode for performing inter-component prediction on chrominance components and the general intra-screen prediction mode in a single syntax form. This may mean that the inter-component prediction mode is conveyed through a mode number of one of N prediction modes of chrominance components.
[0056] Figure 2 illustrates an embodiment that represents directionality information of on-screen prediction through the main direction and the delta direction.
[0057] In the on-screen prediction mode, signaling and parsing can be performed on the on-screen prediction mode information divided into base angle and delta angle as an example of the directional mode.
[0058] Referring to Fig. 2, in an embodiment that indicates directionality information of prediction within a screen through a main direction and a delta direction, D203_PRED, H_PRED, D157_PRED, D135_PRED, D113_PRED, V_PRED, D67_PRED, and D45_PRED may denote main directions. In addition, the main directions may be divided into K, and the directions divided into K may be delta directions.
[0059] In one embodiment, K may be variably varied depending on at least one of the size, number of horizontal pixels, number of vertical pixels, ratio of horizontal and vertical pixels, shape, or component of the current block to be encoded / decoded.
[0060] Alternatively, K may have a specific value depending on at least one condition on the size of a particular block, the aspect ratio of the block, the shape or composition, according to an agreement between the encoder and decoder.
[0061] Alternatively, K can be signaled with higher-level syntax.
[0062] In one embodiment, K between each principal direction may not be the same. For example, K may be 6 between a specific principal direction and K may be 4 between another principal direction, and so on, so that the number of delta intervals between all principal directions may not be the same. This may be adaptively changed based on at least one of the size of the current block to be encoded / decoded, the number of horizontal pixels, the number of vertical pixels, the aspect ratio of the horizontal and vertical pixels, the shape, or the component, depending on the embodiment. In another embodiment, signaling for the delta direction may be indicated by setting the direction from -K / 2 to +K / 2 with respect to the principal direction as the delta direction of the corresponding principal direction and dividing it into the delta direction for the principal direction. For example, when K is 6, the direction from -3 to +3 with respect to the principal direction may be the delta direction of the corresponding principal direction, and accurate directional information for prediction within the picture may be determined through the index for the principal direction and the index for the delta direction.
[0063] Depending on the embodiment, the delta direction interval between the principal and principal directions may be equidistant or unequal, and the angle according to the delta direction index may be transmitted from the encoder to the decoder via a higher level syntax, or may be determined by an agreement between the sub- and decoders without transmission.
[0064] In some embodiments, when the inter-component prediction mode through the luminance component is not applied as the prediction mode of the chrominance component, the prediction of the chrominance component may be performed through a non-directional or directional mode. In this case, in order to signal which mode among the non-directional or directional mode was used for the prediction, the non-directional or directional mode may be managed as a list for the chrominance component prediction mode, and the list may be composed of one or more of the non-directional modes such as DC, SMOOTH, SMOOTH_V, SMOOTH_H, PAETH, or the directional modes such as V, H, D45, D135, D67, D113, D157, D203. In some embodiments, an initial list may be generated from some of the non-directional and directional modes, and the list may be stored and managed in a form in which the list is changed by adding a mode during the encoding / decoding process.
[0065] The order of the indexes in the list can be managed by a fixed index number and a mode mapped to it, as agreed upon between the sub-decoder and the decoder. Alternatively, the order of the list can be variably changed based on the intra-screen prediction mode of the luminance block corresponding to the chrominance block currently being decoded, as agreed upon between the sub-decoder and the decoder.
[0066] In one embodiment, if the intra-screen prediction mode of the co-located luminance block is a directional mode, the primary direction of the directional mode may be added as the first mode in the list, and non-directional modes and directional modes may be sequentially added to the list. At this time, when adding a directional mode, a directional mode that has already been added may be excluded from the list addition. In addition, the order of adding list modes may vary depending on the embodiment and the agreement between the decoder and the decoder.
[0067] In an embodiment, if the parsed index points to the directional mode of the corresponding luminance block, the delta direction can be used identically to the delta direction of the corresponding luminance block.
[0068] At this time, if the parsed index points to a directional mode other than the directional mode of the corresponding luminance block, the delta direction can be implicitly determined as 0.
[0069] In the present disclosure, when inter-component prediction is used based on corresponding luminance components in predicting the current target chrominance block to be encoded / decoded, it is possible to additionally determine which type of inter-component prediction through luminance components is used.
[0070] Whether to use inter-component prediction via luminance components for chroma blocks can be explicitly passed from the encoder to the decoder via related flags, prediction mode information syntax for chroma blocks, etc., or can be implicitly determined in the decoder / sub-decoder via information about surrounding pre-decoded chroma blocks, size, shape, etc. of the chroma blocks.
[0071] Figure 3 illustrates an example of a type of inter-component prediction mode through a luminance component.
[0072] The types of inter-component prediction modes through luminance components may vary depending on the embodiment, and FIG. 3 illustrates one embodiment among various examples of types of inter-component prediction modes through luminance components. That is, this is only a part of the embodiments of the invention, and depending on the applied embodiment, modes may be added, some modes may not be used, and the order of the modes indicated by the index may be changed.
[0073] Depending on the embodiment, information about the mode may be applied in one or more of the following ways: by transmitting related index information from the encoder to the decoder, or by implicitly determining the prediction mode of the corresponding luminance block of the current target chrominance block to be decoded. In the case of implicit determination, if the inter-component prediction mode is determined as one through the mode of the corresponding luminance block, information about the inter-component prediction mode may be determined by the decoder without separately transmitting it from the encoder and without parsing the corresponding information.
[0074] If, according to an embodiment, the inter-component prediction mode is determined in the form of a list including a portion of the mode list of FIG. 3 through the mode of the corresponding luminance block, the encoder transmits index information from a list composed of a portion of the mode list of FIG. 3, and the decoder parses the information to determine the inter-component prediction mode through the luminance component.
[0075] For example, the type of inter-component prediction mode can only include three modes: CFL_EXPLICIT, CFL_DERIVED, and CFL_FILTER.
[0076]
[0077] Figure 4 illustrates an embodiment of prediction in CFL_EXPLICIT mode.
[0078] CFL_EXPLICIT mode is a mode that predicts the current chrominance block based on the restored luminance block, as in the embodiment of Fig. 4. Prediction can be performed with .
[0079] At this time, is the chrominance prediction value at position (x,y), α is a scaling value determined through explicit signaling, may mean an average or median chrominance component value calculated based on surrounding reference sample values of the current chrominance component.
[0080] The calculation may be performed by performing subsampling or downsampling on some or all of the samples of the reference sample area of the restored luminance block corresponding to the current sub / decoding target chrominance block according to the color space indicating the components of luminance, chrominance, etc. of the current frame, such as YUV444, YUV422, YUV420, etc., and calculating based on the corresponding samples. At this time, the subsampling or downsampling method may be determined in units such as sequence, picture, slice, or coding block, depending on the embodiment.
[0081] According to an embodiment, the present invention may be calculated by calculating an average value, a maximum value, a minimum value, a median value or a weighted average value for all or part of samples or downsampled samples in a peripheral restoration reference area of a restoration luma block corresponding to a current sub / decoding target chroma block, and differentiating the calculated average value, maximum value, minimum value, median value or weighted average value from pixel values of the corresponding luminance block. Alternatively, the present invention may be calculated by calculating an average value, a maximum value, a minimum value, a median value or a weighted average value for all or part of samples or downsampled samples in a peripheral restoration luma reference area corresponding to a current sub / decoding target chroma block, and differentiating the calculated average value, maximum value, minimum value, median value or weighted average value from pixel values of the corresponding luminance block.
[0082] In another embodiment The current target chrominance block to be encoded / decoded may be calculated by calculating the average, maximum, minimum, median or weighted average of all or part of the samples or downsampled samples in the luminance block corresponding to the luminance block, and then differentiating the average, maximum, minimum, median or weighted average calculated above from the pixel values of the corresponding luminance block.
[0083] Figure 5 illustrates a list for an embodiment of signaling α of Cb and Cr through one index.
[0084] α can be explicitly signaled for both the Cb component and the Cr component, and the transmission values for Cb and Cr can be transmitted in a form where the sign and absolute value are transmitted separately, or where at least one of the sign and absolute value of the α values for Cb and Cr is shared. That is, α can be transmitted separately with the same sign, or in a form where the sign is transmitted separately and the same α is used. In addition, α can be signaled in a form where a flag is transmitted for whether it has a specific value of 0 or 1, and the signed absolute value is transmitted if it is not the specific value. Depending on the embodiment, α can be transmitted in the form of a direct value or an index into a list for α. In this case, the list can be fixed by a promise between the encoding / decoding periods or can be transmitted through a higher-level syntax. Depending on the embodiment, the list can change variably depending on the previously occurred α, rather than having a fixed size or order. The list can be managed and stored separately for Cb and Cr, or managed as a single list for Cb and Cr.
[0085] If α is parsed through an index for whether it is 0 and the sign value, the final α can be determined by additional index parsing for components that are not 0.
[0086]
[0087] Figure 6 illustrates an embodiment of prediction in CFL_DERIVED mode.
[0088] CFL_DERIVED mode can perform inter-component prediction by implicitly deriving an α value based on the surrounding pre-decoded luminance reference sample of the luminance block corresponding to the current sub / decoding target chrominance block and the surrounding pre-decoded reference sample of the current sub / decoding target chrominance block.
[0089]
[0090] Figure 7 illustrates an example of a surrounding reference sample area used to derive the α value.
[0091] The surrounding reference sample area used to derive the α value may vary depending on the embodiment. Fig. 7 may only illustrate one embodiment thereof.
[0092] Figure 8 illustrates an embodiment of a reference sample area in a corresponding luminance block of a current chrominance block to be encoded / decoded.
[0093] In some embodiments, if a, b, c, d >=1, b=c=1, d=w, a=h may be obtained, and this may be changed by an agreement between the encoder / decoder as an example. In addition, in some embodiments, the values of a, b, c, and d may be variably changed during the encoder / decoder process.
[0094] In one embodiment, the values of a, b, c, and d may be determined based on an agreement between the encoder and decoder, such as the availability of upper and left reference samples and / or the aspect ratio of the current block, and the maximum number of reference samples to use may be determined based on an agreement between the encoder and decoder.
[0095] The reference sample area of the current chrominance block used to derive the α value according to the color space of the current picture may be as shown in Fig. 7. In the case where the current color format is a 4:2:0 format, S x = S y = 2, S for 4:2:2 format x =2, S y =1, S for 4:4:4 format x = S y = 1 can be used.
[0096]
[0097] Figure 9 illustrates an embodiment of a downsampling filter.
[0098] Depending on the color space of the current frame, the luminance reference sample may be downsampled or subsampled to the same size as the reference sample area of the chrominance block. An embodiment of the downsampling filter may be as shown in FIG. 7. A flag or index for selecting which filter among a plurality of downsampling filters to use may be signaled from the encoder to the decoder in units of a sequence, a frame, a slice, a plurality of blocks, or a block. Alternatively, it may be implicitly determined by an agreement between the sub- and decoders. Depending on the embodiment, information on the location at which the downsampling filter is to be performed in the reference sample area may be determined by a method implicitly determined according to the color space relationship between the chrominance component and the luminance component, a method implicitly determined according to an agreement between the sub- and decoders, or a method determined through explicit signaling at a higher level or on a block-by-block basis.
[0099]
[0100] The scale parameter (α) can be derived as in the following equation (1) based on linear regression. In this case, refers to the values obtained by subtracting the mean, maximum, minimum, median, or weighted mean from the surrounding restoration reference samples of the corresponding luminance block for which subsampling or downsampling was performed according to the color space. may mean values obtained by subtracting the mean or median from the surrounding reference samples of the current color difference block.
[0101] Formula (1)
[0102] Depending on the embodiment, it can be used in a simplified form as in Equation (2). In this case, Equation (2) means the kth largest value among the AC values calculated in the reference sample area of the restored luminance component. means the color difference value corresponding to the corresponding luminance pixel location. is the kth smallest value among the values calculated from the reference sample area of the restored luminance component minus the average, maximum, minimum, median, or weighted average. may mean the color difference value corresponding to the corresponding luminance pixel location.
[0103] Formula (2)
[0104] Depending on the embodiment, the scale parameter may be calculated based on the surrounding reference samples that are not differentiating the mean, maximum, minimum, median, or weighted average, as in the embodiment of Equation (2). In this case, M used in deriving the scale parameter may mean the number of left reference samples, N may mean the number of upper reference samples, and K may mean the number of upper left reference samples.
[0105] Depending on the embodiment, the calculation of the scale parameter is performed by downsampling or subsampling the corresponding luminance block as in Equation (3) by taking the maximum value of n among the sample values. ), minimum ( ) and the surrounding chrominance component values corresponding to the corresponding luminance sample value ( , ) can be calculated based on the value. At this time, and The surrounding restoration reference sample values or the mean or median values can be used as calm sample values.
[0106] Formula (3)
[0107] Depending on the embodiment, the CFL mode can be divided into multiple modes depending on the location of the reference sample used to derive the scale parameter.
[0108] Figure 10 illustrates reference samples used according to each mode.
[0109] Referring to Fig. 10, the first mode may be configured to use only the left reference sample, the second mode to use only the upper reference sample, and the third mode to use the left and upper reference samples. In this case, depending on the embodiment, the values of a, b, c, and d may be configured as integers greater than 1, and the reference samples in the b×c area on the upper left may not be used to calculate the scale parameter.
[0110] At this time, it is configured as CFL_DERIVED, CFL_DERIVED_L, CFL_DERIVED_T mode for each reference sample position, and can be parsed and determined for each mode through the CFL mode index parsing unit. At this time, the CFL_DERIVED mode may refer to a mode that uses the left and upper true samples, CFL_DERIVED_L may refer to a mode that uses the left reference sample, and CFL_DERIVED_T may refer to a mode that uses the upper reference sample. According to an embodiment, if the CFL_DERIVED mode is selected in the CFL mode index parsing unit, a 1-bit flag may be additionally parsed, and if the flag is 0, the scale parameter may be derived using the left and upper reference samples, and if the flag is 1, a 1-bit flag may be additionally parsed to determine whether to use only the upper reference sample or only the left reference sample to derive the scale parameter. At this time, when performing inter-component prediction according to the selected mode, the DC value (of the current chrominance block) ) can calculate the DC component through the average value of the left and upper reference samples or the left, upper and upper left reference samples regardless of the currently selected mode, or, depending on the embodiment, can calculate the average value of the area used to derive the scale parameter as the DC value of the current chrominance block. At this time, when calculating the DC component, if the left or upper reference sample does not exist, the calculated DC value of each pixel can be calculated through the DC calculated through one of the reference samples through the same procedure as when the intra-screen prediction mode is DC, and if neither the left nor the upper reference samples exist, the median value according to the pixel bit depth can be calculated as the DC value. Depending on the embodiment, the calculation of the DC value can be performed through the same process as the DC value in the intra-screen prediction mode.
[0111] Depending on the embodiment, the reference sample area to be used for deriving scale parameters may be implicitly determined based on the intra-screen prediction mode of the corresponding luminance block, or the mode or flag or index may be parsed using separate entropy context information based on the intra-screen prediction mode of the corresponding luminance block when parsing the CFL mode or parsing the flag or index for determining the reference sample area.
[0112] Figure 11 illustrates an example of segmentation of a region according to a directional mode.
[0113] Referring to FIG. 11, according to an embodiment, if the prediction mode within the screen of the corresponding luminance block of the current chrominance block is a directional mode, the region may be divided according to the directional mode. At this time, if the directional mode of the corresponding luminance block is a first region mode, a mode for deriving a scale parameter based on a left reference sample region, a mode for deriving a scale parameter based on a left and upper reference sample region if the directional mode of the corresponding luminance block is a second region mode, a mode for deriving a scale parameter based on a left and upper reference sample region if the directional mode of the corresponding luminance block is a third region mode, a mode for deriving a scale parameter based on an upper reference sample region may be implicitly derived, or a flag for determining a mode index or a reference sample region may be decoded using separate entropy context information according to the region. At this time, in one embodiment, the first mode may be H_PRED, the second mode may be D135_PRED, and the third mode may be V_PRED.
[0114] Figure 12 illustrates an example of segmentation of a region according to directional mode.
[0115] Referring to FIG. 12, depending on the embodiment, the image may be divided into two regions, and at this time, a 1-bit flag that determines whether the current mode is CFL_DERIVED mode and whether to derive scale parameters using left and upper reference samples is parsed, and if the flag is 0, the directional mode of the corresponding luminance block may be implicitly derived as a mode for deriving scale parameters using the left reference sample if it exists in the first region, or as a mode for deriving scale parameters using the upper reference sample if it exists in the second region. At this time, if the prediction mode of the corresponding luminance block is a non-directional mode, a 1-bit flag may be additionally parsed to determine a region to be used for deriving scale parameters among the left reference sample and the upper reference sample.
[0116] According to an embodiment, the on-screen prediction mode of the corresponding luminance block is divided into a first region, a second region, and a non-directional mode, and a flag is parsed to determine the region to be used for deriving scale parameters among the left reference sample and the upper reference sample according to each mode, and decoding can be performed using separate entropy context information for each.
[0117] In some embodiments, when there are multiple luminance blocks corresponding to the current chrominance block, the entropy context information determination may be performed based on the intra-screen prediction mode of the luminance block at a specific location determined by an agreement between the decoder and the encoder (implicit mode or flag derivation). In this case, the process may be performed based on the intra-screen prediction mode of the corresponding luminance block at the upper left location of the current chrominance block.
[0118] According to an embodiment, the selection of the corresponding luminance block may be performed by checking the prediction mode of the corresponding luminance block within the screen according to a predetermined scanning order among a predetermined number of positions, and then the block encoded in the directional mode may be determined as the prediction mode of the corresponding luminance block first.
[0119] Depending on the embodiment, the α value for one of the Cb and Cr chrominance components may be explicitly signaled, and the α value for the other component may be implicitly derived based on surrounding reference samples. In this case, a 1-bit flag that determines whether to implicitly derive the α value for either the Cb or Cr component may be additionally parsed and determined. Depending on the embodiment, the flag may be parsed when the current CFL mode is a mode (CFL_EXP_DER) in which the α value is derived implicitly for one component and explicitly for the other component for the two chrominance components.
[0120] According to an embodiment, a plurality of linear models may be applied to a block to perform prediction of a current chrominance block. When predicting a current chrominance block using two linear models, two α values may be derived for each chrominance component based on the surrounding reference samples of the corresponding luminance block and the surrounding reference sample values of the current chrominance block. At this time, according to an embodiment, for the corresponding luminance block, values greater than or equal to a specific threshold may be divided into a first group, and values less than a specific threshold may be divided into a second group, and for each group, a scale parameter may be derived using the same method as the encoder / decoder based on the surrounding reference samples of the corresponding luminance block and the reference sample values of the chrominance block at the corresponding position. At this time, when the scale parameters calculated in each of the two groups are applied to actual prediction, if the AC value obtained by removing the average component from the corresponding luminance is greater than or equal to the threshold, the scale parameter calculated through the first group may be used to perform CFL prediction, and if the AC value is less than the threshold, the scale parameter calculated through the second group may be used to perform CFL prediction.
[0121] When the CFL_MM mode is applied according to the embodiment, the DC value of the current chrominance block ( ) can be calculated based on the surrounding reconstructed reference sample, and at this time, the first DC value is calculated by calculating the average for the chrominance pixel values whose value of the corresponding luminance reference sample is the first group based on the value of the corresponding luminance reference sample, and the second DC value is calculated by calculating the average for the chrominance pixel values whose value of the luminance reference sample is the second group. In an embodiment, when predicting one chrominance block using multiple scale parameters, the DC value of the chrominance block can be calculated in the same manner according to an agreement between the encoder and decoder.
[0122] In some embodiments, when α values are implicitly derived for both chrominance components and CFL prediction is performed, correction can be performed on the implicitly derived α value by signaling the delta slope value (u) through an index. Parsing of the delta slope value can be determined by parsing a 1-bit flag indicating whether to parse the delta slope value if the currently selected CFL mode is a mode that implicitly derives α value (CFL_DERIVED). If the flag is 1, the delta slope value can be parsed, and at this time, an index value indicating which delta slope among n delta slope values determined according to the encoder / decoder agreement will be used can be parsed. In some embodiments, the delta slope value can be determined based on the implicitly derived α value. In some embodiments, whether to parse the delta slope can be determined through one flag for Cb and Cr, or whether to parse the delta slope can be determined through each flag. Depending on the embodiment, when parsing for delta slope values, one delta slope value may be parsed for Cb and Cr, or a separate delta slope value may be parsed for each chrominance component through each index to determine the parsing.
[0123] At this time, if the CFL mode (CFL_MM) that predicts one block using multiple models is selected, CFL prediction can be performed by parsing the delta gradient value for each model or applying one delta gradient value to multiple models through the delta gradient index parsed for each chrominance component (Cb, Cr) or two chrominance components.
[0124] According to an embodiment, after the delta slope (u) is applied, CFL prediction can be performed through an embodiment such as the following equation (4).
[0125] Formula (4)
[0126]
[0127] Figure 13 shows the filter coefficients of CFL_FITER mode.
[0128] In some embodiments, the CFL_FILTER mode may be used to derive filter coefficients using surrounding reference samples and perform inter-component prediction using the filter coefficients. The filter coefficients of the CFL_FITER mode may be as in the embodiment of FIG. 13.
[0129] Depending on the embodiment, the type of filter to be used in CFL_FILTER mode can be determined by parsing at a higher level, or prediction can be performed using the type of filter promised between encoding / decoding.
[0130] Depending on the embodiment, the filter type can be signaled and determined on a block-by-block basis, and in one embodiment, when one of the third or fourth filters is selected on a block-by-block basis, the filter type can be determined by parsing a 1-bit flag.
[0131] For n restoration reference samples of the surrounding reference samples of the corresponding luminance blocks of Figs. 7 and 8 and the surrounding reference sample area of the current chrominance block, The filter coefficient (X) can be calculated through the determinant Gaussian elimination method, LDL decomposition, etc. At this time, is the currently restored reference sample of the chroma block, refers to the restored reference sample of the corresponding luminance block. may mean a restored luminance reference sample for which downsampling has been performed according to the color space of the current picture.
[0132] At this time, for the first filter, color difference component prediction can be performed through the following filter.
[0133]
[0134] For the second filter, color difference component prediction can be performed through the following filter.
[0135]
[0136] For the third and fourth filters, color difference component prediction can be performed through the following filters.
[0137]
[0138] At this time means the sample value of the color difference block that is currently being predicted. represents the restored luminance sample value of the corresponding position, and a~h represent the filter coefficients derived through the surrounding restored reference samples. At this time, may refer to a restored luminance sample that has been downsampled according to the color space. According to an embodiment, is a nonlinear function and can be calculated as follows:
[0139]
[0140] Depending on the embodiment, B can be calculated as an offset value as follows:
[0141] B = 1 << (bitDepth - 1)
[0142] Depending on the embodiment, inter-component prediction may be performed using only some of the filter coefficients among the first to fourth filter coefficients, depending on the agreement between the encoder and decoder. In one embodiment, for the third and fourth filters, inter-component prediction may be performed using only the filter coefficients a, b, d, and e, or b, c, d, and e.
[0143]
[0144] Figure 14 illustrates an embodiment in which CFL_MERGE mode is performed.
[0145] According to an embodiment, before parsing the CFL mode index to be applied to the current chroma block through the CFL mode index parsing unit, a 1-bit flag for determining whether or not the CFL_MERGE mode is used can be parsed, and if the flag is 1, the CFL_MERGE mode can be performed.
[0146] CFL_MERGE_flag is parsed, and if the flag is 0, the CFL mode information of the current chroma block is parsed through the CFL mode index parsing unit. If the flag is 1, the CFL merge list for the current chroma block is constructed through the CFL_merge list construction unit according to the agreement between the decoder and the decoder. The merge list can be constructed through spatial, temporal, history-based candidates, and default candidates based on the current chroma block.
[0147] A spatial merge candidate can be added to the merge list by scanning adjacent blocks surrounding the current chroma block at a position and order determined by an agreement between the encoder and decoder, and if the block is in CFL mode, the mode information and scale parameter or filter parameter information, etc.
[0148] If there is a mode predicted as a CFL mode among the blocks at the corresponding position of the reference frame promised in the encoder / decoder, the temporal merge candidate can be added as a CFL merge candidate for the corresponding mode information and parameters.
[0149] A history-based CFL merge candidate can store the mode information and parameters in a history CFL buffer when a chroma block is decoded in CFL mode, and add the candidate load of the buffer to a CFL merge candidate list. Depending on the embodiment, the history CFL buffer can have information of up to R CFL modes, and if the number of buffers is the maximum number of buffers based on the FIFO structure, the mode that was added to the buffer first can be removed from the buffer. At this time, when adding to the buffer, if a list of the same parameters of the same mode already exists in the buffer through a duplication check, the mode information is not added, and the existing mode information can be managed so that it can be removed last from the buffer. The CFL mode buffer can be initialized in the first superblock of each row.
[0150] If the number of merge candidates promised between the sub-encoder and decoder is not fully filled, a default merge candidate can be added to the candidate list. The default merge candidate can be composed of L scale parameters promised between the sub-encoder and decoder, and CFL prediction can be performed based on the scale parameters. An example of the default candidate can be {1 / 8, -1 / 8, 2 / 8, -2 / 8, 3 / 8, -3 / 8, 4 / 8, -4 / 8, 5 / 8, -5 / 8, 6 / 8}.
[0151]
[0152] The chrominance block prediction unit can generate a prediction signal of the current chrominance block through a restoration value of the corresponding restoration luminance block according to the scale parameter or filter parameter determined by the CFL parameter determination unit when the uv_mode determined by the chrominance block prediction mode determination unit is determined as the CFL prediction mode.
[0153] In the above embodiment, the luminance block corresponding to the current sub / decoding target block, the surrounding reference area of the corresponding luminance block, and the surrounding reference area of the current sub / decoding target block are illustrated as areas having an area for the sake of ease of drawing, but this may be a size including one pixel or a pixel line. In addition, the sample values used in the formulas for calculating the scale value, the median value, the average value, etc. may include all of the methods of using some or all of the values of the samples belonging to the corresponding luminance block, the surrounding reference area of the corresponding luminance block, the current sub / decoding target chrominance block, and the surrounding reference area of the current sub / decoding target chrominance block. In this case, the method of selecting some of the values among the methods of using some of the samples may include methods of selecting a number of samples smaller than the number of samples belonging to the reference area by one or more of the following methods: a method of specifying a location through a specific coordinate of the corresponding sample, a method of excluding values in a specific range or selecting values in a specific range, and a method of using subsampling or downsampling.
[0154]
[0155] The residual block entropy decoding unit can restore quantized second-order transform coefficients when a second-order transform is applied, and can perform restoration of quantized first-order transform coefficients when a second-order transform is not applied.
[0156] The inverse quantization unit can perform inverse quantization by parsing information such as quantization method and quantization parameter information for the restored transform coefficients to obtain inverse quantized transform coefficients.
[0157] The inverse transform unit can perform inverse transform on the inverse quantized transform coefficients using the inverse transform kernel.
[0158] The color difference block restoration unit can generate a final restoration signal by combining the residual signal restored through the inverse transformation unit with the final prediction signal generated through the prediction performing unit.
[0159]
[0160] FIG. 15 illustrates an encoding process according to one embodiment of the present disclosure in an encoder.
[0161] On the encoder side, the residual block acquisition unit can acquire a residual signal based on the prediction signal generated through the chroma block prediction unit and the original signal. Here, the generated prediction signal may be a prediction signal that exhibits optimal encoding / decoding efficiency.
[0162] The transformation unit can perform transformation on the residual signal obtained from the residual block acquisition unit using a transformation kernel.
[0163] The quantization unit can perform quantization on the transformed coefficients to obtain quantized transformed coefficients. Here, the quantization method and quantization parameter information can be encoded and transmitted to the decoder.
[0164] The residual block entropy encoding unit can encode quantized transform coefficients and transmit them to a decoder via a bitstream. If a secondary transform is applied, the quantized secondary transform coefficients can be encoded, and if no secondary transform is applied, the quantized primary transform coefficients can be encoded.
[0165]
[0166] While the exemplary methods of this disclosure are presented as a series of operations for clarity of description, this is not intended to limit the order in which the steps are performed, and individual steps may be performed simultaneously or in different orders, if desired. To implement a method according to this disclosure, additional steps may be included in addition to the steps illustrated, some steps may be excluded and the remaining steps included, or some steps may be excluded and additional steps included.
[0167] The various embodiments of the present disclosure are not intended to list all possible combinations but rather to illustrate representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combinations of two or more.
[0168] Additionally, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, the embodiments may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0169] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.
[0170] The present disclosure may be applicable to an industry relating to a method, device and recording medium for encoding / decoding a video signal.
Claims
1. A step of determining whether to perform inter-component prediction using a corresponding luminance block for the current chrominance block; In response to performing the inter-component prediction on the current chrominance block, determining one inter-component prediction mode among a plurality of inter-component prediction modes available to the current chrominance block; An image decoding method, comprising a step of predicting the current chrominance block based on the determined inter-component prediction mode and generating a prediction block of the current chrominance block.
2. In paragraph 1, An image decoding method, characterized in that the prediction of the current color difference block is performed by applying any one of a plurality of filters.
3. In paragraph 2, An image decoding method, characterized in that the filter coefficients of the above filter are obtained using the surrounding reference sample area of the current chrominance block and the surrounding reference sample area of the luminance block.
4. In paragraph 3, An image decoding method, characterized in that the filter coefficients of the above filter are different depending on the type of the filter.
5. In paragraph 4, An image decoding method, characterized in that the function of the above filter includes a nonlinear function using bit depth.
6. In paragraph 1, A method for decoding an image, characterized in that the prediction of the current chrominance block is performed based on an offset calculated based on a scaling value explicitly signaled from a bitstream and a surrounding reference sample value of the current chrominance block.
7. In paragraph 6, An image decoding method, characterized in that the prediction of the current chrominance block is performed using the average value or median value of all or part of the samples within the surrounding reference area of the luminance block.
8. In paragraph 7, A video decoding method, characterized in that the scaling value is determined by information indicating whether the scaling value explicitly signaled from the bitstream is 0, sign information of the scaling value, and size information of the scaling value.
9. In paragraph 1, An image decoding method, characterized in that the prediction of the current chrominance block is performed using only one of the upper reference sample, the left reference sample, and the upper and left reference samples of the current chrominance block.
10. In paragraph 9, An image decoding method, characterized in that the prediction of the current chrominance block is performed using the calculated DC value of the current chrominance block.
11. In paragraph 10, An image decoding method, characterized in that the calculation of the DC value is performed based on the sample bit depth depending on whether a surrounding reference sample exists.
12. A step of determining whether to perform inter-component prediction using a corresponding luminance block for the current chrominance block; In response to performing the inter-component prediction on the current chrominance block, determining one inter-component prediction mode among a plurality of inter-component prediction modes available to the current chrominance block; An image encoding method, comprising a step of predicting the current chrominance block based on the determined inter-component prediction mode and generating a prediction block of the current chrominance block.
13. A step of determining whether to perform inter-component prediction using a corresponding luminance block for the current chrominance block; In response to performing the inter-component prediction on the current chrominance block, determining one inter-component prediction mode among a plurality of inter-component prediction modes available to the current chrominance block; A step of predicting the current chrominance block based on the determined inter-component prediction mode and generating a prediction block of the current chrominance block; A bitstream transmission method, comprising a step of transmitting a bitstream obtained by encoding the above prediction block.
Citation Information
Patent Citations
Cryopump with improved regeneration speed
KR102661608B1
Intra-prediction mode-based image processing method and apparatus therefor
WO2018236031A1
Image encoding / decoding method and recording medium storing bitstream
WO2024058595A1
KR20230169953A
KR20240042644A