Method and apparatus for video coding and decoding using cross-component prediction based on reconstructed reference samples
By deriving a filter and a cross component prediction model of the brightness component based on the reconstructed reference samples in the video encoding and decoding method, the problem of low cross component prediction efficiency in the prior art is solved, and more efficient video encoding and decoding and better video quality are achieved.
Patent Information
- Application Number
- CN202380079980.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-11
- Filing Date
- 2023-09-13
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively perform cross-component prediction in video encoding and decoding, resulting in low video encoding and decoding efficiency and poor video quality.
In the video encoding and decoding method, filters and cross component prediction models of luminance components are derived based on the reconstructed reference samples, for predicting the current chrominance block when cross component prediction is performed.
Improves video encoding and decoding efficiency and enhances video quality.
Smart Images

Figure CN120226347A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video coding method and apparatus using cross-component prediction based on reconstructed reference samples. Background Art
[0002] The statements in this section merely provide background information related to the present invention and may not necessarily constitute prior art.
[0003] Since video data has a large amount of data compared to audio data or still image data, video data requires a large amount of hardware resources (including memory) to store or transmit the video data without being compressed.
[0004] Accordingly, the encoder is generally used to compress and store or transmit video data. The decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression technologies include H.264 / Advanced Video Codec (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Codec (VVC), which improves the coding efficiency of HEVC by about 30% or more.
[0005] However, as image size, resolution and frame rate gradually increase, the amount of data to be encoded will also increase. Accordingly, new compression technologies are needed that provide higher encoding and decoding efficiency and improved image enhancement effects than existing compression technologies.
[0006] Cross-component prediction or cross-prediction utilizes the reconstructed luminance component to predict the chrominance component. Example cross-component prediction may include a cross-component linear model (CCLM), a gradient linear model (GLM), etc. CCLM applies a linear model to the reconstructed luminance component to predict the chrominance component. The linear model represents a cross-component prediction model. The linear model can be estimated by utilizing the adjacent reference samples of the "corresponding luminance block" co-located with the current chrominance block and the adjacent reference samples of the current chrominance block. GLM applies a linear model to the gradient of the reconstructed luminance component to predict the chrominance component. The same linear model used in CCLM can be used here. A gradient can be generated by applying a gradient filter to the luminance sample at a position corresponding to the current chrominance sample.
[0007] CCLM applies a linear model directly to the luminance component and thus predicts the chrominance component. GLM applies a linear model to which a gradient filter is applied to the luminance component and thus predicts the chrominance component. At this point, the encoder can signal the decoder to indicate the index of the type of the gradient filter. Therefore, there is a need for a method for effectively performing cross-component prediction to improve video encoding and decoding efficiency and enhance video quality. Summary of the Invention
[0008] Technical Problem
[0009] The present invention is dedicated to providing a video encoding and decoding method and apparatus, which are used to derive a filter for a luminance component and a cross-component prediction model for cross-component prediction based on reconstructed reference samples when predicting a current chrominance block according to cross-component prediction.
[0010] Technical Solution
[0011] At least one aspect of the present invention provides a method for a video decoding apparatus to reconstruct a current chrominance block. The method includes establishing a corresponding luminance block of the current chrominance block from a reconstructed region of the luminance component. The corresponding luminance block represents a luminance block at the same position as the current chrominance block. The method further includes deriving, based on the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block, a filter for a linear model or a non-linear model applied to the luminance component from candidate filters, and deriving a cross-component model corresponding to the filter. The method further includes filtering the reconstructed luminance samples in the corresponding luminance block by using filter coefficients according to the derived filter. The method further includes generating a prediction block of the current chrominance block from the filtered reconstructed luminance samples by using the cross-component model.
[0012] Another aspect of the present invention provides a method for a video encoding apparatus to encode a current chrominance block. The method includes establishing a corresponding luminance block of the current chrominance block from a reconstructed region of the luminance component. The corresponding luminance block represents a luminance block at the same position as the current chrominance block. The method further includes deriving, based on the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block, a filter for a linear model or a non-linear model applied to the luminance component from candidate filters, and deriving a cross-component model corresponding to the filter. The method further includes filtering the reconstructed luminance samples in the corresponding luminance block by using filter coefficients according to the derived filter. The method further includes generating a first prediction block of the current chrominance block from the filtered reconstructed luminance samples by using the cross-component model.
[0013] Another aspect of the present invention provides a computer-readable recording medium storing a bitstream generated by a video encoding method. The video encoding method includes establishing a corresponding luma block of a current chroma block from a reconstructed region of a luma component. The corresponding luma block represents a luma block co-located with the current chroma block. The video encoding method further includes deriving, based on the reconstructed luma reference samples of the corresponding luma block and the reconstructed reference samples of the current chroma block, a filter for a linear model or a non-linear model applied to the luma component from candidate filters, and deriving a cross-component model corresponding to the filter. The video encoding method further includes filtering the reconstructed luma samples in the corresponding luma block by using filter coefficients according to the derived filter. The video encoding method further includes generating a first prediction block of the current chroma block from the filtered reconstructed luma samples by using the cross-component model.
[0014] Advantageous Effects
[0015] As described above, the present invention provides a video encoding and decoding method and apparatus for deriving a filter for a luma component and a cross-component prediction model for cross-component prediction based on reconstructed reference samples when predicting a current chroma block. Therefore, the video encoding and decoding method and apparatus improve video encoding and decoding efficiency and enhance video quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a block diagram of a video encoding apparatus that can implement the technology of the present invention.
[0017] Figure 2 illustrates a method of partitioning blocks by using a quadtree plus binary tree plus ternary tree (QTBTTT) structure.
[0018] Figure 3a and Figure 3b illustrates a plurality of intra prediction modes including a wide-angle intra prediction mode.
[0019] Figure 4 illustrates adjacent blocks of a current block.
[0020] Figure 5 is a block diagram of a video decoding apparatus that can implement the technology of the present invention.
[0021] Figure 6 is a schematic diagram showing a gradient filter used in a gradient linear model (GLM).
[0022] Figure 7 is a block diagram detailing a part of a video decoding apparatus according to at least one embodiment of the present invention.
[0023] Figure 8 is a schematic diagram showing filtering of a luma component according to at least one embodiment of the present invention.
[0024] Figure 9 It is a schematic diagram showing the shape of a filter according to at least one embodiment of the present invention.
[0025] Figure 10 It is a schematic diagram showing filter coefficients according to at least one embodiment of the present invention.
[0026] Figure 11 It is a schematic diagram showing a reference sample area of a luminance component and a chrominance component according to at least one embodiment of the present invention.
[0027] Figure 12 It is a schematic diagram showing a gradient filter according to at least one embodiment of the present invention.
[0028] Figure 13a and Figure 13b It is a schematic diagram showing a prediction position for deriving parameters of a cross-component model according to at least one embodiment of the present invention.
[0029] Figure 14 It is a schematic diagram showing the position of a reference sample for deriving model parameters and / or calculating a prediction cost according to at least one embodiment of the present invention.
[0030] Figure 15 It is a flowchart of a method for encoding a current chrominance block by a video encoding device according to at least one embodiment of the present invention.
[0031] Figure 16 It is a flowchart of a method for reconstructing a current chrominance block by a video decoding device according to at least one embodiment of the present invention. Detailed Description
[0032] Hereinafter, some embodiments of the present invention will be described in detail with reference to the accompanying illustrative drawings. In the following description, the same reference numerals denote the same elements, although the elements are shown in different drawings. Further, in the following description of some embodiments, when the detailed description of related known components and functions is considered to obscure the subject matter of the present invention, the detailed description of the related known components and functions may be omitted for clarity and conciseness.
[0033] Figure 1 It is a block diagram of a video encoding device that can implement the technology of the present invention. Hereinafter, with reference to Figure 1 illustrations, the video encoding device and the components of the device will be described.
[0034] The encoding device may include: an image splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.
[0035] Each component of the encoding device may be implemented as hardware or software, or as a combination of hardware and software. Additionally, the functions of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.
[0036] A video consists of one or more sequences including multiple images. Each image is segmented into multiple regions, and encoding is performed on each region. For example, an image is segmented into one or more tiles and / or slices. Here, one or more tiles may be defined as a tile group. Each tile and / or slice is segmented into one or more coding tree units (CTUs). Additionally, each CTU is segmented into one or more coding units (CUs) through a tree structure. The information applied to each coding unit (CU) is encoded into the syntax of the CU, and the information commonly applied to the CUs included in one CTU is encoded into the syntax of the CTU. Additionally, the information commonly applied to all blocks in one slice is encoded into the syntax of the slice header, while the information applied to all blocks constituting one or more images is encoded into the Picture Parameter Set (PPS) or the picture header. Furthermore, the information commonly referred to by multiple images is encoded into the Sequence Parameter Set (SPS). Additionally, the information commonly referred to by one or more SPSs is encoded into the Video Parameter Set (VPS). Moreover, the information commonly applied to one tile or tile group may also be encoded into the syntax of the tile or tile group header. The syntax included in the SPS, PPS, slice header, tile, or tile group header may be referred to as high-level syntax.
[0037] The image splitter 110 determines the size of the coding tree unit (CTU). The information regarding the size (CTU size) of the CTU is encoded into the syntax of the SPS or PPS and is transmitted to the video decoding device.
[0038] The image splitter 110 segments each image constituting the video into multiple coding tree units (CTUs) of a predetermined size, and then recursively segments the CTUs by using a tree structure. The leaf nodes in the tree structure become coding units (CUs), and the CUs are the basic units for encoding.
[0039] The tree structure can be a quadtree (QT), where a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), where a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), where a higher node is divided into three lower nodes in a 1:2:1 ratio. The tree structure can also be a structure that mixes two or more of the QT structure, BT structure, and TT structure. For example, a quadtree plus binarytree (QTBT) structure can be used, or a quadtree plus binarytreeternarytree (QTBTTT) structure can be used. Here, the binarytreeternarytree (BTTT) is added to the tree structure to form a multiple-type tree (MTT).
[0040] Figure 2 is a schematic diagram for describing a method of dividing a block by using the QTBTTT structure.
[0041] As Figure 2 shown, the CTU can first be divided into a QT structure. The quadtree division can be recursive until the size of the divided block reaches the minimum block size (MinQTSize) of the leaf nodes allowed in the QT. The entropy encoder 155 encodes a first flag (QT_split_flag) indicating whether each node of the QT structure is divided into four lower nodes and signals it to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root nodes allowed in the BT, the leaf node can be further divided into at least one of the BT structure or the TT structure. There can be multiple division directions in the BT structure and / or TT structure. For example, there can be two directions, namely, the direction of horizontally dividing the block of the corresponding node and the direction of vertically dividing the block of the corresponding node. As Figure 2 shown, when the MTT division starts, the entropy encoder 155 encodes a second flag (mtt_split_flag) indicating whether the node is divided, and a flag additionally indicating the division direction (vertical or horizontal) and / or a flag indicating the division type (binary or ternary) in the case where the node is divided, and signals it to the video decoding device.
[0042] Alternatively, before encoding a first flag (QT_split_flag) indicating whether each node is split into four lower-layer nodes, a CU split flag (split_cu_flag) indicating whether a node is split may also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU, which is the basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding device starts encoding the first flag first according to the above scheme.
[0043] When QTBT is used as another example of a tree structure, there may be two types, that is, a type in which the block of the corresponding node is horizontally split into two blocks of the same size (i.e., symmetric horizontal split) and a type in which the block of the corresponding node is vertically split into two blocks of the same size (i.e., symmetric vertical split). The entropy encoder 155 encodes a split flag (split_flag) indicating whether each node of the BT structure is split into lower-layer blocks and split type information indicating the split type, and transmits it to the video decoding device. On the other hand, there may additionally be a type in which the block of the corresponding node is split into two asymmetric blocks. The asymmetric form may include a form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or may also include a form in which the block of the corresponding node is split in the diagonal direction.
[0044] A CU may have various sizes according to the QTBT or QTBTTT split from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block". When QTBTTT splitting is adopted, in addition to the square shape, the shape of the current block may also be rectangular.
[0045] The predictor 120 predicts the current block to generate a prediction block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.
[0046] Generally, each current block in the image can be predicted and encoded. Generally, the prediction of the current block can be performed by using an intra prediction technique (which uses data from the image including the current block) or an inter prediction technique (which uses data from the image encoded before the image including the current block). Inter prediction includes both uni-directional prediction and bi-directional prediction.
[0047] The intra predictor 122 predicts the pixels in the current block by using the pixels (reference pixels) adjacent to the current block in the current image including the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as Figure 3aAs shown, multiple intra-prediction modes may include two non-directional modes including Planar mode and DC mode, and may include 65 directional modes. Adjacent pixels to be used and algorithm equations are defined differently according to each prediction mode.
[0048] To perform efficient directional prediction on a current block having a rectangular shape, the directional modes shown by the dashed arrows in Figure 3b (modes #67 to #80, intra-prediction modes #-1 to #-14) may additionally be used. The directional modes may be referred to as "wide angle intra-prediction modes". In Figure 3b the arrows indicate the corresponding reference samples for prediction, rather than representing the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. When the current block has a rectangular shape, the wide angle intra-prediction mode is a mode that performs prediction in the direction opposite to a specific directional mode without additional bit transmission. In this case, in the wide angle intra-prediction mode, some wide angle intra-prediction modes available for the current block may be determined by the ratio of the width to the height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height less than the width, wide angle intra-prediction modes having an angle less than 45 degrees (intra-prediction modes #67 to #80) are available. When the current block has a rectangular shape with a width greater than the height, wide angle intra-prediction modes having an angle greater than -135 degrees are available.
[0049] The intra-predictor 122 may determine the intra-prediction to be used for encoding the current block. In some examples, the intra-predictor 122 may encode the current block by using multiple intra-prediction modes, and may also select an appropriate intra-prediction mode to be used from test modes. For example, the intra-predictor 122 may calculate rate-distortion values by performing rate-distortion analysis on multiple tested intra-prediction modes, and may also select an intra-prediction mode having the best rate-distortion characteristics from the test modes.
[0050] The intra-predictor 122 selects one intra-prediction mode from multiple intra-prediction modes, and predicts the current block by using adjacent pixels (reference pixels) and algorithm equations determined according to the selected intra-prediction mode. Information about the selected intra-prediction mode is encoded by the entropy encoder 155 and transmitted to the video decoding device.
[0051] The inter - frame predictor 124 generates a predicted block of the current block by using motion - compensation processing. The inter - frame predictor 124 searches for the block most similar to the current block in a reference image that has been encoded and decoded earlier than the current image, and generates a predicted block of the current block by using the searched - for block. Additionally, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current image and the predicted block in the reference image. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The entropy encoder 155 encodes the motion information including the information of the reference image and the information about the motion vector used for predicting the current block, and transmits it to the video decoding device.
[0052] The inter - frame predictor 124 can also perform interpolation of the reference image or reference block to increase the prediction accuracy. In other words, sub - samples are interpolated between two consecutive integer samples by applying filter coefficients to a plurality of consecutive integer samples including two integer samples. When performing the process of searching for the block most similar to the current block on the interpolated reference image, the motion vector can represent fractional - unit precision rather than integer - sample - unit precision. For each target region to be encoded, such as units like slices, tiles, CTUs, CUs, etc., the precision or resolution of the motion vector can be set differently. When applying such adaptive motion vector resolution (AMVR), information about the motion vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution can be information representing the precision of the motion - vector difference described below.
[0053] On the other hand, the inter-frame predictor 124 can perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference images and two motion vectors representing the positions of the blocks most similar to the current block in each reference image are used. The inter-frame predictor 124 selects a first reference image and a second reference image from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the blocks most similar to the current block in the corresponding reference images to generate a first reference block and a second reference block. In addition, a predicted block of the current block is generated by averaging or weighted averaging the first reference block and the second reference block. Further, motion information including information about the two reference images used for predicting the current block and information about the two motion vectors is transmitted to the entropy encoder 155. Here, reference picture list 0 may be composed of images in the pre-reconstructed images that are before the current image in the display order, and reference picture list 1 may be composed of images in the pre-reconstructed images that are after the current image in the display order. However, although not particularly limited thereto, pre-reconstructed images after the current image in the display order may be additionally included in reference picture list 0. Conversely, pre-reconstructed images before the current image may also be additionally included in reference picture list 1.
[0054] To minimize the amount of bits consumed for encoding motion information, various methods can be used.
[0055] For example, when the reference image and motion vector of the current block are the same as those of an adjacent block, information identifying the adjacent block is encoded to transmit the motion information of the current block to the video decoding device. This method is called the merge mode.
[0056] In the merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidates (hereinafter referred to as "merge candidates") from the adjacent blocks of the current block.
[0057] As the adjacent blocks for deriving the merge candidates, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image can be used, as Figure 4 shown. In addition, in addition to the current image where the current block is located, blocks within the reference image (which may be the same as or different from the reference image used for predicting the current block) can also be used as merge candidates. For example, a co-located block of the current block within the reference image or a block adjacent to the co-located block can be additionally used as a merge candidate. If the number of merge candidates selected by the above method is less than the preset number, zero vectors are added to the merge candidates.
[0058] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates by using adjacent blocks. A merge candidate to be used as motion information of the current block is selected from among the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding device.
[0059] Merge skip mode is a special case of merge mode. After quantization, when all transform coefficients for entropy coding are close to zero, only the neighbor block selection information is transmitted without the residual signal. By using merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, screen content images, etc.
[0060] Hereinafter, merge mode and merge skip mode are collectively referred to as merge / skip mode.
[0061] Another method for encoding motion information is the advanced motion vector prediction (AMVP) mode.
[0062] In the AMVP mode, the inter-frame predictor 124 derives a motion vector prediction candidate for a motion vector of a current block by using neighboring blocks of the current block. As neighboring blocks for deriving motion vector prediction candidates, the neighboring blocks may be used. Figure 4 All or some of the left block A0, lower left block A1, upper block B0, upper right block B1 and upper left block B2 adjacent to the current block in the current image shown. In addition, in addition to the current image where the current block is located, blocks located in a reference image (which may be the same as or different from the reference image used to predict the current block) may also be used as adjacent blocks for deriving motion vector prediction candidates. For example, a co-located block of the current block in the reference image or a block adjacent to the co-located block may be used. If the number of motion vector candidates selected by the above method is less than the preset number, a zero vector is added to the motion vector candidates.
[0063] The inter-frame predictor 124 derives a motion vector prediction candidate by using the motion vectors of the neighboring blocks, and determines a motion vector prediction of the motion vector of the current block by using the motion vector prediction candidate. In addition, a motion vector difference is calculated by subtracting the motion vector prediction from the motion vector of the current block.
[0064] Motion vector prediction can be obtained by applying a predefined function (e.g., median and average calculations, etc.) to motion vector prediction candidates. In this case, the video decoding device also knows the predefined function. In addition, since the neighboring blocks used to derive the motion vector prediction candidates are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information used to identify the motion vector prediction candidates. Accordingly, in this case, the information about the motion vector difference and the information about the reference image used to predict the current block are encoded.
[0065] On the other hand, motion vector prediction can also be determined by a scheme of selecting any one of the motion vector prediction candidates. In this case, the information used to identify the selected motion vector prediction candidate is additionally encoded together with the information about the motion vector difference and the information about the reference image used to predict the current block.
[0066] The subtractor 130 generates a residual block by subtracting the prediction block generated by the intra predictor 122 or the inter predictor 124 from the current block.
[0067] The transformer 140 transforms the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 can transform the residual signal in the residual block by using the entire size of the residual block as the transform unit, or the residual block can also be divided into multiple sub-blocks, and the transform can be performed by using the sub-blocks as the transform units. Alternatively, the residual block is divided into two sub-blocks, namely a transform region and a non-transform region, to transform the residual signal by using only the transform region sub-block as the transform unit. Here, the transform region sub-block can be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or the vertical axis). In this case, the entropy encoder 155 encodes a flag (cu_sbt_flag) indicating only the transform sub-block, and the direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag), and signals them to the video decoding device. Additionally, the size of the transform region sub-block can have a size ratio of 1:3 based on the horizontal axis (or the vertical axis). In this case, the entropy encoder 155 additionally encodes a flag (cu_sbt_quad_flag) for dividing the corresponding segmentation, and signals it to the video decoding device.
[0068] On the other hand, the transformer 140 can perform the transformation of the residual block separately in the horizontal direction and the vertical direction. For this transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a multiple transform set (MTS). The transformer 140 can select a pair of transformation functions in the MTS with the highest transformation efficiency, and can transform the residual block on each of the horizontal direction and the vertical direction. The entropy encoder 155 encodes the information (mts_idx) about the pair of transformation functions in the MTS and signals it to the video decoding device.
[0069] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using quantization parameters and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also quantize the relevant residual block immediately without transforming any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions can be encoded and signaled to the video decoding device.
[0070] The rearrangement unit 150 can perform rearrangement of the coefficient values on the quantized residual values.
[0071] The rearrangement unit 150 can change the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can scan the coefficients from the DC coefficient to the high-frequency region using a zig-zag scan or a diagonal scan to output a 1D coefficient sequence. According to the size of the transform unit and the intra prediction mode, a vertical scan that scans the 2D coefficient array in the column direction and a horizontal scan that scans the 2D block type coefficients in the row direction can also be used instead of the zig-zag scan. In other words, according to the size of the transform unit and the intra prediction mode, the scan method to be used can be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.
[0072] The entropy encoder 155 encodes the sequence of 1D quantized transform coefficients output from the rearrangement unit 150 by using various coding schemes including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc. to generate a bitstream.
[0073] In addition, the entropy encoder 155 encodes information related to block partitioning (e.g., CTU size, CTU partitioning flag, QT partitioning flag, MTT partitioning type, and MTT partitioning direction, etc.) so that the video decoding device can partition blocks in the same way as the video encoding device. In addition, the entropy encoder 155 encodes information about the prediction type indicating whether the current block is encoded by intra prediction or inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information about the intra prediction mode) or inter prediction information (merge index in the case of the merge mode, and information about the reference image index and motion vector difference in the case of the AMVP mode) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information about the quantization parameter and information about the quantization matrix).
[0074] The inverse quantizer 160 inverse quantizes the quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct the residual block.
[0075] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. When performing intra prediction on the next block, the pixels in the reconstructed current block are used as reference pixels.
[0076] The loop filter unit 180 performs filtering on the reconstructed pixels to reduce block artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transform / quantization. The loop filter unit 180, as an in-loop filter, may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.
[0077] The deblocking filter 182 filters the boundaries between the reconstructed blocks to remove blocking artifacts that occur due to block-based coding / decoding, and the SAO filter 184 and the ALF 186 perform additional filtering on the deblocked video. The SAO filter 184 and the ALF 186 are filters for compensating the difference between the reconstructed pixels and the original pixels that occur due to lossy coding. The SAO filter 184 applies an offset in units of CTUs to enhance the subjective image quality and coding efficiency. On the other hand, the ALF 186 performs block-based filtering and applies different filters by dividing the boundaries and the degree of variation of the corresponding blocks to compensate for distortion. Information about the filter coefficients to be used for the ALF can be encoded and signaled to the video decoding device.
[0078] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter prediction of the blocks within the image to be encoded subsequently.
[0079] The video encoding device can store the bitstream of the encoded video data in a non-volatile storage medium or transmit the bitstream to the video decoding device through a communication network.
[0080] Figure 5 is a functional block diagram of a video decoding device that can implement the technology of the present invention. Hereinafter, with reference to Figure 5 , the video decoding device and the components of the device are described.
[0081] The video decoding device may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transform unit 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.
[0082] Similar to Figure 1 the video encoding device, each component of the video decoding device can be implemented as hardware or software, or implemented as a combination of hardware and software. In addition, the functions of each component can be implemented as software, and the microprocessor can also be implemented to execute the functions of the software corresponding to each component.
[0083] The entropy decoder 510 extracts information related to block partitioning by decoding the bitstream generated by the video encoding device to determine the current block to be decoded, and extracts the prediction information and the information about the residual signal required for reconstructing the current block.
[0084] The entropy decoder 510 determines the size of a coding tree unit (CTU) by extracting information about the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS), and divides an image into CTUs having the determined size. In addition, a CTU is determined as the highest layer (i.e., the root node) of a tree structure, and the segmentation information of the CTU can be extracted to divide the CTU by using the tree structure.
[0085] For example, when dividing a CTU by using a QTBTTT structure, a first flag (QT_split_flag) related to the segmentation of a quad tree (QT) is first extracted to divide each node into four lower-layer nodes. In addition, a second flag (mtt_split_flag) related to the segmentation of a multi-type tree (MTT), a segmentation direction (vertical / horizontal), and / or a segmentation type (binary / trinary) are extracted with respect to a node corresponding to a leaf node of the QT to divide the corresponding leaf node into an MTT structure. As a result, each node below the leaf node of the QT is recursively divided into a binary tree (BT) or a ternary tree (TT) structure.
[0086] As another example, when dividing a CTU by using a QTBTTT structure, a CU segmentation flag (split_cu_flag) indicating whether to divide a coding unit (CU) is extracted. When dividing the corresponding block, the first flag (QT_split_flag) may also be extracted. During the division process, for each node, zero or more recursive MTT divisions may occur after zero or more recursive QT divisions. For example, for a CTU, an MTT division may occur immediately, or conversely, only multiple QT divisions may occur.
[0087] As another example, when dividing a CTU by using a QTBT structure, a first flag (QT_split_flag) related to the segmentation of a quad tree (QT) is extracted to divide each node into four lower-layer nodes. In addition, a segmentation flag (split_flag) indicating whether to further divide the node corresponding to the leaf node of the QT into a binary tree (BT) and segmentation direction information are extracted.
[0088] On the other hand, when the entropy decoder 510 determines a current block to be decoded by using the segmentation of a tree structure, the entropy decoder 510 extracts information about a prediction type indicating whether the current block is intra-frame predicted or inter-frame predicted. When the prediction type information indicates intra-frame prediction, the entropy decoder 510 extracts a syntax element for intra-frame prediction information (intra-frame prediction mode) of the current block. When the prediction type information indicates inter-frame prediction, the entropy decoder 510 extracts information about syntax elements representing inter-frame prediction information, i.e., a motion vector and a reference image to which the motion vector refers.
[0089] In addition, the entropy decoder 510 extracts quantization-related information and extracts information on the quantized transform coefficients of the current block as information on the residual signal.
[0090] The rearrangement unit 515 can change the sequence of 1D quantized transform coefficients entropy decoded by the entropy decoder 510 back into a 2D coefficient array (i.e., a block) in the reverse order of the coefficient scan order performed by the video coding device.
[0091] The inverse quantizer 520 inverse quantizes the quantized transform coefficients and inverse quantizes the quantized transform coefficients by using the quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video coding device to the 2D array of quantized transform coefficients.
[0092] The inverse transformer 530 reconstructs the residual signal by inverse-transforming the inverse quantized transform coefficients from the frequency domain to the spatial domain to generate a residual block of the current block.
[0093] In addition, when the inverse transformer 530 inverse-transforms a partial region (sub-block) of the transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) for inverse-transforming only the sub-block of the transform block, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the non-inverse-transformed region with the value "0" as the residual signal to generate the final residual block of the current block.
[0094] In addition, when applying MTS, the inverse transformer 530 determines the transform index or transform matrix to be applied in each of the horizontal and vertical directions by using the MTS information (mts_idx) signaled from the video coding device. The inverse transformer 530 also performs inverse transformation on the transform coefficients in the transform block in the horizontal and vertical directions by using the determined transform function.
[0095] The predictor 540 can include an intra predictor 542 and an inter predictor 544. When the prediction type of the current block is intra prediction, the intra predictor 542 is activated, and when the prediction type of the current block is inter prediction, the inter predictor 544 is activated.
[0096] The intra predictor 542 determines the intra prediction mode of the current block among multiple intra prediction modes according to the syntax element of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block according to the intra prediction mode by using the adjacent reference pixels of the current block.
[0097] The inter - frame predictor 544 determines the motion vector of the current block and the reference image for motion - vector reference by using the syntax element of the inter - frame prediction mode extracted from the entropy decoder 510.
[0098] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 to the prediction block output from the inter - frame predictor 544 or the intra - frame predictor 542. When performing intra - frame prediction on a block to be decoded subsequently, the pixels within the reconstructed current block are used as reference pixels.
[0099] The loop filter unit 560 as an in - loop filter may include a de - blocking filter 562, a SAO filter 564, and an ALF 566. The de - blocking filter 562 performs de - blocking filtering on the boundaries between the reconstructed blocks to remove block artifacts that occur due to block - based unit decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after de - blocking filtering to compensate for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The filter coefficients of the ALF are determined by using the information about the filter coefficients decoded from the bitstream.
[0100] The reconstructed blocks filtered by the de - blocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter - frame prediction of the blocks within the image to be encoded subsequently.
[0101] In some embodiments, the present invention relates to encoding and decoding video images as described above. More specifically, the present invention provides a video encoding and decoding method and apparatus for deriving a filter for the luminance component and a cross - component prediction model for cross - component prediction based on reconstructed reference samples when predicting a current chrominance block according to cross - component prediction.
[0102] The following embodiments may be performed by the intra - frame predictor 122 in a video encoding apparatus. The following embodiments may also be performed by the intra - frame predictor 542 in a video decoding apparatus.
[0103] When encoding a current block, the video encoding apparatus may generate signaling information associated with this embodiment from the perspective of optimizing rate - distortion. The video encoding apparatus may encode the signaling information using the entropy encoder 155 and send the encoded signaling information to the video decoding apparatus. The video decoding apparatus may decode the signaling information associated with the decoding of the current block from the bitstream using the entropy decoder 510.
[0104] In the following description, the term "target block" may be used interchangeably with the current block or coding unit (CU), or may refer to some regions of the coding unit.
[0105] In addition, a value of true for a flag indicates the case where the flag is set to 1. Additionally, a value of false for a flag indicates the case where the flag is set to 0.
[0106] I. Intra Prediction of Chrominance Components
[0107] Introduce several techniques for improving coding and decoding efficiency by utilizing intra prediction.
[0108] In the Versatile Video Coding (VVC) technology, a luminance block has 65 subdivided angular intra prediction modes or direction modes (i.e., modes 2 to mode 66) in addition to non - angular or non - directional modes (i.e., planar and DC modes), as Figure 3a shown. The 65 direction modes, planar, and DC modes are collectively referred to as 67 IPMs.
[0109] According to the prediction direction utilized by the luminance block, the chrominance block can also have limited access to the intra prediction of these subdivided direction modes. However, except for the horizontal and vertical directions, the intra prediction of the chrominance block may not always utilize the various direction modes available to the luminance block. To be able to use these different direction modes, the prediction mode of the current chrominance block needs to be set to the direct mode (DM). By setting the prediction mode to DM, the current chrominance block can utilize the direction modes other than the horizontal and vertical modes of the luminance block.
[0110] When the chrominance block is encoded, the most commonly used or default intra prediction modes that maintain video quality include planar, DC, vertical, horizontal modes, and DM. Under DM, the intra prediction mode of the luminance block that is spatially corresponding to the current chrominance block is used as the intra prediction mode of the chrominance block.
[0111] The video coding device can signal to the video decoding device whether the intra prediction mode of the chrominance block is DM. The video coding device can transmit DM to the video decoding device in many ways. For example, the video coding device can indicate whether the intra prediction mode of the chrominance block is DM by setting intra_chroma_pred_mode, which is the information for indicating the intra prediction mode of the chrominance block, to a specific value and sending this information to the video decoding device.
[0112] When encoding the chrominance block in the intra prediction mode, the video coding device can set the intra prediction mode IntraPredModeC of the chrominance block according to Table 1.
[0113] Hereinafter, in order to distinguish intra_chroma_pred_mode and IntraPredModeC, which are information related to the intra prediction mode of a chroma block, intra_chroma_pred_mode and IntraPredModeC are respectively referred to as a chroma intra prediction mode indicator and a chroma intra prediction mode.
[0114] [Table 1]
[0115]
[0116] Here, lumaIntraPredMode is an intra prediction mode (hereinafter, “luma intra prediction mode”) of a luma block corresponding to a current chroma block. lumaIntraPredMode represents Figure 3a one of the prediction modes shown. For example, in Table 1, lumaIntraPredMode = 0 indicates a planar prediction mode, and lumaIntraPredMode = 1 indicates a DC prediction mode. The cases where lumaIntraPredMode is 18, 50, and 66 respectively indicate direction modes called horizontal, vertical, and VDIA. On the other hand, the cases where intra_chroma_pred_mode = 0, 1, 2, and 3 respectively indicate planar, vertical, horizontal, and DC prediction modes. The case where intra_chroma_pred_mode = 4 is DM, in which the value of the chroma intra prediction mode IntraPredModeC is set to be equal to the value of lumaIntraPredMode.
[0117] On the other hand, the process of analyzing the intra prediction mode of a chroma block performed by a video decoding device is shown in Table 2.
[0118] [Table 2]
[0119] if(CclmEnabled) cclm_mode_flag if(cclm_mode_flag) cclm_mode_idx else intra_chroma_pred_mode
[0120] The video decoding device analyzes cclm_mode_flag, which indicates whether to use a cross-component linear model (CCLM) mode. If cclm_mode_flag is 1 to enable the CCLM mode, the video decoding device analyzes cclm_mode_idx indicating the CCLM mode. According to the value of cclm_mode_idx, the CCLM mode can indicate one of three modes. On the other hand, if cclm_mode_flag is 0 (indicating that the CCLM mode is not used), the video decoding device analyzes intra_chroma_pred_mode indicating the intra prediction mode, as described above.
[0121] If the CCLM mode is applied to the intra prediction of the current chrominance block, the video decoding device determines a corresponding region in the luma image corresponding to the current chrominance block (hereinafter referred to as the "corresponding luma region"). To predict the linear model, the left reference pixel and the upper reference pixel of the corresponding luma region, as well as the left reference pixel and the upper reference pixel of the target chrominance block, can be utilized. Hereinafter, the left reference pixel and the upper reference pixel are collectively referred to as reference pixels, neighboring pixels, or adjacent pixels. Additionally, the reference pixels in the chrominance channel are referred to as chrominance reference pixels, and the reference pixels in the luma channel are referred to as luma reference pixels.
[0122] In CCLM prediction, a linear model between the reference pixels in the corresponding luma region and the reference pixels of the chrominance block is derived, and then the linear model is applied to the reconstructed pixels in the corresponding luma region to generate a prediction block as a predictor for the target chrominance block. For example, four pairs of pixels (i.e., combining the pixels in the neighboring pixel line of the current chrominance block and the pixels in the corresponding luma region) can be used to derive the linear model. In response to the four pairs of pixels, the video decoding device can derive α and β representing the linear model, as shown in Equation 1.
[0123] [Equation 1]
[0124]
[0125] Here, X a and X b each represent the average of the minimum and second minimum values and the average of the maximum and second maximum values of the corresponding luma pixels in the four pairs of pixels. In addition, Y a and Y b each represent the average of the minimum and second minimum values and the average of the maximum and second maximum values of the chrominance pixels in the four pairs of pixels. Then, the video decoding device can use the linear model to generate a predictor pred L (i,j) for the current chrominance block from the pixel value rec' C (i,j) of the corresponding luma region, as shown in Equation 2.
[0126] [Equation 2]
[0127] pred C (i,j) = α · rec′ L (i,j) + β
[0128] As described above, according to the positions of adjacent pixels used in deriving the linear model, the CCLM mode is divided into three modes: CCLM_LT, CCLM_L, and CCLM_T. The CCLM_LT mode uses two pixels in each direction of adjacent pixels adjacent to the left and top sides of the current chrominance block. The CCLM_L mode uses four pixels of adjacent pixels adjacent to the left side of the current chrominance block. Finally, the CCLM_T mode utilizes four pixels of adjacent pixels adjacent to the top side of the current chrominance block.
[0129] The GLM mode (Gradient Linear Model mode) applies a linear model to the gradient G(i,j) of the reconstructed luminance component to predict the chrominance component, as shown in Equation 3.
[0130] [Equation 3]
[0131] pred C (i,j) = α·G(i,j) + β
[0132] This process can utilize the same linear model as that used in CCLM (Cross-Component Linear Model). The gradient G(i,j) can be generated by applying a gradient filter to the luminance samples co-located with the current chrominance sample. Sixteen types of gradient filters are defined, and Figure 6 the examples in show some of the gradient filters. The video coding device can signal the video decoding device an index indicating the type of the gradient filter.
[0133] As described above, CCLM directly applies a linear model to the luminance component to predict the chrominance component. Additionally, GLM predicts the chrominance component by applying a linear model to the filtered luminance component. Since GLM uses a linear model derived in the same manner as CCLM, GLM does not consider the filter when deriving the cross-component model. Moreover, it also needs to signal an index indicating the type of the filter. To solve the above problems, the embodiments of the present invention implicitly derive the filter type and apply the filter to the luminance component during the derivation of the cross-component model.
[0134] The following embodiments are described around the video decoding device, but can also be implemented in the same or similar manner in the video coding device.
[0135] II. Embodiments According to the Present Invention
[0136] Figure 7 is a block diagram of a part detailing a video decoding device according to at least one embodiment of the present invention.
[0137] The video decoding device according to the present embodiment can decode the syntax elements required for prediction and generate a reconstructed block of the current chrominance block based on the decoded syntax. Figure 7The operations shown in can be performed by the entropy decoder 510 and the predictor 540 of the video decoding device. On the other hand, similar to Figure 7 The operations shown in can be performed by the predictor 120 of the video encoding device. In this case, the video decoding device uses the encoded information parsed from the bitstream, while the video encoding device can use the encoded information preset at a higher level in terms of minimizing rate distortion. Hereinafter, for convenience of description, the present embodiment will be described centering on the video decoding device.
[0138] As Figure 5 shown, the predictor 540 includes an intra predictor 542 and an inter predictor 544 according to the prediction technique. However, for the intra prediction of the chrominance component, the intra predictor 544 may include all or part of an implicit filter derivator 710, a luminance component filter applicator 720, and a prediction executor 730, as Figure 7 shown.
[0139] When the color format of the input video is the YUV format (YUV420, YUV411, YUV422, YUV444, etc.), the video decoding device can perform the prediction and reconstruction of the luminance component, and then can perform the prediction and reconstruction of the chrominance component. On the other hand, when the color format of the input video is RGB, the video encoding device can perform the color format conversion from RGB to YUV, and then can encode the converted video. Here, in the case of the YUV format, the color format represents the association between the pixels in the luminance component and the pixels in the chrominance component.
[0140] The entropy decoder 510 decodes the chrominance intra prediction mode indicator. The chrominance intra prediction mode indicator indicates the prediction mode of the current chrominance block for intra prediction.
[0141] When the chrominance intra prediction mode indicator indicates the cross-component prediction mode, the entropy decoder 510 decodes a flag indicating whether to derive an implicit filter (hereinafter referred to as "implicit filter derivation flag"). On the other hand, the cross-component prediction mode can be applied using the cross-component prediction model. According to the decoded cross-component prediction mode, a linear model or a non-linear model can be derived as the cross-component prediction model. Hereinafter, the cross-component prediction model can be used interchangeably with the cross-component model.
[0142] When the chrominance intra prediction mode indicator does not indicate the cross-component prediction mode, the video decoding device can generate a prediction block of the current chrominance block according to the corresponding chrominance prediction mode.
[0143] At this time, when an implicit filter derivation flag is received from a higher level such as a sequence, a slice, etc., the flag at the CU level may be implicitly derived as the same value as the flag received from the higher level. Alternatively, if the flag received from the higher level is 1, the entropy decoder 510 may further decode the flag at the CU level to determine whether to derive an implicit filter.
[0144] As an example, if the color format is YUV444, the implicit filter derivation flag may be derived as 0.
[0145] If the implicit filter derivation flag is true, the implicit filter deriver 710 derives the type of filter (or form of filter) to be applied to the luminance component of the cross-component prediction. In the process of deriving the filter type, the implicit filter deriver 710 can derive the best cross-component model consistent with the derived filter type. In this case, the filter type applied to the luminance component can include the gradient filter used in the above-mentioned GLM technology.
[0146] On the other hand, if the implicit filter derivation flag is false, the implicit filter deriver 710 does not derive a filter, and the video decoding apparatus may perform cross-component prediction according to a conventional method such as CCLM. For example, when CCLM is applied, the intra predictor 542 may derive a linear model suitable for CCLM as described above.
[0147] The luma component filter applier 720 performs filtering on the reconstructed luma samples by applying the filter coefficient according to the derived filter type to a reconstructed luma block co-located with the current chroma block (hereinafter, referred to as a 'corresponding luma block').
[0148] Furthermore, if the implicit filter derivation flag is derived as false, the operation of the luma component filter applier 720 may be omitted.
[0149] The prediction performer 730 uses the derived optimal cross-component model to generate a prediction block for the current chrominance block from the filtered reconstructed luma samples of the corresponding luma block.
[0150] On the other hand, if the implicit filter derivation flag is false, the prediction executor 730 may utilize the unfiltered reconstructed luma samples of the cross component prediction. That is, the prediction executor 730 may input the unfiltered reconstructed luma samples into the derived linear model to generate a prediction block for the current chroma block.
[0151] The operation of the implicit filter deriver 710 is described in detail below. The implicit filter deriver 710 may include all or part of the luma component filtering unit 712, the cross-component model estimator 714, and the prediction cost calculator 716.
[0152] On the other hand, if the implicit filter derivation flag is false, the operation of the implicit filter deriver 710 can be omitted.
[0153] The luma component filtering unit 712 uses all candidate filters to filter the corresponding luma block and luma reference sample region, as Figure 8 shown. The luma component filtering unit 712 generates a filtered corresponding luma block from the corresponding luma block, and generates a filtered reconstructed luma reference sample from the reconstructed luma reference sample region or "reconstructed luma reference samples". In this case, the candidate filters can be predefined based on the protocol between the video encoding device and the video decoding device.
[0154] The number of candidate filters is one or more. For example, a single candidate filter generates a filtered corresponding luma block and a filtered reconstructed luma reference sample region.
[0155] To derive the filter, the implicit filter deriver 710 utilizes the filtered reconstructed luma reference sample region to allow the luma component filtering unit 712 to omit the filtering of the corresponding luma block.
[0156] As Figure 8 shown in one example, the reconstructed reference sample regions represented by A, B, C, and D can vary. For example, A = W or A = 2W, and B = H or B = 2H. Figure 8 The example of Figure 8 shows the luma component filtering when the color format is YUV420. In the example of
[0157] Figure 9
[0158] Figure 9 The shape of the candidate filters and the corresponding filter coefficients can vary. For example, a rectangular or square filter with F W ×F H filter coefficients and / or an F×F diamond-shaped filter with F W ×F H filter coefficients can be used as candidates. In the example of Figure 9 shown, a filter with F W ×F HRectangular or square-shaped filters of filter coefficients, 3×3 rhombus filters, and 5×5 rhombus filters. According to an embodiment, a specific shape of a filter defined by a protocol between a video encoding device and a video decoding device may be used. In addition, for each filter shape, there may be multiple candidate filters with different filter coefficients.
[0159] For example, F W ×F H The candidate filters in the form may be an edge filter for extracting a specific directionality, a downsampling filter for extracting a specific pixel value, and / or a subsampling filter. Figure 10 The example in W shows some available filter coefficients when F H = 3 and F
[0160] As described above, if the color format is YUV444, the implicit filter derivation flag may be derived as 0. Alternatively, if the color format is YUV444, the downsampling filter or the subsampling filter may be excluded from the candidate filters.
[0161] As Figure 11 shown, the cross-component model estimator 714 derives the parameters of the cross-component model based on the reconstructed pixel values of the filtered reference sample region of the corresponding luminance block and the reconstructed pixel values of the reference sample region of the current chrominance block. This may derive the same number of candidate cross-component models as the candidate filters. For example, a single candidate filter results in a single candidate cross-component model.
[0162] As described above, the cross-component model derived according to the decoded cross-component prediction mode of the current chrominance block may be a linear model or a non-linear model.
[0163] As an example, in the case of a linear model, the cross-component model estimator 714 may estimate the parameters α and β of the linear model by using the linear regression equation shown in Equation 4.
[0164] [Equation 4]
[0165]
[0166] Here, L(n) represents the reconstructed pixel values of the filtered reference sample region of the corresponding luminance block, and C(n) represents the reconstructed pixel values of the reference sample region of the current chrominance block. N indicates the number of sample pairs used in the linear regression.
[0167] Alternatively, the cross-component model estimator 714 can estimate the parameters α and β of the linear model as shown in Equation 5.
[0168] [Equation 5]
[0169]
[0170] Here, x max and x min represent the average of the maximum and the second maximum values and the average of the minimum and the second minimum values of the corresponding filtered luminance reference pixels, respectively. In addition, y max and y min represent the average of the maximum and the second maximum values and the average of the minimum and the second minimum values of the corresponding filtered chrominance reference pixels, respectively.
[0171] Based on the linear model estimated as described above, the chrominance component can be predicted according to Equation 6.
[0172] [Equation 6]
[0173] pred C (i,j) = α·L′(i,j) + β
[0174] Here, L′(i,j) represents the filtered reconstructed luminance sample at position (i,j).
[0175] As another example, in the case of a non-linear model, the cross-component model estimator 714 can derive the model parameters representing the non-linear model based on the reference samples of the reconstructed luminance reference region and the reconstructed reference samples of the current chrominance block. By using the derived model parameters, the chrominance component can be predicted according to Equation 7. In this case, a subset model that uses only some of the model parameters can be used.
[0176] [Equation 7]
[0177] pred C (i,j) = c0L′(i,j) + c1i + c2j + c3P(i,j) + c4B + c5G x (i,j) + c6G y (i,j)
[0178] Here, c i represents the non-linear model parameters. L′(i,j) represents the filtered reconstructed luminance sample at position (i,j). The non-linear term P and the bias term B can be calculated as shown in Equation 8.
[0179] [Equation 8]
[0180] P(i,j) = (L'(i,j)·L'(i,j) + (1 << (bit depth - 1))) >> bit depth
[0181] B = 1 << (bit depth - 1)
[0182] In addition, G can be calculated by applying a gradient filter of size k x × k y to the corresponding filtered or unfiltered luminance sample values along the x and y axes, respectively. In this case, the gradient filter can be set according to the protocol between the video encoding device and the video decoding device. For example, when k x = 3 and k y = 3, the filters in the x-axis and y-axis directions for extracting the gradient can be represented as shown in the illustration of x y Figure 12
[0183] In this case, by using the 3×3 gradient filters E x and E y , the values of G x and G y can be calculated as shown in Equation 9 and Equation 10.
[0184] [Equation 9]
[0185] G x (i,j) = L(2i - 1, 2j - 1)E x (0, 0) + L(2i, 2j - 1)E x (1, 0)
[0186] + L(2i + 1, 2j - 1)E x (2, 0) + L(2i - 1, 2j)E x (0, 1) + L(2i, 2j)E x (1, 1)
[0187] + L(2i + 1, 2j)E x (2, 1) + L(2i - 1, 2j + 1)E x (0, 2)
[0188] + L(2i, 2j + 1)E x (1, 2) + L(2i + 1, 2j + 1)E x (2, 2)
[0189] [Equation 10]
[0190] G y (i,j) = L(2i - 1, 2j - 1)E y (0,0) + L(2i, 2j - 1)E y (1,0)
[0191] + L(2i + 1, 2j - 1)E y (2,0) + L(2i - 1, 2j)E y (0,1) + L(2i, 2j)E y (1,1)
[0192] + L(2i + 1, 2j)E y (2,1) + L(2i - 1, 2j + 1)E y (0,2)
[0193] + L(2i, 2j + 1)E y (1,2) + L(2i + 1, 2j + 1)E y (2,2)
[0194] Here, L′(i,j) represents the corresponding unfiltered luminance sample at position (i,j).
[0195] On the other hand, when the prediction value S is defined as S (S ∈ {L', i, j, P, B, Gx, Gy}), for the number N of pixels used for parameter derivation, the cross-component model estimator 714 can calculate the autocorrelation matrix AC of S and the cross-correlation between S and the reconstructed chrominance reference samples. The cross-component model estimator 714 can estimate the parameter vector that minimizes Equation 11 based on the estimated autocorrelation and cross-correlation
[0196] [Equation 11]
[0197]
[0198] Here, a method such as LDL decomposition can be used to estimate the parameter vector
[0199] Alternatively, the cross-component model estimator 714 can use the reference samples of the reconstructed luminance reference region to generate as many linear equations as the number of parameters according to Equation 7. In this case, the prediction value on the left side of Equation 7 can be replaced by the reconstructed chrominance reference samples. The cross-component model estimator 714 can derive the parameters of the non-linear model by solving the linear equations using Gaussian elimination.
[0200] As another example, by using different non-linear model parameters, the chrominance component can be predicted according to Equation 12.
[0201] [Equation 12]
[0202] pred C(i,j) = c0L′(i - 1,j - 1) + c1L ′ (i,j - 1) + c2L ′ (i + 1,j - 1) + c3L ′ (i - 1,j) + c4b ′ (i,j) + c5L ′ (i + 1,j) + c6L ′ (i - 1,j + 1) + c7L ′ (i,j + 1) + c8L ′ (i + 1,j + 1) + c9P + c 10 B
[0203] In this case, a subset model that uses only some of the model parameters can be used. Additionally, the position information (i,j) of the currently predicted pixel, gradients (Gx and Gy), etc. can be added as inputs to the prediction model, and parameters for the added inputs can be added.
[0204] The prediction cost calculator 716 inputs N reference samples of the filtered and reconstructed reference sample region into the cross-component model to perform chrominance component prediction, thereby generating a predicted value, and then calculates the deviation between the predicted value and the actually reconstructed chrominance reference sample, i.e., the prediction cost. The prediction cost calculator 716 can apply this error calculation process to all candidate filters and the corresponding candidate cross-component models, and then can determine the candidate filter and candidate cross-component model that generate the minimum prediction cost as the filter and cross-component model for predicting the current chrominance block.
[0205] On the other hand, if there is a single candidate filter, which means there is a single candidate cross-component model, the operation of the prediction cost calculator 716 can be omitted.
[0206] Figure 13a and Figure 13b are schematic diagrams showing the prediction positions for deriving the parameters of the cross-component model according to some embodiments of the present invention.
[0207] As another example, the implicit filter derivator 710 can derive the type of filter to be applied to the luminance component and the prediction positions for deriving the parameters of the cross-component model.
[0208] Regarding the prediction positions, the AL mode that uses the upper and left reference samples of the current chrominance block, the A mode that uses the upper reference sample of the current chrominance block, and the L mode that uses the left reference sample of the current chrominance block are defined, as Figure 13a and Figure 13bAs shown. Hereinafter, the AL mode, A mode and L mode associated with the predicted position are collectively referred to as position modes. When M is the number of candidate filters, the implicit filter deriver 710 derives parameters of the cross-component model for M×3 combinations of candidate filters and candidate position modes. Therefore, M×3 cross-component filters can be derived for M×3 combinations. The implicit filter deriver 710 calculates the prediction cost based on the derived model parameters. The implicit filter deriver 710 can determine the candidate filter, candidate position mode and candidate cross-component model with the minimum cost as the filter, position mode and cross-component model of the current chroma block.
[0209] At this time, the location of the reference samples used to derive the model parameters can be different from the location of the reference samples used to calculate the prediction cost, such as Figure 14 shown.
[0210] In addition, when there is a single candidate filter, that is, when the candidate cross-component models are equal in number to the candidate position patterns, the implicit filter deriver 710 calculates the prediction cost according to the candidate position patterns. The implicit filter deriver 710 may determine the candidate position pattern and the candidate cross-component model with the minimum cost as the position pattern and the cross-component model of the current chrominance block.
[0211] On the other hand, if the position of the reference sample for implicitly deriving the parameters of the cross-component model is derived, the derivation of the implicit filter can be omitted. The video decoding device can omit the parsing of the implicit filter derivation flag and the derivation of the implicit filter because it uses the filter determined at a higher level. The video decoding device can determine whether to derive the implicit filter and whether to derive the implicit parameter estimation position by decoding the implicit filter derivation flag and the implicit parameter estimation position derivation flag, respectively.
[0212] Reference now Figure 15 and Figure 16 , describes a method for predicting a current chrominance block. In the following, it is assumed that the current chrominance block is predicted based on cross-component prediction.
[0213] Figure 15 is a flowchart of a method for encoding a current chrominance block by a video encoding apparatus according to at least one embodiment of the present invention.
[0214] The video encoding apparatus establishes a luma region corresponding to a current chroma block from a reconstructed region of a luma component (S1500). Here, the corresponding luma block refers to a luma block co-located with the current chroma block.
[0215] The video encoding device derives a filter to be applied to the luminance component in a candidate filter based on the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block, and derives a cross-component model corresponding to the derived filter (S1502). Here, the filter can be a linear model or a non-linear model.
[0216] The video encoding device filters the reconstructed luminance reference samples by using each candidate filter. The video encoding device derives a candidate cross-component model corresponding to each candidate filter based on the filtered reconstructed luminance reference samples and the reference samples of the current chrominance block. Therefore, the same number of candidate cross-component models as the candidate filters can be derived.
[0217] The video encoding device inputs all or part of the filtered reconstructed luminance reference samples into the candidate cross-component model to generate a predicted value of the chrominance component, and calculates a prediction cost. Here, the prediction cost indicates the deviation between the predicted value and the reconstructed chrominance reference samples. Among the candidate filters and the corresponding candidate cross-component models, the video encoding device can determine the candidate filter and the cross-component model with the minimum prediction cost as the filter and the cross-component model for predicting the current chrominance block.
[0218] The video encoding device filters the reconstructed luminance samples in the corresponding luminance block by using the filter coefficients according to the derived filter (S1504).
[0219] The video encoding device generates a first predicted block of the current chrominance block from the filtered reconstructed luminance samples by using the cross-component model (S1506).
[0220] The video encoding device derives a linear model based on the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block (S1508).
[0221] The video encoding device generates a second predicted block of the current chrominance block from the reconstructed luminance samples by using the linear model (S1510).
[0222] The video encoding device determines an implicit filter derivation flag based on the first predicted block and the second predicted block (S1512). Here, the implicit filter derivation flag indicates whether the filter is derived implicitly.
[0223] In terms of rate-distortion optimization, the video encoding device can determine the implicit filter derivation flag. For example, if the first predicted block is the best, the video encoding device sets the implicit filter derivation flag to true. On the other hand, if the second predicted block is the best, the video encoding device sets the implicit filter derivation flag to false.
[0224] The video encoding apparatus encodes the implicit filter derivation flag (S1514).
[0225] In addition, the video encoding apparatus generates a residual block of the current chrominance block by subtracting the first prediction block or the second prediction block according to the implicit filter derivation flag from the current chrominance block. Then, the video encoding apparatus encodes the residual block of the current chrominance block.
[0226] Figure 16 is a flowchart of a method for reconstructing a current chrominance block by a video decoding device according to at least one embodiment of the present invention.
[0227] The video decoding apparatus establishes a luma region corresponding to the current chroma block from the reconstructed region of the luma component (S1600). Here, the corresponding luma block refers to a luma block co-located with the current chroma block.
[0228] The video decoding apparatus decodes an implicit filter derivation flag from a bitstream (S1602). Here, the implicit filter derivation flag indicates whether a filter is derived implicitly.
[0229] The video decoding apparatus checks the implicit filter derivation flag (S1604).
[0230] If the implicit filter derivation flag is true (S1604: Yes), the video decoding apparatus performs the following steps.
[0231] The video decoding apparatus derives a filter applied to the luminance component from the candidate filters based on the reconstructed luminance reference sample of the corresponding luminance block and the reconstructed reference sample of the current chrominance block, and derives a cross-component model corresponding to the derived filter (S1606). Here, the filter can be a linear model or a nonlinear model.
[0232] The video decoding device filters the reconstructed luminance reference samples by using each of the candidate filters. The video decoding device derives a candidate cross-component model corresponding to each candidate filter based on the filtered reconstructed luminance reference samples and the reference samples of the current chrominance block. Therefore, the same number of candidate cross-component models as the candidate filters can be derived.
[0233] The video decoding device inputs all or part of the filtered reconstructed luminance reference sample into the candidate cross-component model to generate a prediction value of the chrominance component and calculate the prediction cost. Here, the prediction cost indicates the deviation between the prediction value and the reconstructed chrominance reference sample. The video decoding device can determine the candidate filter and the candidate cross-component model with the minimum prediction cost among the candidate filters and the corresponding candidate cross-component models as the filter and the cross-component model for predicting the current chrominance block.
[0234] The video decoding device filters the reconstructed luminance samples in the corresponding luminance block by using the filter coefficients of the derived filter (S1608).
[0235] The video decoding device generates a prediction block of the current chrominance block from the filtered reconstructed luminance samples by using a cross-component model (S1610).
[0236] On the other hand, if the implicit filter derivation flag is false (S1604 is NO), the video decoding device performs the following steps.
[0237] The video decoding device derives a linear model based on the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block (S1620).
[0238] The video decoding device generates a prediction block of the current chrominance block from the reconstructed luminance samples by using the linear model (S1622).
[0239] In addition, the video decoding device decodes the residual block of the current chrominance block. Thereafter, the video decoding device may generate a reconstructed block of the current chrominance block by summing the prediction block and the decoded residual block of the current chrominance block.
[0240] Although the steps in the respective flowcharts described are executed sequentially, these steps merely illustrate the technical ideas of some embodiments of the present invention. Therefore, those of ordinary skill in the art to which the present invention pertains may execute the steps by changing the order described in the respective drawings or by executing two or more steps in parallel. Therefore, the steps in the respective flowcharts are not limited to the order shown in the order of occurrence.
[0241] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or combinations thereof. It should also be understood that the functional components described in the present invention are labeled as "…… unit" to emphasize the possibility of their independent implementation.
[0242] On the other hand, various methods or functions described in some embodiments can be implemented as instructions stored in a non-volatile recording medium, and the instructions can be read and executed by one or more processors. The non-volatile recording medium may include various types of recording devices that store data in a form readable by a computer system. For example, the non-volatile recording medium may include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical disk drives, magnetic hard disk drives, and solid-state drives (SSD), and the like.
[0243] Although exemplary embodiments of the present invention have been described for illustrative purposes, those of ordinary skill in the art to which the present invention pertains should understand that various modifications, additions, and substitutions can be made without departing from the spirit and scope of the present invention. Accordingly, the embodiments of the present invention have been described for simplicity and clarity. The scope of the technical idea of the embodiments of the present invention is not limited by the examples. Accordingly, those of ordinary skill in the art to which the present invention pertains should understand that the scope of the present invention should not be limited by the embodiments explicitly described above, but by the claims and their equivalents.
[0244] Reference Numeral
[0245] 122: Intra Predictor
[0246] 542: Intra Predictor
[0247] 710: Implicit Filter Deriver
[0248] 712: Luminance Component Filtering Unit
[0249] 714: Cross-Component Model Estimator
[0250] 716: Prediction Cost Calculator
[0251] 720: Luminance Component Filter Applicator
[0252] 730: Prediction Executor.
[0253] Cross-Reference to Related Applications
[0254] This application claims the priority and benefit of Korean Patent Application No. 10-2022-0156717, filed on November 21, 2022, and Korean Patent Application No. 10-2023-0120682, filed on September 11, 2023, the entire contents of each of which are incorporated herein by reference.
Claims
1. A method for reconstructing a current chrominance block by a video decoding device, the method comprising: establishing a corresponding luminance block of the current chrominance block from a reconstructed region of a luminance component, the corresponding luminance block representing a luminance block collocated with the current chrominance block; deriving a filter of a linear model or a nonlinear model applied to the luminance component and a cross-component model corresponding to the filter from candidate filters based on the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block; filtering the reconstructed luminance samples in the corresponding luminance block by using filter coefficients of the derived filter; and generating a prediction block of the current chrominance block from the filtered reconstructed luminance samples by using the cross-component model.
2. The method according to claim 1, further comprising: decoding an implicit filter derivation flag from a bitstream indicating whether to implicitly derive a filter; and checking the implicit filter derivation flag, wherein, when the implicit filter derivation flag is true, the method further comprises: deriving a filter and a cross-component model corresponding to the filter, filtering the reconstructed luminance samples, and generating a prediction block of the current chrominance block.
3. The method according to claim 1, further comprising, when the implicit filter derivation flag is false: deriving a linear model based on the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block; and generating a prediction block of the current chrominance block from the reconstructed luminance samples by using the linear model.
4. The method according to claim 1, wherein, The candidate filters include: a rectangular-shaped filter, a square-shaped filter, and / or a diamond-shaped filter, wherein the rectangular-shaped filter or the square-shaped filter includes: a filter configured to extract edges of a specific directionality, a downsampling filter configured to extract specific pixel values, and / or a subsampling filter.
5. The method according to claim 1, wherein Deriving a cross-component model corresponding to a filter includes: filtering the reconstructed luminance reference samples by using each of the candidate filters; and deriving a candidate cross-component model corresponding to each of the candidate filters based on the filtered reconstructed luminance reference samples and the reference samples of the current chrominance block.
6. The method according to claim 5, wherein Deriving a cross-component model corresponding to a filter includes: inputting all or part of the filtered reconstructed luminance reference samples into a candidate cross-component model to generate a predicted value of a chrominance component, and calculating a prediction cost, the prediction cost being a deviation between the predicted value and the reconstructed chrominance reference samples; and determining both the candidate filter and the candidate cross-component model that minimize the prediction cost among the candidate filters and the corresponding candidate cross-component models as the filter and the cross-component model.
7. The method according to claim 5, wherein Deriving a cross-component model corresponding to a filter includes: deriving a linear model or a nonlinear model as a candidate cross-component model based on a preset number of sample pairs of the filtered reconstructed luminance reference samples and the reconstructed reference samples of the current chrominance block.
8. The method according to claim 1, wherein The non-linear model is configured to generate a predicted value of a chrominance component by using all or part of the following: filtered luminance samples to which a filter has been applied, coordinates in the horizontal direction of the filtered luminance samples, coordinates in the vertical direction of the filtered luminance samples, non-linear terms of the filtered luminance samples, bias terms, gradients in the horizontal direction of the luminance samples, and gradients in the vertical direction of the luminance samples.
9. The method according to claim 1, wherein, The non-linear model is configured to generate a predicted value of a chrominance component by using all or part of the following: filtered luminance samples to which a filter has been applied, filtered adjacent luminance samples of the filtered luminance samples, coordinates in the horizontal direction of the filtered luminance samples, coordinates in the vertical direction of the filtered luminance samples, non-linear terms of the filtered luminance samples, bias terms, gradients in the horizontal direction of the luminance samples, and gradients in the vertical direction of the luminance samples.
10. The method according to claim 1, wherein, Deriving a cross-component model corresponding to a filter includes: filtering the reconstructed luminance reference samples according to each of the candidate position patterns by using each of the candidate filters; and deriving a candidate cross-component model corresponding to each of the candidate filters and each of the candidate position patterns based on the filtered reconstructed luminance reference samples and the reference samples of the current chrominance block.
11. According to the method of claim 10, wherein, Deriving a cross-component model corresponding to a filter includes: inputting all or part of the filtered reconstructed luminance reference samples into a candidate cross-component model to generate a predicted value of a chrominance component, and calculating a prediction cost, where the prediction cost is the deviation between the predicted value and the reconstructed chrominance reference samples; and among the candidate filters, candidate position patterns, and candidate cross-component models corresponding to combinations of the candidate filters and candidate position patterns, determining the candidate filter, candidate position pattern, and candidate cross-component model that minimize the prediction cost as the filter, the position pattern of the current chrominance mode, and the cross-component model.
12. The method according to claim 10, wherein Each of the candidate position patterns includes: an AL mode using the upper-side and left-side reference samples of the current chrominance block, an A mode using the upper-side reference sample of the current chrominance block, or an L mode using the left-side reference sample of the current chrominance block.
13. A method for encoding a current chrominance block by a video encoding device, the method including: establishing a corresponding luminance block of the current chrominance block from a reconstructed region of a luminance component, where the corresponding luminance block represents a luminance block that is co-located with the current chrominance block; deriving a filter of a linear model or a non-linear model applied to the luminance component from candidate filters based on the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block, and deriving a cross-component model corresponding to the filter; filtering the reconstructed luminance samples in the corresponding luminance block by using filter coefficients according to the derived filter; and generating a first prediction block of the current chrominance block from the filtered reconstructed luminance samples by using the cross-component model.
14. The method according to claim 13, further including: deriving a linear model based on the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block; and Generating a second prediction block of a current chrominance block from the reconstructed luminance samples by using a linear model.
15. The method according to claim 14, further comprising: Determining an implicit filter derivation flag indicating whether to implicitly derive a filter based on the first prediction block and the second prediction block; And Encoding the implicit filter derivation flag.
16. A computer-readable recording medium storing a bitstream generated by a video encoding method, wherein, The video coding method includes: Establishing a corresponding luminance block of a current chrominance block from a reconstructed region of a luminance component, the corresponding luminance block representing a luminance block co-located with the current chrominance block; Deriving a filter of a linear model or a non-linear model applied to the luminance component from the reconstructed luminance reference samples of the corresponding luminance block and the reconstructed reference samples of the current chrominance block, and deriving a cross-component model corresponding to the filter; Filtering the reconstructed luminance samples in the corresponding luminance block by using the filter coefficients of the derived filter; and Generating a first prediction block of the current chrominance block from the filtered reconstructed luminance samples by using the cross-component model.
Citation Information
Patent Citations
Preparation and Composition of Flavor Oil with Enhanced Shiitake Mushroom Flavor
KR1020220156717A
Power quality improvement and smart electric fire detection system and method
KR1020230120682A