Method and apparatus for cross-component prediction based on inter prediction signal
By generating chroma block prediction signals using motion information in inter-frame prediction, intra-frame block copying, or intra-frame template matching modes, and deriving filter coefficients to improve the coding performance of chroma components, the problem of insufficient coding efficiency and image quality in existing technologies is solved, and more efficient video coding is achieved.
Patent Information
- Application Number
- CN202480032778.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2024-03-27
- Publication Date
- 2025-12-12
Smart Images

Figure CN121128177A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a video encoding method and apparatus using cross-component prediction based on inter-prediction signal. BACKGROUND
[0002] The following description is provided solely to assist in understanding the present application and is not admitted to be prior art.
[0003] Since video data has more data amount than voice data or still image data, etc., a large amount of hardware resources including a memory are required to directly store or transmit video data without compression processing.
[0004] Therefore, when storing or transmitting video data, it is common to compress video data using an encoder to store or transmit, and a decoder receives the compressed video data, decompresses and plays. Such video compression techniques include H.264 / AVC, High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), etc., and VVC improves the coding efficiency by more than about 30% compared to HEVC.
[0005] However, as the size, resolution, and frame rate of video gradually increase, the amount of data to be encoded also increases, and thus a new compression technique having higher coding efficiency than the existing compression technique and having more improved quality of image is required.
[0006] In an Enhanced Compression Model (ECM) as a next-generation technology, when a corresponding luma block is predicted according to an inter-prediction or Intra Block Copy (IBC) mode in a single tree, a Cross-component Residual Model (CCRM) mode is applied to prediction of a chroma component. The CCRM mode derives convolutional filter coefficients based on a relationship between a prediction signal of a corresponding luma component and a prediction signal of a chroma component. The CCRM mode applies the derived filter to a reconstructed signal of the corresponding luma component to predict the chroma component.
[0007] In AV1, which is an open-source based video codec, a Chroma from Luma (CFL) mode utilizes an Alternative Current (AC) component of a reconstructed luma component to predict an AC component of a chroma component when intra-predicting. The CFL mode subtracts a mean, i.e., a Direct Current (DC), from the reconstructed luma component to generate the AC component of the luma component. The CFL mode multiplies the AC component of the luma component by a scaling parameter to predict the AC component of the chroma component, and utilizes reference samples to estimate a DC component of the chroma component. The CFL mode adds the predicted DC component to the AC component to generate a final prediction signal of the chroma component.
[0008] As described above, the CCRM mode is applied to cross-component prediction in a single coding tree structure, while the CFL mode is applied to cross-component prediction in intra-prediction. Therefore, in order to improve video coding efficiency and improve video quality, when cross-component prediction is performed for a chroma component, a scheme capable of effectively utilizing a corresponding luma component in various environments is needed. SUMMARY
[0009] Technical Problem The purpose of the present disclosure is to provide a video coding method and apparatus for cross-component prediction of a chroma component based on a signal predicted according to motion information when a corresponding luma block is coded in an inter-prediction, Intra Block Copy (IBC) mode, or Intra Template Matching Prediction (IntraTMP) mode.
[0010] Means for Solving Technical Problem According to an embodiment of the present disclosure, a method of reconstructing a current chroma block performed by a video decoding apparatus is provided, including the steps of: generating a first prediction signal of the current chroma block using motion information, wherein the motion information is a motion vector or a block vector; obtaining a luma prediction signal and a luma reconstructed sample of a luma block in a corresponding luma region; deriving a filter coefficient for cross-component prediction of the current chroma block using the luma prediction signal and the first prediction signal; and applying the filter coefficient to the luma reconstructed sample to generate a second prediction signal of the current chroma block.
[0011] According to another embodiment of the present disclosure, a method of encoding a current chroma block performed by a video encoding apparatus is provided, including the steps of: generating a first prediction signal of the current chroma block using motion information, wherein the motion information is a motion vector or a block vector; obtaining a luma prediction signal and a luma reconstructed sample of a luma block in a corresponding luma region; deriving filter coefficients for cross-component prediction of the current chroma block using the luma prediction signal and the first prediction signal; and applying the filter coefficients to the luma reconstructed sample to generate a second prediction signal of the current chroma block.
[0012] According to another embodiment of the present disclosure, a recording medium storing a bitstream generated by a video encoding method and capable of being read by a computer is provided, the video encoding method including the steps of: generating a first prediction signal of a current chroma block using motion information, wherein the motion information is a motion vector or a block vector; obtaining a luma prediction signal and a luma reconstructed sample of a luma block in a corresponding luma region; deriving filter coefficients for cross-component prediction of the current chroma block using the luma prediction signal and the first prediction signal; and applying the filter coefficients to the luma reconstructed sample to generate a second prediction signal of the current chroma block.
[0013] Inventive Effects As described above, according to the present embodiment, by providing a video encoding method and apparatus, when a corresponding luma block is encoded in an inter prediction, IBC mode or Intra TMP mode, a cross-component prediction is performed on a chroma component based on a signal predicted according to motion information, so that the effect of improving the encoding performance of the chroma component, improving the video encoding efficiency and improving the video quality can be achieved. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 is an exemplary block diagram of a video encoding apparatus capable of implementing the technology of the present disclosure.
[0015] Figure 2 is a diagram for explaining a method of partitioning a block using a quadtree plus binary tree ternary tree (QTBTTT) structure.
[0016] Figure 3a and Figure 3b is a diagram showing a plurality of intra prediction modes including a wide angle intra prediction mode.
[0017] Figure 4 is an exemplary diagram of a surrounding block of a current block.
[0018] Figure 5 is an exemplary block diagram of a video decoding apparatus capable of implementing the technology of the present disclosure.
[0019] Figure 6 is an example diagram illustrating prediction of chroma components based on a Cross-component Residual Model (CCRM) mode.
[0020] Figure 7 is an example diagram illustrating a reconstructed signal of a luma component to which a convolutional filter is applied.
[0021] Figure 8 is an example diagram illustrating prediction of chroma components based on a CFL (Chroma from Luma) mode.
[0022] Figure 9 is a block diagram illustrating a part of a video decoding apparatus based on an embodiment of the present disclosure in detail.
[0023] Figure 10 is a block diagram illustrating a prediction performing unit based on an embodiment of the present disclosure in detail.
[0024] Figure 11 is an example diagram illustrating a reconstructed sample region used for scaling parameter derivation based on an embodiment of the present disclosure.
[0025] Figures 12a to 12c is an example diagram illustrating a reconstructed sample region used for scaling parameter derivation based on another embodiment of the present disclosure.
[0026] Figure 13 is an example diagram illustrating derivation of scaling parameters based on an embodiment of the present disclosure.
[0027] Figure 14 is an example diagram illustrating a reconstructed signal of a luma component to which a convolutional filter is applied based on an embodiment of the present disclosure.
[0028] Figure 15 is an example diagram illustrating cross-component prediction of chroma components based on an embodiment of the present disclosure.
[0029] Figure 16 is an example diagram illustrating a predetermined position within a corresponding luma region based on an embodiment of the present disclosure.
[0030] Figure 17 is an example diagram illustrating a position for confirming a prediction mode of a surrounding block based on an embodiment of the present disclosure.
[0031] Figure 18 is a flowchart illustrating a method of encoding, by a video encoding apparatus, a current chroma block based on an embodiment of the present disclosure.
[0032] Figure 19is a vehicle diagram illustrating a method of reconstructing a current chroma block by a video decoding apparatus according to an embodiment of the present disclosure.
[0033] Figure 20 is a flowchart illustrating a method of encoding a current chroma block by a video encoding apparatus according to another embodiment of the present disclosure.
[0034] Figure 21 is a flowchart illustrating a method of reconstructing a current chroma block by a video decoding apparatus according to another embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that, when referring to portions of the drawings, the same reference numerals are used for the same portions throughout the several drawings. Also, in describing the present embodiments, if it is determined that a detailed description of related known structures or functions can obscure the gist of the present embodiments, a detailed description thereof will be omitted.
[0036] Figure 1 is an exemplary block diagram of a video encoding apparatus capable of implementing the technology of the present disclosure. Hereinafter, referring to the drawing of Figure 1 , a video encoding apparatus and a lower structure of the apparatus will be described.
[0037] The video encoding apparatus can include a picture partitioning section 110, a prediction section 120, a subtracter 130, a transform section 140, a quantization section 145, a reordering section 150, an entropy encoding section 155, an inverse quantization section 160, an inverse transform section 165, an adder 170, a loop filter section 180, and a memory 190.
[0038] Each component of the video encoding apparatus can be implemented by hardware or software, or can be implemented by a combination of hardware and software. Also, the function of each component can be implemented by software, and a microprocessor can execute the software corresponding to each component.
[0039] A video consists of a sequence of more than one image. Each image is segmented into multiple regions, and encoding is performed on each region. For example, an image may be segmented into more than one tile and / or slice. Here, more than one tile can be defined as a tile group. Each tile or slice is segmented into more than one coding tree unit (CTU). Furthermore, each CTU is segmented into more than one coding unit (CU) through a tree structure. Information applied to each CU is encoded as the syntax of the CU, while information commonly applied to multiple CUs included in a CTU is encoded as the syntax of the CTU. In addition, information commonly applied to all blocks within a slice is encoded as the syntax of the slice header, while information applied to all blocks constituting more than one image is encoded in the picture parameter set (PPS) or picture header. Moreover, information commonly referenced by multiple images is encoded in the sequence parameter set (SPS). Furthermore, information jointly referenced by more than one SPS is encoded in the Video Parameter Set (VPS). Additionally, information commonly applied to a tile or tile group can also be encoded as the syntax of the tile or tile group header. The syntax contained within SPS, PPS, slice headers, tiles, or tile group headers can be referred to as high-level syntax.
[0040] The image segmentation unit 110 determines the size of the CTU. Information about the size of the CTU (CTU size) is encoded into SPS or PPS syntax and transmitted to the video decoding device.
[0041] After the image segmentation unit 110 segments each picture that constitutes the image into multiple CTUs of a predetermined size, it recursively segments the CTUs using a tree structure. The leaf nodes in the tree structure become the basic unit of encoding, i.e., CUs.
[0042] As a tree structure, it can be a quadtree (QT) where the upper-level node (or parent node) is divided into four lower-level nodes (or child nodes) of the same size; a binary tree (BT) where the upper-level node is divided into two lower-level nodes; a ternary tree (TT) where the upper-level node is divided into three lower-level nodes in a 1:2:1 ratio; or a structure that combines two or more of these QT, BT, and TT structures. For example, a quadtree plus binary tree (QTBT) structure or a quadtree plus binary tree ternary tree (QTBTTT) structure can be used. Here, BTTT can be combined and called a multiple-type tree (MTT).
[0043] Figure 2 This is a diagram used to illustrate the method of segmenting blocks using the QTBTTT structure.
[0044] like Figure 2 As shown, the CTU can first be segmented into a QT structure. Quadtree segmentation can be repeated until the size of the splitting block reaches the minimum block size (MinQTSize) allowed for leaf nodes in the QT. A first flag (QT_split_flag) indicating whether each node in the QT structure has been split into the next four nodes is encoded by the entropy coding unit 155 and sent as a signal to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) allowed for the root node in the BT, further segmentation can be performed using either a BT structure or a TT structure. Multiple segmentation directions can exist in the BT structure and / or TT structure. For example, there can be both directions where the block of the corresponding node is segmented horizontally and directions where it is segmented vertically. Figure 2 As shown, when MTT splitting begins, a second flag (mtt_split_flag) indicating whether a node is split, a flag indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (binary or triangular) are encoded by the entropy coding unit 155 and sent to the video decoding device.
[0045] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node has been split into the four lower-level nodes, the CU splitting flag (split_cu_flag) indicating whether the node has been split can be encoded first. When the CU splitting flag (split_cu_flag) value indicates that it has not been split, the block of the corresponding node becomes a leaf node in the split tree structure, becoming a CU (coding unit) as the basic unit of encoding. When the CU splitting flag (split_cu_flag) value indicates that it has been split, the video encoding device encodes from the first flag in the manner described above.
[0046] As another example of a tree structure, when using QTBT, there are two types: horizontally splitting the corresponding node's block into two blocks of the same size (i.e., symmetric horizontal splitting) and vertically splitting the block into two blocks of the same size (i.e., symmetric vertical splitting). A splitting flag (split_flag) indicating whether each node in the BT structure is split into lower-level blocks, and splitting type information indicating the splitting type, are encoded by the entropy coding unit 155 and transmitted to the video decoding device. On the other hand, there is also a type that splits the corresponding node's block into two asymmetrical blocks. Asymmetrical splitting can include splitting the corresponding node's block into two rectangular blocks with a 1:3 size ratio, or splitting the corresponding node's block along a diagonal direction.
[0047] Depending on the QTBT or QTBTTT segmentation from the CTU, the CU can have various sizes. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) will be referred to as the "current block". When using QTBTTT segmentation, the current block can be either square or rectangular.
[0048] The prediction unit 120 predicts the current block to generate a prediction block. The prediction unit 120 includes an intra-frame prediction unit 122 and an inter-frame prediction unit 124.
[0049] Typically, predictive coding can be performed on each current block within an image. Generally, prediction of the current block can be performed using intra-frame prediction techniques (using data from the image containing the current block) or inter-frame prediction techniques (using data from images encoded before the image containing the current block). Inter-frame prediction includes one-way prediction and two-way prediction.
[0050] The intra-frame prediction unit 122 uses pixels (reference pixels) located around the current block in the current image, including the current block, to predict pixels within the current block. Depending on the prediction direction, multiple intra-frame prediction modes exist. For example, such as... Figure 3a As shown, multiple intra-frame prediction modes can include two non-directional modes (Planar mode and DC mode) and 65 directional modes. The surrounding pixels to be used and the calculation formulas are defined differently for each prediction mode.
[0051] To perform effective directional prediction for the current rectangular block, additional methods can be used. Figure 3b The dashed arrows in the diagram illustrate several directional modes (67 to 80, -1 to -14 intra-prediction modes). These can be referred to as "wide angle intra-prediction modes". Figure 3b The middle arrow indicates the corresponding reference sample used in the prediction and does not indicate the prediction direction. The prediction direction is opposite to the direction the arrow points. Wide-angle intra-frame prediction mode is a mode that performs prediction in the opposite direction of a specific directional mode without transmitting additional bits when the current block is rectangular. In this case, based on the ratio of the width to the height of the current rectangular block, a subset of wide-angle intra-frame prediction modes that can be used for the current block can be determined. For example, wide-angle intra-frame prediction modes with an angle less than 45 degrees (67 to 80 intra-frame prediction modes) are available when the current block is a rectangle with a height less than its width; while wide-angle intra-frame prediction modes with an angle greater than -135 degrees (-1 to -14 intra-frame prediction modes) are available when the current block is a rectangle with a width greater than its height.
[0052] The intra prediction unit 122 is capable of determining the intra prediction mode for encoding the current block. In some examples, the intra prediction unit 122 can use multiple intra prediction modes to encode the current block, and can also select a suitable intra prediction mode to use from the tested modes. For example, the intra prediction unit 122 uses rate-distortion analysis for multiple tested intra prediction modes to calculate rate-distortion values, and selects the intra prediction mode with the best rate-distortion characteristics from the multiple tested modes.
[0053] The intra-prediction unit 122 selects one intra-prediction mode from a plurality of intra-prediction modes and uses the surrounding pixels (reference pixels) determined according to the selected intra-prediction mode and the calculation formula to predict the current block. Information about the selected intra-prediction mode is encoded by the entropy coding unit 155 and transmitted to the video decoding device.
[0054] The inter-frame prediction unit 124 generates a prediction block for the current block using a motion compensation process. The inter-frame prediction unit 124 searches for the most similar block to the current block within reference images that have been encoded and decoded prior to the current image, and generates a prediction block for the current block using the searched block. Furthermore, it generates a motion vector (MV) corresponding to the displacement between the current block in the current image and the prediction block in the reference image. Typically, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma and chroma components. Information about the reference image used for predicting the current block, and motion information including information about the motion vector, are encoded by the entropy coding unit 155 and transmitted to the video decoding device.
[0055] To improve prediction accuracy, the inter-frame prediction unit 124 can also perform interpolation on a reference image or reference block. That is, subsamples between two consecutive integer samples are obtained by interpolation using filter coefficients applied to multiple consecutive integer samples, including those two integer samples. If a process of searching for the most similar block to the current block is performed on the interpolated reference image, the motion vector can be represented with precision down to decimal units, rather than integer sample units. The precision or resolution of the motion vector can be set differently depending on the target region to be encoded, such as a slice, tile, CTU, CU, etc. When applying the Adaptive Motion Vector Resolution (AMVR) as described above, information about the motion vector resolution applied to each target region should be sent via a signal for each target region. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is sent via a signal. The information about the motion vector resolution can be information representing the precision of the differential motion vector, which will be described later.
[0056] On the other hand, the inter-frame prediction unit 124 can perform inter-frame prediction using bi-prediction. In the case of bi-prediction, two reference images and two motion vectors representing the positions of the most similar blocks to the current block within each reference image are used. The inter-frame prediction unit 124 selects a first reference image and a second reference image from reference image list 0 (RefPicList0) and reference image list 1 (RefPicList1), respectively, and searches for blocks similar to the current block within each reference image to generate a first reference block and a second reference block. Furthermore, the first reference block and the second reference block are averaged or weighted averaged to generate the predicted block for the current block. Information about the two reference images used for predicting the current block and motion information including information about the two motion vectors are transmitted to the entropy coding unit 155. Here, reference image list 0 can consist of images in the reconstructed images that are displayed before the current image, and reference image list 1 can consist of images in the reconstructed images that are displayed after the current image. However, this is not a strict limitation. Reconstructed images that are displayed after the current image can be additionally included in reference image list 0, and conversely, reconstructed images that are displayed before the current image can also be additionally included in reference image list 1.
[0057] Various methods can be used to minimize the number of bits required to encode motion information.
[0058] For example, when the reference image and motion vector of the current block are the same as those of surrounding blocks, the motion information of the current block can be transmitted to the video decoding device by encoding the information that can identify the surrounding blocks. This method can be called "merge mode".
[0059] In the merge mode, the inter-frame prediction unit 124 selects a predetermined number of merge candidate blocks (hereinafter referred to as "merge candidates") from the surrounding blocks of the current block.
[0060] like Figure 4 As shown, as peripheral blocks used to derive merge candidates, all or some of the left-side block (A0), lower-left-side block (A1), upper-side block (B0), upper-right-side block (B1), and upper-left-side block (B2) adjacent to the current block in the current image can be used. In addition to using the current image containing the current block, blocks located in a reference image (which may be the same as or different from the reference image used for predicting the current block) can also be used as merge candidate pieces. For example, blocks located in the same position as the current block in the reference image (co-located block) or blocks adjacent to that block in the same position can be additionally used as merge candidates. When the number of merge candidates selected by the method described above is less than a preset number, a 0 vector is added to the merge candidates.
[0061] The inter-frame prediction unit 124 uses these peripheral blocks to construct a merge list including a predetermined number of merge candidates. From the multiple merge candidates included in the merge list, it selects a merge candidate to be used as motion information for the current block and generates merge index information to identify the selected candidate. The generated merge index information is encoded by the entropy coding unit 155 and transmitted to the video decoding device.
[0062] Merge skip mode is a special case of merging mode. After quantization, when the transform coefficients used for entropy coding are close to zero, only the peripheral block selection information is transmitted instead of the residual signal. Merge skip mode can achieve relatively high coding efficiency in images with little motion, still images, and screen content images.
[0063] Hereinafter, merge mode and merge skip mode will be collectively referred to as merge / skip mode.
[0064] Another method for encoding motion information is the Advanced Motion Vector Prediction (AMVP) mode.
[0065] In AMVP mode, the inter-frame prediction unit 124 uses multiple peripheral blocks of the current block to derive multiple predicted motion vector candidates for the motion vector of the current block. As the peripheral blocks used to derive the predicted motion vector candidates, multiple predicted motion vector candidates can be derived using... Figure 4 The diagram shows all or part of the left block (A0), lower left block (A1), upper block (B0), upper right block (B1), and upper left block (B2) adjacent to the current block within the current image. Alternatively, blocks located within a reference image (which may be the same as or different from the reference image used for predicting the current block) can be used as peripheral blocks for deriving motion vector candidates, instead of using the current image containing the current block. For example, a block located at the same position as the current block within the reference image (co-located block) or a block adjacent to that block at the same position can be used. Using the method described above, if the number of motion vector candidates is less than a preset number, a zero vector is added to the motion vector candidate list.
[0066] The inter-frame prediction unit 124 uses the motion vectors of these surrounding blocks to derive candidate motion vectors for prediction, and uses these candidate motion vectors to determine the predicted motion vector for the current block. Furthermore, it subtracts the predicted motion vector from the motion vector of the current block to calculate the differential motion vector.
[0067] The predicted motion vectors can be obtained by applying a predefined function (e.g., median, mean, etc.) to the predicted motion vector candidates. The video decoding device is also aware of this predefined function. Furthermore, the surrounding blocks used to derive the predicted motion vector candidates are already encoded and decoded, so the video decoding device is also aware of the motion vectors of these surrounding blocks. Therefore, the video encoding device does not need to encode the information used to identify the predicted motion vector candidates. Thus, what is encoded at this point is information about the differential motion vectors and information about the reference image used to predict the current block.
[0068] On the other hand, the predicted motion vector can also be determined by selecting one from multiple predicted motion vector candidates. In this case, along with information about the differential motion vector and information about the reference image used to predict the current block, additional encoding is performed on the information used to identify the selected predicted motion vector candidate.
[0069] Subtractor 130 subtracts the prediction block generated by intra-frame prediction unit 122 or inter-frame prediction unit 124 from the current block to generate a residual block.
[0070] The transformation unit 140 transforms the residual signal within the residual block, which has pixel values in the spatial domain, into transform coefficients in the frequency domain. The transformation unit 140 can use the entire size of the residual block as a transformation unit to transform the residual signal within the residual block, or it can divide the residual block into multiple sub-blocks and use these sub-blocks as transformation units. Alternatively, it can divide the block into two sub-blocks: a transformation region and a non-transformation region, and only use the transformation region sub-block as a transformation unit to transform the residual signal. Here, the transformation region sub-block can be one of two rectangular blocks with a 1:1 size ratio based on the horizontal (or vertical) axis. In this case, a flag indicating that only the sub-block is transformed (cu_sbt_flag), directional (vertical / horizontal) information (cu_sbt_horizontal_flag), and / or position information (cu_sbt_pos_flag) are encoded by the entropy encoding unit 155 and sent as signals to the video decoding device. In addition, the size of the transformed region sub-blocks can have a 1:3 size ratio based on the horizontal axis (or vertical axis). In this case, the flag (cu_sbt_quad_flag) for dividing the corresponding segment is additionally encoded by the entropy coding unit 155 and sent to the video decoding device as a signal.
[0071] On the other hand, the transform unit 140 can perform transforms on the residual block in both the horizontal and vertical directions. Various types of transform functions or transform matrices can be used for the transform. For example, the transform function pairs used for horizontal and vertical transforms can be defined as a Multiple Transform Set (MTS). The transform unit 140 can select the transform function pair with the highest transform efficiency from the MTS and perform transforms on the residual block in both the horizontal and vertical directions. Information (mts_idx) about the transform function pair selected from the MTS is encoded by the entropy encoding unit 155 and sent as a signal to the video decoding device.
[0072] The quantization unit 145 quantizes the transform coefficients output from the transform unit 140 using quantization parameters and outputs the quantized transform coefficients to the entropy coding unit 155. For a specific block or frame, the quantization unit 145 can directly quantize the relevant residual block without transformation. The quantization unit 145 can also apply different quantization coefficients (scaling values) to the transform coefficients based on their positions within the transform block. The quantization matrix applied to the two-dimensionally arranged quantized transform coefficients can be encoded and sent as a signal to the video decoding device.
[0073] The reordering unit 150 is capable of reordering the coefficient values of the quantized residual values.
[0074] The reordering unit 150 can transform a two-dimensional coefficient array into a one-dimensional coefficient sequence using coefficient scanning. For example, in the reordering unit 150, a zig-zag scan or a diagonal scan can be used to scan from DC coefficients to coefficients in the high-frequency region, and output a one-dimensional coefficient sequence. Depending on the size of the transform unit and the intra-frame prediction mode, a vertical scan that scans the two-dimensional coefficient array in the column direction or a horizontal scan that scans the two-dimensional block shape coefficients in the row direction can be used instead of a zig-zag scan. That is, the scanning method to be used can be determined from zig-zag scan, diagonal scan, vertical scan, and horizontal scan based on the size of the transform unit and the intra-frame prediction mode.
[0075] The entropy coding unit 155 uses various coding methods such as context-based adaptive binary arithmetic code (CABAC) and exponential Golomb coding to encode the sequence of multiple one-dimensional quantized transform coefficients output from the reordering unit 150, thereby generating a bit stream.
[0076] Furthermore, the entropy coding unit 155 encodes information related to block segmentation, such as CTU size, CU segmentation flag, QT segmentation flag, MTT segmentation type, and MTT segmentation direction, so that the video decoding device can segment blocks in the same way as the video encoding device. Additionally, the entropy coding unit 155 encodes information indicating whether the current block is encoded using intra-frame prediction or inter-frame prediction, and encodes intra-frame prediction information (i.e., information about the intra-frame prediction mode) or inter-frame prediction information (motion information encoding mode (merging mode or AMVP mode), merging index in merging mode, reference image index and differential motion vector information in AMVP mode) according to the prediction type. Furthermore, the entropy coding unit 155 encodes information related to quantization, i.e., information about quantization parameters and information about the quantization matrix.
[0077] The inverse quantization unit 160 performs inverse quantization on the quantized transform coefficients output from the quantization unit 145 to generate transform coefficients. The inverse transform unit 165 transforms the transform coefficients output from the inverse quantization unit 160 from the frequency domain to the spatial domain to reconstruct the residual block.
[0078] Adder 170 adds the reconstructed residual block to the prediction block generated by prediction unit 120, thereby reconstructing the current block. When performing intra-frame prediction for the next sequence of blocks, the pixels in the reconstructed current block are used as reference pixels.
[0079] The loop filter unit 180 performs filtering on the reconstructed pixels to reduce blocking artifacts, ringing artifacts, and blurring artifacts caused by block-based prediction and transform / quantization. The loop filter unit 180, as an in-loop filter, can include all or part of the deblocking filter 182, the Sample Adaptive Offset (SAO) filter 184, and the Adaptive Loop Filter (ALF) 186.
[0080] Deblocking filter 182 performs filtering on the boundaries between reconstructed blocks to remove blocking artifacts caused by the encoding / decoding of block units. SAO filter 184 and ALF 186 perform additional filtering on the deblocked image. SAO filter 184 and ALF 186 are filters used to compensate for differences between reconstructed pixels and original pixels caused by lossy coding. SAO filter 184 applies offsets at CTU units, thereby improving not only subjective image quality but also coding efficiency. In contrast, ALF 186 performs filtering on a block-by-block basis, applying different filters by dividing the edges and varying degrees of change of the corresponding blocks to compensate for distortion. Information about the filter coefficients used in the ALF can be encoded and sent as a signal to the video decoding device.
[0081] The reconstructed blocks filtered by deblocking filter 182, SAO filter 184, and ALF 186 are stored in memory 190. When all blocks in an image have been reconstructed, the reconstructed image can be used as a reference image for inter-frame prediction of blocks in subsequent images to be encoded.
[0082] Video encoding devices can store the bit stream of encoded video data on a non-transitory recording medium or transmit it to a video decoding device via a communication network.
[0083] Figure 5 This is an exemplary block diagram of a video decoding apparatus capable of implementing the technology disclosed herein. Hereinafter, reference is made to... Figure 5 The video decoding device and its underlying structure are described.
[0084] The video decoding device may include: an entropy decoding unit 510, a reordering unit 515, an inverse quantization unit 520, an inverse transform unit 530, a prediction unit 540, an adder 550, a loop filter unit 560, and a memory 570.
[0085] and Figure 1 Similarly, the components of a video decoding device can be implemented in hardware or software, or a combination of both. Furthermore, the functions of each component can also be implemented in software, with a microprocessor executing the corresponding software functions.
[0086] The entropy decoding unit 510 decodes the bitstream generated by the video encoding device to extract information related to block segmentation, thereby determining the current block to be decoded and extracting prediction information, residual signal information, etc. required to reconstruct the current block.
[0087] The entropy decoding unit 510 extracts information about the size of the CTU from the sequence parameter set (SPS) or the picture parameter set (PPS) to determine the size of the CTU, and segments the image into CTUs of the determined size. Furthermore, the CTU is identified as the top level (i.e., the root node) of the tree structure, and segmentation information for the CTU is extracted, thereby utilizing the tree structure to segment the CTU.
[0088] For example, when using the QTBTT structure to segment the CTU, the first flag (QT_split_flag) related to the QT split is extracted to divide each node into four lower-level nodes. Then, for the node corresponding to a leaf node of the QT, the second flag (mtt_split_flag) related to the MTT split and the splitting direction (vertical / horizontal) and / or splitting type (binary / ternary) information are extracted, and the corresponding leaf node is segmented using the MTT structure. Thus, each node below the leaf node of the QT is recursively segmented using either a BT or TT structure.
[0089] As another example, when using the QTBTTT structure to split the CTU, the CU splitting flag (split_cu_flag) indicating whether the CU should be split is first extracted. If the corresponding block is split, the first flag (QT_split_flag) can also be extracted. During the splitting process, each node can undergo zero or more repeated QT splits followed by zero or more repeated MTT splits. For example, the CTU can be directly split into MTTs, or conversely, only multiple QT splits can be performed.
[0090] As another example, when segmenting a CTU using a QTBT structure, the first flag (QT_split_flag) related to the QT split is extracted, and each node is split into four lower-level nodes. Furthermore, for nodes corresponding to leaf nodes of the QT, a split flag (split_flag) indicating whether they can be further split using BT and the split direction information are extracted.
[0091] On the other hand, when the entropy decoding unit 510 determines the current block to be decoded using tree structure segmentation, it extracts information about the prediction type, indicating whether the current block undergoes intra-frame prediction or inter-frame prediction. When the prediction type information indicates intra-frame prediction, the entropy decoding unit 510 extracts the syntax elements of the intra-frame prediction information (intra-frame prediction mode) for the current block. When the prediction type information indicates inter-frame prediction, the entropy decoding unit 510 extracts the syntax elements of the inter-frame prediction information, namely, information representing the motion vector and the reference image referenced by that motion vector.
[0092] In addition, the entropy decoding unit 510 extracts information related to quantization, as well as information about the quantized transform coefficients of the current block as information about the residual signal.
[0093] The reordering unit 515 can reorder the sequence of one-dimensional quantized transform coefficients entropy decoded in the entropy decoding unit 510 into a two-dimensional coefficient array (i.e., a block) in the reverse order of coefficient scanning performed by the video encoding device.
[0094] The inverse quantization unit 520 performs inverse quantization on the quantized transform coefficients using quantization parameters. The inverse quantization unit 520 can apply different quantization coefficients (scaling values) to the two-dimensional array of quantized transform coefficients. The inverse quantization unit 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video encoding device to the two-dimensional array of quantized transform coefficients.
[0095] The inverse transform unit 530 inversely transforms the inversely quantized transform coefficients from the frequency domain to the spatial domain to reconstruct the residual signal, thereby generating the residual block of the current block.
[0096] In addition, when performing inverse transformation only on a portion of the transform block (sub-block), the inverse transformation unit 530 extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transform block is transformed, the directionality (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or the position information (cu_sbt_pos_flag) of the sub-block. The residual signal is reconstructed by inversely transforming the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain. For the regions that have not undergone inverse transformation, the residual signal is filled with "0" values, thereby generating the final residual block of the current block.
[0097] In addition, when applying MTS, the inverse transform unit 530 uses the MTS information (mts_idx) sent from the video encoding device to determine the transform function or transform matrix to be applied in the horizontal and vertical directions respectively, and performs inverse transform on the transform coefficients in the transform block in the horizontal and vertical directions using the determined transform function.
[0098] The prediction unit 540 may include an intra-prediction unit 542 and an inter-prediction unit 544. The intra-prediction unit 542 is activated when the prediction type of the current block is intra-prediction, and the inter-prediction unit 544 is activated when the prediction type of the current block is inter-prediction.
[0099] The intra prediction unit 542 determines the intra prediction mode of the current block from multiple intra prediction modes based on the grammatical elements of the intra prediction mode extracted from the entropy decoding unit 510, and predicts the current block using reference pixels around the current block based on the intra prediction mode.
[0100] The inter-frame prediction unit 544 uses the syntax elements of the inter-frame prediction mode extracted from the entropy decoding unit 510 to determine the motion vector of the current block and the reference picture to which the motion vector is referenced, and uses the motion vector and the reference picture to predict the current block.
[0101] Adder 550 adds the residual block output by inverse transform unit 530 to the prediction block output by inter-frame prediction unit 544 or intra-frame prediction unit 542, thereby reconstructing the current block. The pixels in the reconstructed current block will be used as reference pixels when performing intra-frame prediction on the block to be decoded later.
[0102] The loop filtering unit 560, acting as an in-loop filter, can include a deblocking filter 562, a SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between reconstructed blocks to remove blocking artifacts caused by the decoding of block units. The SAO filter 564 and ALF 566 perform additional filtering on the reconstructed blocks after deblocking to compensate for differences between reconstructed pixels and original pixels caused by lossy coding. The filter coefficients of the ALF are determined using information about the filter coefficients decoded from the bitstream.
[0103] The reconstructed blocks filtered by deblocking filter 562, SAO filter 564, and ALF 566 are stored in memory 570. When all blocks in an image are reconstructed, the reconstructed image is used as a reference image for inter-frame prediction of blocks in subsequent images to be encoded.
[0104] This embodiment relates to the encoding and decoding of images (videos) as described above. More specifically, a video encoding method and apparatus are provided, wherein when a corresponding luma block is encoded in inter-frame prediction, intra-block copy (IBC) mode, or intra-template matching prediction (IntraTMP) mode, cross-component prediction of the chroma component is performed based on the signal predicted according to motion information.
[0105] The following embodiments can be executed by the prediction unit 120 within a video encoding apparatus. Alternatively, they can be executed by the prediction unit 540 within a video decoding apparatus.
[0106] When encoding the current block, the video encoding apparatus can generate signaling information relevant to this embodiment from a rate-distortion optimization perspective. The video encoding apparatus can then transmit the signaling information to the video decoding apparatus after encoding it using the entropy encoding unit 155. The video decoding apparatus can use the entropy decoding unit 510 to decode the signaling information related to the decoding of the current block from the bitstream.
[0107] In the following description, the term "target block" may be used in the same sense as the current block or coding unit (CU). Alternatively, "target block" may also refer to a portion of a coding unit.
[0108] Additionally, a flag value of "true" indicates that the flag is set to 1. Conversely, a flag value of "false" indicates that the flag is set to 0.
[0109] I. Intra Block Copy (IBC) and IntraTMP IBC performs intra-frame prediction of the current block by using block vectors to copy reference blocks within the same frame to generate the prediction block for the current block.
[0110] The video encoding device derives the optimal block vector by performing block matching. Here, the block vector represents the displacement from the current block to the reference block. To improve encoding efficiency, the video encoding device does not directly transmit the block vector, but instead divides it into a Block Vector Predictor (BVP) and a Block Vector Difference (BVD), encodes the BVP and BVD, and then transmits them to the video decoding device.
[0111] In terms of using block vectors, IBC exhibits characteristics of inter-frame prediction. Therefore, IBC can be distinguished into IBC merge / skip mode and IBC AMVP mode.
[0112] In IBC merge / skip mode, the video encoding unit constructs an IBC merge list. From a coding efficiency optimization perspective, after selecting a block vector from among several candidates included in the IBC merge list, the video encoding unit can use the selected candidate as a block vector predictor (BVP). The video encoding unit determines a merge index indicating the selected block vector. However, the video encoding unit does not generate a block vector predictor (BVD). The video encoding unit encodes the merge index and transmits it to the video decoding unit. The IBC merge list can be constructed by both the video encoding and video decoding units using the same method. After decoding the merge index, the video decoding unit can generate block vectors from the IBC merge list using the merge index.
[0113] In IBC skip mode, the video encoding device uses the same block vector transmission method as in IBC merge mode, but does not transmit the residual block corresponding to the difference between the current block and the predicted block.
[0114] In IBC AMVP mode, from the perspective of optimizing coding efficiency, the video encoder determines block vectors and constructs an IBC AMVP list. The video encoder determines a candidate index, which indicates one of the candidate block vectors included in the IBC AMVP list as a Block Vector Verifier (BVP). The video encoder calculates the difference between the BVP and the block vector, i.e., the Block Vector Value (BVD). Then, the video encoder encodes the candidate index and the BVD and transmits them to the video decoder.
[0115] The video decoding device decodes the candidate index and the block vector (BVD). After retrieving the BVP indicated by the candidate index from the IBC AMVP list, the video decoding device can add the BVP to the BVD to reconstruct the block vector.
[0116] In the Enhanced Compression Model (ECM), Intra Template Matching Prediction (IntraTMP) sets L-shaped / left / top templates around the current block. After searching for the most similar template in the reconstructed region of the current frame, it uses the block adjacent to the searched template and having the same size as the current block as the predicted block. IntraTMP searches for similar templates based on a loss function, using the Sum of Absolute Differences (SAD) as the loss function.
[0117] II. CCRM and CFL (Chroma from Luma) In ECM, when the luma block corresponding to a single coding tree is predicted according to inter-frame prediction or intra-block copy (IBC) mode, the cross-component residual model (CCRM) mode is applied to the prediction of the chroma component. For example... Figure 6 As shown in the example, the CCRM mode derives the convolutional filter coefficients based on the relationship between the predicted signals of the corresponding luminance components and the predicted signals of the chrominance components. Figure 6In the example, predY represents the predicted signal for the luminance component, and predCb and predCr represent the predicted signals for the chrominance component. resY represents the residual signal for the luminance component, and resCb and resCr represent the residual signals for the chrominance component.
[0118] As shown in Equation 1, the CCRM mode applies the derived filter to... Figure 7 The example uses the reconstructed signal of the corresponding luminance component to predict the chrominance component. Figure 7 In the example, C represents the position of the chromaticity sample generated based on the filter derived by the application.
[0119] Formula 1 Here, the nonlinear function nonlinear(X) and the offset B are defined using the bit depth as shown in Equation 2.
[0120]
Formula 2
[0121] In the open-source video codec AV1, such as Figure 8 As shown in the example, the CFL (Chroma from Luma) mode uses the Alternative Current (AC) component of the reconstructed luma component to predict the AC component of the chroma component during intra-frame prediction. The CFL mode generates the AC component of the luma component by subtracting the average value, i.e., the Direct Current (DC), from the reconstructed luma component. The CFL mode predicts the AC component of the chroma component by multiplying the AC component of the luma component by a scaling parameter, and estimates the DC component of the chroma component using reference samples. The CFL mode generates the final predicted signal of the chroma component by adding the predicted DC component to the AC component. The scaling parameter is transmitted to the video decoding device after being generated by the video encoding device. At this time, the optimal scaling parameters are transmitted separately for the Cb and Cr components.
[0122] As mentioned above, the CCRM mode is used for cross-component prediction in a single coding tree structure, and the CFL mode is used for cross-component prediction in intra-frame prediction. The following describes a scheme for applying the CFL mode to cross-component prediction in inter-frame prediction or IBC mode, and a scheme for applying the CCRM mode to cross-component prediction in a dual coding tree structure.
[0123] Although the following embodiments are described with a focus on a video decoding device, they can also be implemented in the same or similar way in a video encoding device.
[0124] III. Embodiments based on this disclosure The video decoding apparatus based on this embodiment determines prediction and transform units, and for the current block corresponding to the determined units, performs prediction and inverse transform using the determined prediction techniques and prediction modes, thereby ultimately generating a reconstructed block of the current block. Figure 9 The example shown can be performed using the inverse transform unit 530, prediction unit 540, and adder 550 of the video decoding apparatus. On the other hand, it can be performed using the inverse transform unit 165, image segmentation unit 110, prediction unit 120, and adder 170 of the video encoding apparatus. Figure 9 The same operation is shown in the example. In this case, the video decoding device uses the encoded information parsed from the bitstream, while the video encoding device, considering rate-distortion minimization, uses the encoded information set from the upper layer. Hereinafter, for ease of explanation, this embodiment will be described focusing on the video decoding device.
[0125] like Figure 5 As shown in the example, according to the prediction technique, the prediction unit 540 can include an intra-frame prediction unit 542 and an inter-frame prediction unit 544, however, as Figure 9 As shown, the prediction unit 540 may also include all or part of the prediction unit determination unit 902, the prediction technology determination unit 904, the prediction mode determination unit 906, and the prediction execution unit 908.
[0126] When the input video's color format is YUV (YUV420, YUV411, YUV422, YUV444, etc.), the video decoding device can perform prediction and reconstruction of the chrominance component after predicting and reconstructing the luminance component. That is, the luminance and chrominance components can be... Figure 9 The components shown are reconstructed sequentially. On the other hand, when the input video's color format is RGB, the video encoding device can perform a color format conversion from RGB to YUV and then encode the converted video. Here, in YUV format, the color format represents the correspondence between pixels of the luminance component and pixels of the chrominance component.
[0127] The prediction unit determination unit 902 determines a prediction unit (PU). The prediction technique determination unit 904 determines a prediction technique for the prediction unit (e.g., intra-frame prediction, inter-frame prediction, or intra-block copy (IBC) mode, palette mode, etc.). The prediction mode determination unit 906 determines a specific prediction mode for the prediction technique. The prediction execution unit 908 generates a prediction block for the current block based on the determined prediction mode.
[0128] The inverse transformation unit 530 includes a transformation unit determination unit 910 and an inverse transformation execution unit 912. The transformation unit determination unit 910 determines a transformation unit (TU) for the inverse quantization signal of the current block, and the inverse transformation execution unit 912 performs an inverse transformation on the transformation unit represented by the inverse quantization signal to generate a residual signal.
[0129] Adder 550 adds the predicted block to the residual signal to generate a reconstructed block. The reconstructed block can be stored in memory and later used for the prediction of other blocks.
[0130] The prediction unit determined in the prediction unit determination unit 902 can be the current block, or one of the multiple sub-blocks after the current block is divided. In this case, depending on the color format, the prediction unit for the chroma component can be the same size as the prediction unit for the luminance component. Furthermore, after determining the prediction units for both the luminance and chroma components, prediction can be performed for the prediction unit of the chroma component.
[0131] The prediction technique determination unit 904 determines a prediction technique for the prediction unit. As described above, the prediction technique can be one of inter-frame prediction, intra-frame prediction, IBC mode, and palette mode. In this case, the prediction technique for the chroma component can be determined to be the same as the prediction technique for the corresponding luma component without separate information signaling and parsing.
[0132] As an example, when the prediction technique for the current block is not intra-frame prediction, the video decoding device parses the 1-bit flag information. For instance, if the parsed flag indicates a skip mode, the video decoding device determines the prediction mode for the current block as either inter-frame prediction merge mode or IBC merge mode. In skip mode, the video decoding device can use the predicted signal as the reconstructed signal without omitting the inverse transform process (i.e., without parsing the residual signal).
[0133] Conversely, when the parsed flags do not indicate that the current block is in Skip mode, the prediction technique determination unit 904 can parse a series of 1-bit flags to determine the prediction technique of the current block as one of the following: inter-frame prediction, intra-frame prediction, IBC mode, palette mode, etc.
[0134] For example, if a skip is not applied to the current block, and the prediction technique is determined to be inter-frame prediction or IBC mode, the video decoder parses a 1-bit flag. Based on the parsed flag, the prediction mode for the current block can be determined to be either regular merge mode or Advanced Motion Vector Prediction (AMVP) mode.
[0135] The prediction model determination unit 906 determines a specific prediction model for the prediction technology.
[0136] As an example, when the prediction technique for the current block is intra-frame prediction, the prediction mode determination unit 906 can determine a directional prediction mode, a planar mode (horizontal plane, vertical plane, or conventional plane), or a DC mode as the prediction mode for the current block.
[0137] As another example, when the prediction technique for the current block is intra-frame prediction, the prediction mode determination unit 906 can determine the intra-frame TMP mode as the prediction mode for the current block.
[0138] The prediction execution unit 908 generates a prediction block for the current block based on the determined prediction technology and prediction mode.
[0139] As an example, the prediction execution unit 908 generates a prediction block for the current block using motion information, and the adder 550 generates a reconstructed block for the luminance component. The prediction execution unit 908 performs cross-component prediction (CCP) based on the prediction block for the luminance component, the reconstructed block for the luminance component, and the prediction block for the chrominance component to generate a prediction block for the chrominance component, and the adder 550 generates a reconstructed block for the chrominance component.
[0140] Figure 10 This is a block diagram showing in detail a predictive execution unit based on an embodiment of the present disclosure.
[0141] like Figure 10 As shown in the example, for cross-component prediction, the prediction execution unit 908 includes a current block predictor 1002, a CCP performing determiner 1004, a CCP parameter determiner 1006, and a CCP performer 1008. Figure 10 In the example, in addition to the components of the prediction execution unit mentioned above, an adder 550 for generating reconstruction blocks is also included.
[0142] The current block prediction unit 1002 uses motion information to generate prediction blocks for the current block, namely, prediction blocks for the luma block and prediction blocks for the chroma block. During inter-frame prediction, the motion information is the motion vector of a block or sub-block; in IBC mode, the motion information is a block vector. In IntraTMP mode, the motion information is the displacement between multiple templates, i.e., a block vector.
[0143] According to the embodiments, when the inter-frame prediction mode is skip mode or merge / skip mode, cross-component prediction for the inter-frame prediction block can be further performed.
[0144] Adder 550 adds the prediction block of the luma block to the residual signal of the luma block to generate a reconstructed block of the luma component. As described above, the residual signal of the luma component is generated by applying entropy decoding, inverse quantization, and inverse transform to the bitstream. Furthermore, after generating the prediction block of the chroma block, adder 550 adds the prediction block of the chroma block to the residual signal of the chroma block to generate a reconstructed block of the chroma component. As described above, the residual signal of the chroma component is generated by applying entropy decoding, inverse quantization, and inverse transform to the bitstream.
[0145] The CCP execution determination unit 1004 explicitly determines whether to perform cross-component prediction for the current chroma block by parsing an additional flag. If the prediction mode for the corresponding luma block is inter-frame prediction or a block vector-based prediction mode (e.g., IBC or IntraTMP), the flag can be transmitted. According to an embodiment, the flag can be parsed on a chroma block or transform block basis.
[0146] When the execution of cross-component prediction is determined on a block-by-block basis, and the residual signals of the corresponding luma blocks are all zero, cross-component prediction can be implicitly not performed. Furthermore, when the execution of cross-component prediction is determined on a transform block-by-transform basis, and the Code Block Flag (CBF) of the corresponding luma transform unit is 0 (when the residual signals are all zero), cross-component prediction can be implicitly not performed.
[0147] CCP parameter determination unit 1006 determines the parameters used for cross-component prediction.
[0148] As an example, the video decoding device parses the scaling parameter β of the AC component transmitted through the video encoding device.
[0149] Scaling parameters can be resolved at each unit performing cross-component prediction. The video decoding device can receive scaling parameters separately for the Cb / Cr components. The scaling parameters can be determined by resolving an index indicating one of a preset number of values defined according to an agreement between the video encoding and decoding devices. Multiple scaling parameter types can exist, and for each type, the value and number of scaling parameters can be defined differently. Which type of scaling parameter among the multiple types is used can be determined based on surrounding information or relevant syntax elements.
[0150] As another example, the video decoding device derives scaling parameters based on the reconstructed luminance and chrominance signals. In this case, the video decoding device can derive the scaling parameters using the same method as the video encoding device.
[0151] In cross-component prediction, when predicting the AC component of the chrominance block from the AC component of the reconstructed luma block, the video decoding device derives scaling parameters based on the surrounding reconstructed luma and chrominance signals. This allows for... Figure 11 The example shows the setting of reconstructed sample regions for calculating scaling parameters for the corresponding luma and chroma blocks. Here, A, B, C, and D are integer values greater than or equal to 0, which can be defined according to the agreement between the video encoding and decoding devices, the size of the current block, and / or the availability of reference samples. Figure 11 In the example, k and l are determined based on the current image's chroma format. For instance, when the chroma format is 4:2:0, k and l are set to 2.
[0152] According to the embodiments, it is possible to explicitly send signals: such as Figures 12a to 12c The example shows the regions used to derive scaling parameters in the first, second, and third regions. Here, A, B, C, and D are integer values greater than or equal to 0, which can be defined according to the agreement between the video encoding and decoding devices, the size of the current block, and / or the availability of reference samples. Figures 12a to 12c In the example, k and l are determined based on the current image's chroma format. For instance, when the chroma format is 4:2:0, k and l are set to 2.
[0153] Figure 13 This is an example diagram illustrating the derivation of scaling parameters based on an embodiment of the present disclosure.
[0154] The video decoding device calculates scaling parameters based on a reference sample region. The reference sample region used to derive the scaling parameters can be predefined according to an agreement between the video encoding and decoding devices. Alternatively, the reference sample region can be explicitly determined based on signaling.
[0155] The video decoding device calculates the AC value (ReconY) of the luminance component by subtracting the average value of the reference sample area from the reconstructed luminance value (ReconY) of the reference sample area corresponding to the luminance block. AC The video decoder calculates the AC value (ReconC) of the chroma component by subtracting the average value of the reference sample region from the reconstructed chroma value (ReconC) of the chroma block from the reference sample region. AC ReconC is either the Cb or Cr signal. The video decoding device can derive the scaling parameters based on the two AC components.
[0156] At this point, the video decoding device can downsample the reconstructed luminance value according to the chroma format of the current image (e.g., YCbCr 4:2:0, YCbCr 4:2:2, etc.) and use the downsampled reconstructed luminance value to calculate the AC value of the luminance component.
[0157] The scaling parameter can be calculated based on linear regression according to Equation 3.
[0158]
Formula 3
[0159] Formula 4 In Equation 4, X AC max (k) represents the k-th largest AC value calculated in the reference sample region of the reconstructed luminance component, Y AC max (k) represents the chromaticity value corresponding to the position of the luminance pixel. X AC min (k) represents the k-th smallest AC value calculated in the reference sample region of the reconstructed luminance component, Y AC min (k) represents the chromaticity value corresponding to the position of the luminance pixel.
[0160] As another example, the video decoding device derives the filter coefficients based on the reconstructed luminance and chrominance signals. In this case, the video decoding device can derive the filter coefficients using the same method as the video encoding device.
[0161] The video decoding device calculates filter coefficients based on one or more of the following: inter-frame predicted luminance samples (predL), offsets (B) determined based on bit depth, the output of a nonlinear function (NL) using more than one inter-frame predicted luminance sample as input, and inter-frame predicted chrominance samples (predC). Here, the nonlinear function, the input to the nonlinear function (… The offset (B) can be defined in a variety of ways, for example, as in Equation 5.
[0162]
Formula 5
[0163] The video decoding device predicts the chroma signal according to Equation 6.
[0164]
Formula 6
[0165]
Formula 7
[0166] As an inter-frame prediction sample within the corresponding luma block used in the prediction of chroma components and the derivation of filter coefficients, the video decoding device can use the value obtained by subtracting the offset from the prediction sample. In this case, the offset can be calculated according to Equation 8, based on the bit depth of the luma signal.
[0167]
Form 8
[0168] Figure 15 This is an example diagram illustrating cross-component prediction of chromaticity components based on an embodiment of the present disclosure.
[0169] As an example, the video decoding apparatus generates a prediction signal for the chroma component using the AC scaling parameter β. The video decoding apparatus predicts the AC value of the chroma component using the scaling parameter derived from the CCP parameter determination unit 1006 or parsed from the bitstream, and the corresponding reconstructed luminance signal. The video decoding apparatus, as... Figure 15 As shown in the example, the prediction of the chromaticity component is performed using the AC scaling parameter.
[0170] When downsampling the reconstructed luminance block (ReconY) according to the current image's chroma format, the video decoding device calculates the AC value of the luminance component by subtracting the average value of the corresponding luminance block from the downsampled luminance reconstructed sample. The video decoding device calculates the reconstructed luminance AC value (ReconY) for each location. AC (i, j) is multiplied by the scaling parameter to calculate the AC value (ReconC) of the chromaticity component. AC (i, j)). The video decoding device adds the AC value of the chroma component to the average value of the chroma component to generate the final chroma prediction signal. At this time, the video decoding device calculates the average value of the chroma component based on the surrounding reconstructed chroma reference samples. Alternatively, the video decoding device may perform motion compensation of the chroma component using motion information used for the prediction of the corresponding luma block, and then use the average value of the chroma prediction signal as the average value of the chroma component. Here, as mentioned above, the motion information can be a motion vector based on inter-frame prediction or a block vector based on IBC mode or IntraTMP mode.
[0171] As another example, the video decoding apparatus generates a prediction signal for the chroma component using a filter derived in the CCP parameter determination unit 1006. For example, the video decoding apparatus can generate the prediction signal for the chroma component using a filter according to Equation 6. In Equation 6, L0 to L5 represent reconstructed luminance samples generated by adding the inter-frame prediction signal corresponding to the luminance component to the reconstructed residual signal. At this time, the inter-frame prediction signal for the luminance component is a signal predicted based on the motion vector. The video decoding apparatus applies the derived filter to the sample values of the reconstructed corresponding luminance block to generate the final chroma prediction signal.
[0172] As another example, when the corresponding luma block is predicted according to Intra Block Copy (IBC) or IntraTMP mode, the video decoding device performs the derivation of filter coefficients and the prediction of chroma components in the same manner as for inter-frame predicted blocks. In this case, the video decoding device uses the luma prediction signal and chroma prediction signal predicted based on the block vector for the derivation of the filter coefficients.
[0173] The following explains the case where one chroma block corresponds to multiple luma blocks when a dual tree structure is applied. Multiple luma blocks corresponding to a chroma region are represented as corresponding luma regions.
[0174] For example, when determining whether to perform cross-component prediction on a block-by-block basis, and the luma block corresponding to the center position of the current chroma block is predicted using inter-frame prediction with motion vectors or prediction modes using block vectors (IBC or IntraTMP), the video decoding device can additionally parse a flag to determine whether to perform cross-component prediction. On the other hand, when all residual signals of the luma block corresponding to the center position of the current chroma block are zero, cross-component prediction can be implicitly not performed.
[0175] Furthermore, when determining whether to perform prediction for each transform unit (TU), a prediction mode corresponding to the luma block at the TU center position is used. For example, if the luma block corresponding to the TU center position is predicted using a prediction mode that uses motion vectors or block vectors, the video decoding device can parse the flags to determine whether to perform cross-component prediction. On the other hand, when the CBF of the luma block corresponding to the TU center position is 0 (all residual signals are zero), cross-component prediction can be implicitly not performed.
[0176] As an example, after determining the prediction modes for predetermined positions within a corresponding luminance region in a predetermined order, if there are luminance blocks predicted using a prediction mode employing motion vectors or block vectors, the video decoding device can parse the flags to determine whether to perform cross-component prediction. In this case, the predetermined order and the predetermined positions within the corresponding luminance region can be set according to an agreement between the video encoding device and the video decoding device. For example... Figure 16 As shown in the example, the predetermined positions within the corresponding brightness area can include the center, upper left corner, upper right corner, lower left corner, lower right corner, etc.
[0177] On the other hand, when multiple blocks are predicted using prediction modes employing motion vectors or block vectors at a specified location, the video decoding apparatus utilizes the block predicted by the first prediction mode employing motion vectors or block vectors in a predetermined order. That is, the video decoding apparatus can derive filter coefficients using the signal predicted based on the motion vectors or block vectors of the first block, and perform cross-component prediction using the derived filter.
[0178] As an example, when deriving filter coefficients based on a signal predicted from motion information and using the derived filter to predict chroma components, the video decoding device can average or weightedly sum the chroma prediction signal generated using the filter with the chroma prediction signal generated from the motion information to generate the final prediction signal. Here, as mentioned above, the motion information can be a motion vector based on inter-frame prediction or a block vector based on IBC mode or IntraTMP mode. When performing cross-component prediction, the video decoding device can always perform weighted sum prediction. Alternatively, the video decoding device can explicitly resolve a flag to determine whether to further perform weighted sum prediction.
[0179] The weight values used for weighted summation can be predefined according to an agreement between the video encoding and decoding devices. Alternatively, the video decoding device calculates weight values for the motion-based prediction signal and the cross-component prediction signal based on the prediction mode information of the peripheral blocks. When calculating weight values based on the prediction modes of the peripheral blocks, the location of the prediction modes can vary depending on the embodiment. Figure 17 As shown in the example, the video decoding device can determine the prediction mode of the surrounding blocks at positions such as TR / BL or TR / BL / TL0 / TL1, thereby deriving the weight values.
[0180] Based on the current chroma block, the number of inter-frame predicted blocks in the prediction mode of surrounding blocks is defined as N. inter The number of non-inter-frame predicted blocks is defined as N. noninter At this time, N inter It can include blocks predicted according to a prediction mode using block vectors. The video decoding device can determine the weight values (W) of the prediction signal based on motion information as shown in Table 1. inter ) and the weight values (W) of the cross-component prediction signal interCC ).
[0181] Table 1 In Table 1, 2 P Used to divide by the predicted value based on weighted summation. q02 p - q0> q0 and q12 p - q1> q1 is a value defined according to the agreement between the video encoding device and the video decoding device.
[0182] When performing weighted summation, the video decoding device generates the final predicted value of the chroma component according to Equation 9.
[0183]
Form 9
[0184] The following uses Figures 18 to 21 The diagram illustrates a method for cross-component prediction of the current chroma block based on derived filter coefficients. The following describes the case where one chroma block corresponds to multiple luma blocks when a dual-tree structure is applied. Multiple luma blocks corresponding to a chroma region are represented as corresponding luma regions.
[0185] exist Figure 18 and Figure 19 The diagram uses a flag indicating whether a weighted summation is performed on the two predictors.
[0186] Figure 18 This is a flowchart illustrating a method for encoding a current chroma block using a video encoding apparatus based on an embodiment of the present disclosure.
[0187] The video encoding device uses motion information to generate the first prediction signal for the current chroma block (S1800). Here, in the case of inter-frame prediction, the motion information is the motion vector of the block or sub-block; in the case of IBC mode, the motion information is the block vector; and in the case of IntraTMP mode, the motion information is the displacement between multiple templates, i.e., the block vector.
[0188] The video encoding device acquires the luminance prediction signal and luminance reconstruction sample of the luminance block in the corresponding luminance area (S1802).
[0189] When the luminance block corresponding to the center position of the current chroma block is predicted according to the prediction mode using motion information, the video coding device performs cross-component prediction.
[0190] When the residual signals of the luminance blocks corresponding to the center position of the current chroma block are all zero, the video coding device implicitly omits cross-component prediction.
[0191] When determining whether to perform prediction for each transform unit and the luminance block corresponding to the center position of the transform unit is predicted according to a prediction mode using motion information, the video coding device performs cross-component prediction.
[0192] When the residual signals of the luminance blocks corresponding to the center position of the transform unit are all zero, the video coding device implicitly omits cross-component prediction.
[0193] The video encoding device determines the prediction pattern at a predetermined position within the corresponding luminance area according to a predetermined order. If there is a luminance block that is predicted based on the prediction pattern using motion information, cross-component prediction is performed.
[0194] The video encoding device uses the luminance prediction signal and the first prediction signal to derive the filter coefficients for cross-component prediction of the current chroma block (S1804).
[0195] The video encoding device derives filter coefficients based on the luminance prediction signal, the offset (B) determined based on the bit depth, the output of a nonlinear function (NL) that uses one or more luminance prediction signals as input, and one or more of the first prediction signals.
[0196] The video encoding device applies filter coefficients to the luminance reconstruction signal to generate a second prediction signal for the current chroma block (S1806).
[0197] The video encoding device applies the derived filter coefficients to the luminance reconstruction sample corresponding to the current chroma sample, the output of the nonlinear function, and the offset to calculate the predicted signal of the current chroma sample.
[0198] The video encoding device calculates weight values (S1808). The video encoding device calculates the weight values of the first prediction signal and the second prediction signal based on the prediction modes of the surrounding blocks of the current chroma block.
[0199] The video encoding device uses weight values to perform a weighted summation of the first prediction signal and the second prediction signal to generate the final prediction signal of the current chroma block (S1810).
[0200] The video encoding device determines a flag indicating whether to perform cross-component prediction based on the first prediction signal, the second prediction signal, and the final prediction signal (S1812).
[0201] From a rate-distortion optimization perspective, the video coding device determines a flag indicating whether to perform cross-component prediction. For example, the flag is determined to be false when the first prediction signal is optimal. Conversely, the flag is determined to be true when the second or final prediction signal is optimal.
[0202] The video encoding device encodes a flag indicating whether to perform cross-component prediction (S1814).
[0203] The video encoding device confirms the flag indicating whether to perform cross-component prediction (S1816).
[0204] When the flag indicating whether to perform cross-component prediction is true ("Yes" in S1816), the video coding device performs the following steps.
[0205] The video encoding device determines a flag indicating whether to perform a weighted summation based on the second prediction signal and the final prediction signal (S1818).
[0206] From a rate-distortion optimization perspective, the video coding device determines a flag indicating whether to perform a weighted summation. For example, the flag is set to false when the second predicted signal is optimal. Conversely, the flag is set to true when the final predicted signal is optimal.
[0207] The video encoding device encodes a flag indicating whether to perform a weighted summation (S1820).
[0208] When the flag indicating whether to perform cross-component prediction is false ("No" in S1816), the video coding device omits the determination and encoding of the flag indicating whether to perform weighted summation.
[0209] Subsequently, the video encoding device generates a residual signal by subtracting the first prediction signal, the second prediction signal, or the final prediction signal from the current chroma block, based on flags indicating whether to perform cross-component prediction and flags indicating whether to perform weighted summation. The video encoding device then transforms / quantizes the residual signal to generate quantized transform coefficients and encodes the quantized transform coefficients.
[0210] Figure 19 This is a flowchart illustrating a method for reconstructing a current chroma block using a video decoding apparatus based on an embodiment of the present disclosure.
[0211] The video decoding device uses motion information to generate a first prediction signal for the current chroma block (S1900). Here, in the case of inter-frame prediction, the motion information is the motion vector of the block or sub-block; in the case of IBC mode, the motion information is the block vector; and in the case of IntraTMP mode, the motion information is the displacement between multiple templates, i.e., the block vector.
[0212] The video decoding device acquires the luminance prediction signal and luminance reconstruction sample of the luminance block in the corresponding luminance area (S1902).
[0213] The video decoding device decodes from the bitstream an indication of whether to perform cross-component prediction (S1904).
[0214] When the luminance block corresponding to the center position of the current chroma block is predicted according to the prediction mode using motion information, the video decoding device decodes the marker.
[0215] When the residual signals of the luminance blocks corresponding to the center position of the current chroma block are all zero, the video decoding device implicitly omits cross-component prediction.
[0216] When determining whether to perform prediction for each transform unit, and when the luminance block corresponding to the center position of the transform unit is predicted according to the prediction mode using motion information, the video decoding device decodes the flag.
[0217] When the residual signals of the luminance blocks corresponding to the center position of the transform unit are all zero, the video decoding device implicitly omits cross-component prediction.
[0218] The video decoding device confirms the prediction pattern at a predetermined position within the corresponding brightness area according to a predetermined order. If there is a brightness block that is predicted based on the prediction pattern using motion information, the marker is decoded.
[0219] The video decoding device confirms the flag indicating whether to perform cross-component prediction (S1906).
[0220] When the flag indicating whether to perform cross-component prediction is false ("No" in S1906), the video decoding device determines the first prediction signal as the final prediction signal for the current chroma block.
[0221] Conversely, when the flag indicating whether to perform cross-component prediction is true ("Yes" in S1906), the video decoding device performs the following steps.
[0222] The video decoding device uses the luminance prediction signal and the first prediction signal to derive the filter coefficients for cross-component prediction of the current chroma block (S1908).
[0223] The video decoding device derives filter coefficients based on the luminance prediction signal, the offset (B) determined based on the bit depth, the output of a nonlinear function (NL) that uses one or more luminance prediction signals as input, and one or more first prediction signals.
[0224] The video decoding device applies filter coefficients to the luminance reconstruction sample to generate a second prediction signal for the current chroma block (S1910).
[0225] The video decoding device applies the derived filter coefficients to the luminance reconstruction sample corresponding to the current chroma sample, the output of the nonlinear function, and the offset to calculate the predicted signal of the current chroma sample.
[0226] The video decoding device decodes the flag indicating whether to perform a weighted summation (S1912).
[0227] The video decoding device confirms the flag indicating whether to perform a weighted summation (S1914).
[0228] When the flag indicating whether to perform weighted summation is false ("No" in S1914), the video decoding device is able to determine the second prediction signal as the final prediction signal for the current chroma block.
[0229] When the flag indicating whether to perform weighted summation is true ("Yes" in S1914), the video decoding device performs the following steps.
[0230] The video decoding device calculates the weight values (S1916). The video decoding device calculates the weight values of the first prediction signal and the second prediction signal based on the prediction modes of the surrounding blocks of the current chroma block.
[0231] The video decoding device uses weight values to perform a weighted summation of the first prediction signal and the second prediction signal to generate the final prediction signal of the current chroma block (S1918).
[0232] The video decoding device decodes the quantized transform coefficients from the bitstream and performs inverse quantization / inverse transform on the quantized transform coefficients to reconstruct the residual signal. The video decoding device adds the final predicted signal to the residual signal to generate the reconstructed signal of the current chroma block.
[0233] exist Figure 20 and Figure 21 The illustration does not use a flag indicating whether a weighted summation is performed on the two predictors; instead, a weighted summation is always performed.
[0234] Figure 20 This is a flowchart illustrating a method for encoding a current chroma block using a video encoding apparatus based on another embodiment of the present disclosure.
[0235] exist Figure 20 In the diagram, the steps from generating the first prediction signal of the current chroma block to generating the final prediction signal of the current chroma block (S2000 to S2010) are... Figure 18 Perform the same procedure. The following explains the procedure in the same way. Figure 18 The diagram illustrates different steps.
[0236] The video encoding device determines a flag indicating whether to perform cross-component prediction based on the first prediction signal and the final prediction signal (S2012).
[0237] From a rate-distortion optimization perspective, the video coding device determines a flag indicating whether to perform cross-component prediction. For example, the flag is set to false when the first predicted signal is optimal. Conversely, the flag is set to true when the final predicted signal is optimal.
[0238] The video encoding device encodes a flag indicating whether to perform cross-component prediction (S2014).
[0239] Subsequently, the video encoding device subtracts the first or final prediction signal from the current chroma block to generate a residual signal, based on a flag indicating whether to perform cross-component prediction. The video encoding device then transforms / quantizes the residual signal to generate quantized transform coefficients and encodes these quantized transform coefficients.
[0240] Figure 21 This is a flowchart illustrating a method for reconstructing a current chroma block using a video decoding apparatus based on another embodiment of the present disclosure.
[0241] exist Figure 21 In the diagram, the steps from generating the first prediction signal of the current chroma block to generating the second prediction signal of the current chroma block (S2100 to S2110) are... Figure 19 Perform the same procedure. The following is unrelated and appears to be a separate instruction: "Nothing is wrong." Figure 19 The diagram illustrates different steps.
[0242] The video decoding device calculates weight values (S2112). Based on the prediction modes of the surrounding blocks of the current chroma block, the video decoding device calculates the weight values of the first prediction signal and the second prediction signal.
[0243] The video decoding device uses weight values to perform a weighted summation of the first prediction signal and the second prediction signal to generate the final prediction signal of the current chroma block (S2114).
[0244] The video decoding device decodes the quantized transform coefficients from the bitstream and performs inverse quantization / inverse transform on the quantized transform coefficients to reconstruct the residual signal. The video decoding device adds the final predicted signal to the residual signal to generate the reconstructed signal of the current chroma block.
[0245] Although the flowcharts / timing diagrams in this specification describe the processes as being executed sequentially, this is merely an illustrative representation of the technical concepts of the embodiments of this disclosure. In other words, those skilled in the art can modify the order of execution in the flowcharts / timing diagrams, or execute one or more processes in parallel, without departing from the essential characteristics of the embodiments of this disclosure, thereby making various modifications and variations and applications. Therefore, the flowcharts / timing diagrams are not limited to a time sequence order.
[0246] It should be understood that the exemplary embodiments described above can be implemented in many different ways. The functions or methods described in one or more examples can be implemented by hardware, software, firmware, or any combination thereof. It should be understood that functional components described in this specification are labeled "...unit" to particularly emphasize their implementation independence.
[0247] On the other hand, the various functions or methods described in this embodiment can also be implemented by one or more instructions stored in a processor-readable and executable non-transitory recording medium. Non-transitory recording media include all kinds of storage devices that store data in a form readable by a computer system. For example, non-transitory recording media include erasable programmable read-only memory (EPROM), flash memory drives, optical disc drives, magnetic hard disk drives, solid-state drives (SSDs), and other recording media.
[0248] The above description is merely an exemplary depiction of the technical concept of this embodiment. Those skilled in the art can make various modifications and variations without departing from the essential characteristics of this embodiment. Therefore, the embodiments disclosed herein are intended to illustrate the technical concept and not to limit it. Any interpretation based on these embodiments should not constitute a limitation on the scope of the technical concept of this embodiment. The scope of protection of this embodiment should be interpreted according to the claims, and all technical concepts within their equivalent scope should be interpreted as being included within the scope of the rights of this embodiment.
[0249] (Explanation of reference numerals in the attached image) 120: Forecasting Department 540: Forecasting Department 902: Prediction Unit Determination Section 904: Prediction Technology Determination Department 906: Prediction Model Determination Department 908: Forecasting and Execution Department 1002: Current block prediction department 1004: CCP Implementation Determination Department 1006: CCP Parameter Determination Department 1008: CCP Executive Department Cross-references to related applications This patent application claims priority to Korean Patent Application No. 10-2023-0062589, filed on May 15, 2023, and Korean Patent Application No. 10-2024-0041180, filed on March 26, 2024, the contents of which are incorporated herein by reference.
Claims
1. A method for reconstructing a current chroma block, performed by a video decoding device, characterized in that, Includes the following steps: A first prediction signal for the current chroma block is generated using motion information, wherein the motion information is a motion vector or a block vector; Obtain the brightness prediction signal and brightness reconstruction sample of the brightness block within the corresponding brightness region; The filter coefficients for cross-component prediction of the current chroma block are derived using the luminance prediction signal and the first prediction signal; and The filter coefficients are applied to the luminance reconstruction sample to generate a second prediction signal for the current chroma block.
2. The method according to claim 1, characterized in that, Further steps include the following: A flag indicating whether to perform the cross-component prediction is used from the bitstream decoding; and Confirm the flag. When the flag is true, the step of deriving the filter coefficients is performed.
3. The method according to claim 2, characterized in that, When the luminance block corresponding to the center position of the current chroma block is predicted according to the prediction mode using the motion information, the step of decoding the flag is performed.
4. The method according to claim 2, characterized in that, When the residual signals of the luminance blocks corresponding to the center position of the current chroma block are all zero, the cross-component prediction is implicitly omitted.
5. The method according to claim 2, characterized in that, When determining whether to perform prediction for each transform unit and the luminance block corresponding to the center position of the transform unit is predicted according to the prediction mode using the motion information, the step of decoding the flag is performed.
6. The method according to claim 5, characterized in that, When the residual signals of the luminance blocks corresponding to the center position of the transformation unit are all zero, the cross-component prediction is implicitly omitted.
7. The method according to claim 2, characterized in that, The prediction pattern of a predetermined position within the corresponding brightness area is confirmed according to a predetermined order. When there is a brightness block predicted according to the prediction pattern using the motion information, the step of decoding the flag is performed.
8. The method according to claim 1, characterized in that, Further steps include the following: Obtain the weight value; and The first prediction signal and the second prediction signal are weighted and summed using the weight values to generate the final prediction signal for the current chroma block.
9. The method according to claim 8, characterized in that, In the step of obtaining the weight value: Use predefined weight values based on the agreement between the video encoding device and the video decoding device.
10. The method according to claim 1, characterized in that, In the step of deriving the filter coefficients: The filter coefficients are derived based on the brightness prediction signal, the offset (B) determined based on the bit depth, the output of a nonlinear function (NL) using one or more brightness prediction signals as input, and one or more of the first prediction signals.
11. The method according to claim 1, characterized in that, In the step of generating the second predicted signal: The filter coefficients are applied to the luminance reconstruction sample corresponding to the current chroma sample, the output of the nonlinear function, and the offset to calculate the predicted signal of the current chroma sample.
12. A method for encoding a current chroma block, performed by a video encoding device, characterized in that, Includes the following steps: A first prediction signal for the current chroma block is generated using motion information, wherein the motion information is a motion vector or a block vector; Obtain the brightness prediction signal and brightness reconstruction sample of the brightness block within the corresponding brightness region; The filter coefficients for cross-component prediction of the current chroma block are derived using the luminance prediction signal and the first prediction signal; and The filter coefficients are applied to the luminance reconstruction sample to generate a second prediction signal for the current chroma block.
13. The method according to claim 12, characterized in that, When the luminance block corresponding to the center position of the current chroma block is predicted according to the prediction mode using the motion information, the step of deriving the filter coefficients is performed.
14. The method according to claim 12, characterized in that, Further steps include the following: Obtaining weights; and The weighted summation of the first prediction signal and the second prediction signal is performed using the weights to generate the final prediction signal for the current chroma block.
15. The method according to claim 14, characterized in that, Further steps include the following: A flag indicating whether to perform the cross-component prediction is determined based on the first prediction signal and the final prediction signal; and The flag indicating whether to perform the cross-component prediction is encoded.
16. A recording medium that stores a bitstream generated by a video encoding method and is readable by a computer, characterized in that, The video encoding method includes the following steps: A first prediction signal for the current chroma block is generated using motion information, wherein the motion information is a motion vector or a block vector; Obtain the brightness prediction signal and brightness reconstruction sample of the brightness block within the corresponding brightness region; Using the luminance prediction signal and the first prediction signal, filter coefficients for cross-component prediction of the current chroma block are derived; and The filter coefficients are applied to the luminance reconstruction sample to generate a second prediction signal for the current chroma block.
Citation Information
Patent Citations
Creating state-aware panoramic images
KR1020230062589A
Provides an interior app that can provide a construction estimate that fits construction practices and applications
KR1020240041180A