Method for extrapolation-based intra prediction

By using an intra-frame prediction method based on extrapolation filters, the prediction block of the current block is generated using the reconstructed reference samples and prediction samples. This solves the problem of insufficient encoding and decoding efficiency for high-resolution and high-frame-rate videos, and improves both video quality and encoding/decoding efficiency.

CN121128179APending Publication Date: 2025-12-12HYUNDAI MOTOR CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202480032771.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-04-16
Filing Date
2024-04-18
Publication Date
2025-12-12

Smart Images

  • Figure CN121128179A_ABST
    Figure CN121128179A_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose a video encoding and decoding method and apparatus using extrapolation-based intra prediction. In an embodiment of the present invention, an image decoding device determines a reference sample region for deriving filter coefficients. Here, the reference sample region includes reconstructed reference samples of the current block. An image decoding device determines a filter shape for extrapolation-based intra prediction. Here, the filter has a rectangular shape and includes a prediction pixel region and an input pixel region. The image decoding device derives a filter coefficient from the reference sample region and the filter shape. The image decoding device generates a prediction block of the current block by applying the derived filter coefficient to the reconstructed reference sample and the prediction sample of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a video encoding and decoding method and apparatus that utilizes extrapolation-based intra-frame prediction. Background Technology

[0002] The statements in this section are merely background information relating to the present invention and do not necessarily constitute prior art.

[0003] Because video data contains a large amount of data compared to audio or still image data, video data requires significant hardware resources (including memory) to store or transmit uncompressed video data.

[0004] Accordingly, encoders are typically used to compress and store or transmit video data. Decoders receive compressed video data, decompress the received compressed video data, and play the decompressed video data. Video compression technologies include H.264 / Advanced Video Codec (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Codec (VVC), which is approximately 30% or more more efficient than HEVC.

[0005] However, as image size, resolution, and frame rate gradually increase, the amount of data to be encoded also increases. Accordingly, new compression techniques are needed that offer higher encoding / decoding efficiency and improved image enhancement compared to existing compression technologies.

[0006] In next-generation technologies, namely the enhanced compression model (ECM), extrapolation-based intra prediction (EIP) mode uses extrapolation filters to perform intra prediction based on neighboring pre-reconstructed reference samples. EIP mode uses reconstructed reference samples and predicted samples as input, performing extrapolation sequentially one pixel at a time. Here, the region of the reference sample and the shape of the extrapolation filter are determined by explicit signaling.

[0007] During intra-frame prediction based on extrapolation filters, EIP mode can use the reconstructed reference sample as input alone to generate the predicted sample. Alternatively, EIP mode can generate the predicted sample by using both the reconstructed reference sample and the predicted signal as input, or by using only the predicted signal as input. In EIP mode, the extrapolation filter uses a single filtering operation to generate a single predicted sample from the input sample.

[0008] Therefore, in order to improve video encoding and decoding efficiency and enhance video quality, an efficient method is needed to utilize reconstructed reference samples and prediction signals during intra-frame prediction based on extrapolation filters. Summary of the Invention

[0009] Technical issues This invention aims to provide a video encoding and decoding method and apparatus that uses an extrapolation filter to effectively predict the current sample from reconstructed reference samples and predicted samples.

[0010] Technical solution At least one aspect of the present invention provides a method for reconstructing a current block using a video decoding apparatus. The method includes determining a reference sample region for deriving filter coefficients. Here, the reference sample region includes reference samples for the reconstruction of the current block. The method further includes determining a filter shape for extrapolation-based intra-frame prediction. Here, the filter shape is rectangular and includes a prediction pixel region and an input pixel region. The method further includes deriving filter coefficients based on the reference sample region and the filter shape. The method further includes generating a prediction block of the current block by applying the derived filter coefficients to the reference samples and prediction samples for the reconstruction of the current block.

[0011] Another aspect of the invention provides a method for encoding a current block using a video coding apparatus. The method includes determining a reference sample region for deriving filter coefficients. Here, the reference sample region includes reference samples for the reconstruction of the current block. The method also includes determining a filter shape for extrapolation-based intra-frame prediction. Here, the filter shape is rectangular and includes a prediction pixel region and an input pixel region. The method further includes deriving filter coefficients based on the reference sample region and the filter shape. The method also includes generating a first prediction block of the current block by applying the derived filter coefficients to the reference samples and prediction samples for the reconstruction of the current block.

[0012] Another aspect of the invention provides a method for providing video data to a video decoding apparatus. The method includes encoding the video data into a bitstream and transmitting the bitstream to the video decoding apparatus. Encoding the video data includes determining a reference sample region for deriving filter coefficients. Here, the reference sample region includes reference samples for the reconstruction of the current block. Encoding the video data also includes determining a filter shape for extrapolation-based intra-frame prediction. Here, the filter shape is rectangular and includes a prediction pixel region and an input pixel region. Encoding the video data further includes deriving filter coefficients based on the reference sample region and the filter shape. Encoding the video data also includes generating a prediction block for the current block by applying the derived filter coefficients to the reference samples and prediction samples for the reconstruction of the current block.

[0013] Beneficial effects As described above, the present invention provides a video encoding / decoding method and apparatus that uses an extrapolation filter to efficiently predict the current sample from reconstructed reference samples and predicted samples. Therefore, the video encoding / decoding method and apparatus improve video encoding / decoding efficiency and enhance video quality. Attached Figure Description

[0014] Figure 1 This is a block diagram of a video encoding device that can implement the technology of this invention.

[0015] Figure 2 This demonstrates a method for partitioning blocks using a quadtree plus binary tree ternary tree (QTBTTT) structure.

[0016] Figure 3a and Figure 3b Multiple intra-prediction modes, including the wide-angle intra-prediction mode, are shown.

[0017] Figure 4 Show the adjacent blocks of the current block.

[0018] Figure 5 This is a block diagram of a video decoding device that can implement the technology of this invention.

[0019] Figure 6 This is a schematic diagram showing the region used for reconstruction based on extrapolation-based intra-frame prediction.

[0020] Figure 7 This is a schematic diagram showing the shape of the extrapolation filter.

[0021] Figure 8 This is a schematic diagram illustrating an ongoing extrapolation-based intra-frame prediction.

[0022] Figure 9 This is a block diagram showing in detail a portion of a video decoding apparatus according to at least one embodiment of the present invention.

[0023] Figure 10 This is a block diagram detailing a predictive actuator according to at least one embodiment of the present invention.

[0024] Figures 11a to 11c This is a schematic diagram illustrating a reference sample region for deriving filter coefficients according to some embodiments of the present invention.

[0025] Figure 12 This is a schematic diagram illustrating the shape of a filter according to at least one embodiment of the present invention.

[0026] Figure 13 This is a schematic diagram illustrating the prediction order of pixels according to at least one embodiment of the present invention.

[0027] Figures 14 to 16This is a schematic diagram illustrating the prediction order of pixels according to another embodiment of the present invention.

[0028] Figure 17 This is a schematic diagram illustrating the shape of a filter according to at least one embodiment of the present invention.

[0029] Figure 18 This is a flowchart of a method for encoding the current block by a video encoding device according to at least one embodiment of the present invention.

[0030] Figure 19 This is a flowchart of a method for reconstructing the current block by a video decoding device according to at least one embodiment of the present invention. Detailed Implementation

[0031] In the following description, some embodiments of the invention will be described in detail with reference to the accompanying illustrative drawings. In the description below, the same reference numerals denote the same elements, although the elements are shown in different drawings. Furthermore, in the following description of some embodiments, detailed descriptions of relevant known components and functions may be omitted for clarity and brevity when it is considered that such detailed descriptions obscure the subject matter of the invention.

[0032] Figure 1 This is a block diagram of a video encoding apparatus that can implement the technology of this invention. In the following text, reference is made to... Figure 1 The diagram illustrates the video encoding apparatus and its components.

[0033] The encoding device may include: an image segmenter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.

[0034] Each component of the encoding device can be implemented as hardware or software, or a combination of hardware and software. Furthermore, the function of each component can be implemented as software, and the microprocessor can be implemented to execute the software functions corresponding to each component.

[0035] A video consists of one or more sequences of images. Each image is segmented into multiple regions, and encoding is performed on each region. For example, an image is segmented into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile and / or slice is segmented into one or more coding tree units (CTUs). Additionally, each CTU is segmented into one or more coding units (CUs) through a tree structure. Information applied to each coding unit (CU) is encoded as the syntax of the CU, and information commonly applied to CUs included in a CTU is encoded as the syntax of the CTU. Furthermore, information commonly applied to all blocks within a slice is encoded as the syntax of the slice header, while information applied to all blocks constituting one or more images is encoded as a Picture Parameter Set (PPS) or picture header. Moreover, information commonly referenced by multiple images is encoded as a Sequence Parameter Set (SPS). Additionally, information commonly referenced by one or more SPSs is encoded as a Video Parameter Set (VPS). Furthermore, information commonly applied to a tile or tile group can also be encoded as the syntax of the tile or tile group header. The syntax included in SPS, PPS, slice headers, and tile or tile group headers can be called high-level syntax.

[0036] Image splitter 110 determines the size of the coding tree unit (CTU). Information about the size of the CTU (CTU dimensions) is encoded into SPS or PPS syntax and transmitted to the video decoding device.

[0037] Image segmenter 110 segments each image constituting the video into multiple coding tree units (CTUs) of a predetermined size, and then recursively segments the CTUs using a tree structure. The leaf nodes in the tree structure become coding units (CUs), which are the basic units of coding.

[0038] A tree structure can be a quadtree (QT), where a higher node (or parent node) is split into four lower nodes (or child nodes) of the same size. A tree structure can also be a binary tree (BT), where a higher node is split into two lower nodes. A tree structure can also be a ternary tree (TT), where a higher node is split into three lower nodes in a 1:2:1 ratio. A tree structure can also be a combination of two or more of the following structures: QT, BT, and TT. For example, a quadtree plus binary tree (QTBT) structure can be used, or a quadtree plus binary tree ternary tree (QTBTTT) structure can be used. Here, a binary tree ternary tree (BTTT) is added to the tree structure to form a multiple-type tree (MTT).

[0039] Figure 2 This is a schematic diagram illustrating a method for segmenting blocks using a QTBTTT structure.

[0040] like Figure 2 As shown, the CTU can first be segmented into a QT structure. Quadtree segmentation can be recursive until the size of the segmented blocks reaches the minimum block size (MinQTSize) allowed for leaf nodes in the QT. The entropy encoder 155 encodes a first flag (QT_split_flag) indicating whether each node of the QT structure is segmented into the four nodes below it and signals this flag to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) allowed for the root node in the BT, the leaf node can be further segmented using at least one of the BT or TT structures. Multiple segmentation directions can exist in the BT and / or TT structures. For example, two directions can exist: a horizontal direction for segmenting the blocks of the corresponding node and a vertical direction for segmenting the blocks of the corresponding node. Figure 2 As shown, when MTT splitting begins, the entropy encoder 155 encodes a second flag (mtt_split_flag) indicating whether a node has been split, as well as a flag indicating the splitting direction (vertical or horizontal) and / or the splitting type (binary or ternary) if the node has been split, and signals this to the video decoding device.

[0041] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node has been split into the four lower-level nodes, the CU split flag (split_cu_flag) indicating whether a node has been split can also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node has not been split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU, which is the basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that each node has been split, the video encoding device begins encoding the first flag using the above scheme.

[0042] When QTBT is used as another example of a tree structure, two types can exist: one where the corresponding node's block is horizontally divided into two blocks of the same size (i.e., symmetrical horizontal division), and the other where the corresponding node's block is vertically divided into two blocks of the same size (i.e., symmetrical vertical division). The entropy encoder 155 encodes a split flag (split_flag) indicating whether each node of the BT structure is divided into blocks of the lower layer, and split type information indicating the split type, and transmits this information to the video decoding device. Alternatively, the type where the corresponding node's block is divided into two blocks that are asymmetrically divided can also exist. Asymmetrical forms can include the form where the corresponding node's block is divided into two rectangular blocks with a size ratio of 1:3, or it can also include the form where the corresponding node's block is divided diagonally.

[0043] The CU can have various sizes depending on the QTBT or QTBTTT segmentation from the CTU. In the following text, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is called the "current block". When using QTBTTT segmentation, the current block can also be rectangular in shape in addition to a square shape.

[0044] Predictor 120 predicts the current block to generate a prediction block. Predictor 120 includes an intra-frame predictor 122 and an inter-frame predictor 124.

[0045] Typically, each block of the current image can be predicted and encoded. This prediction can usually be performed using either intra-frame prediction (which utilizes data from the image containing the current block) or inter-frame prediction (which utilizes data from an image encoded before the image containing the current block). Inter-frame prediction includes both one-way and two-way prediction.

[0046] Intra-predictor 122 predicts pixels in the current block by utilizing pixels (reference pixels) that are adjacent to the current block in the current image, including the current block. Depending on the prediction direction, multiple intra-prediction modes exist. For example, such as... Figure 3aAs shown, multiple intra-frame prediction modes can include two non-directional modes: Planar mode and DC mode, and can include 65 directional modes. The neighboring pixels and algorithm equations to be used are defined differently for each prediction mode.

[0047] To perform efficient orientation prediction for the current block with a rectangular shape, additional methods can be used. Figure 3b The directional modes indicated by the dashed arrows (intra-prediction modes #67 to #80, #-1 to #-14) can be referred to as "wide-angle intra-prediction modes." Figure 3b In the diagram, the arrow indicates the corresponding reference sample used for prediction, not the prediction direction. The prediction direction is opposite to the direction indicated by the arrow. When the current block has a rectangular shape, the wide-angle intra-frame prediction mode is a mode that performs prediction in the opposite direction to a specific directional mode without additional bit transmission. In this case, in the wide-angle intra-frame prediction mode, some wide-angle intra-frame prediction modes available for the current block can be determined by the ratio of the width to the height of the current block with a rectangular shape. For example, when the current block has a rectangular shape with a height less than its width, wide-angle intra-frame prediction modes with angles less than 45 degrees (intra-frame prediction modes #67 to #80) are available. When the current block has a rectangular shape with a width greater than its height, wide-angle intra-frame prediction modes with angles greater than -135 degrees (intra-frame prediction modes #-1 to #-14) are available.

[0048] Intra-predictor 122 can determine the intra-prediction to be used for encoding the current block. In some examples, intra-predictor 122 can encode the current block by utilizing multiple intra-prediction modes, and can also select the appropriate intra-prediction mode to use from test modes. For example, intra-predictor 122 can calculate rate-distortion values ​​by utilizing rate-distortion analysis of multiple test intra-prediction modes, and can also select the intra-prediction mode with the best rate-distortion characteristics from the test modes.

[0049] Intra-predictor 122 selects one intra-prediction mode from multiple intra-prediction modes and predicts the current block by utilizing neighboring pixels (reference pixels) determined according to the selected intra-prediction mode and an algorithmic equation. Entropy encoder 155 encodes the information about the selected intra-prediction mode and transmits it to the video decoding device.

[0050] Inter-frame predictor 124 generates a predicted block for the current block by utilizing motion compensation processing. Inter-frame predictor 124 searches for the most similar block to the current block in a reference image that was encoded and decoded earlier than the current image, and generates the predicted block for the current block using the searched block. Additionally, a motion vector (MV) is generated, corresponding to the displacement between the current block in the current image and the predicted block in the reference image. Typically, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma and chroma components. The motion information, including information from the reference image and information about the motion vector used to predict the current block, is encoded by entropy encoder 155 and transmitted to the video decoding device.

[0051] The inter-frame predictor 124 can also perform interpolation of a reference picture or reference block to increase prediction accuracy. In other words, subsamples are interpolated between two consecutive integer samples by applying filter coefficients to multiple consecutive integer samples comprising two integer samples. When performing the process of searching for the most similar block to the current block on the interpolated reference picture, the motion vector can be represented with fractional-unit precision instead of integer-sample-unit precision. For each target region to be encoded, such as a cell like a slice, tile, CTU, CU, etc., the precision or resolution of the motion vector can be set differently. When this adaptive motion vector resolution (AMVR) is applied, information about the motion vector resolution to be applied to each target region should be signaled for each target region. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution can be information representing the precision of the motion vector difference to be described below.

[0052] On the other hand, the inter-frame predictor 124 can perform inter-frame prediction by utilizing bidirectional prediction. In the case of bidirectional prediction, two reference images and two motion vectors representing the positions of the blocks most similar to the current block in each reference image are used. The inter-frame predictor 124 selects a first reference image and a second reference image from reference image list 0 (RefPicList0) and reference image list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the block most similar to the current block in the corresponding reference image to generate a first reference block and a second reference block. Furthermore, the predicted block for the current block is generated by averaging or weighted averaging the first reference block and the second reference block. In addition, motion information including information about the two reference images used to predict the current block and information about the two motion vectors is transmitted to the entropy encoder 155. Here, reference image list 0 may consist of images in the pre-reconstructed images that precede the current image in display order, and reference image list 1 may consist of images in the pre-reconstructed images that follow the current image in display order. However, although not particularly limited to this, pre-reconstructed images that follow the current image in display order may be additionally included in reference image list 0. Conversely, pre-reconstructed images preceding the current image can also be additionally included in the reference image list 1.

[0053] Various methods can be used to minimize the number of bits consumed in encoding motion information.

[0054] For example, when the reference image and motion vector of the current block are the same as those of the adjacent blocks, the information that can identify the adjacent blocks is encoded to transmit the motion information of the current block to the video decoding device. This method is called merge mode.

[0055] In merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidate blocks (hereinafter referred to as "merge candidates") from the neighboring blocks of the current block.

[0056] As adjacent blocks used to derive merge candidates, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image can be used, such as Figure 4 As shown. In addition to the current image containing the current block, blocks located within a reference image (which may be the same as or different from the reference image used to predict the current block) can also be used as merge candidates. For example, a co-located block of the current block within the reference image, or a block adjacent to that co-located block, can be additionally used as a merge candidate. If the number of merge candidates selected using the above method is less than a preset number, a zero vector is added to the merge candidates.

[0057] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates by utilizing neighboring blocks. It selects a merge candidate from the merge candidates included in the merge list to be used as motion information for the current block, and generates merge index information for identifying the selected candidate. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding device.

[0058] The merge skip mode is a special case of the merge mode. After quantization, when all transform coefficients used for entropy coding are close to zero, only the adjacent block selection information is transmitted, and the residual signal is not transmitted. By utilizing the merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, and screen content images.

[0059] From then on, merge mode and merge skip mode were collectively referred to as merge / skip mode.

[0060] Another approach for encoding motion information is the Advanced Motion Vector Prediction (AMVP) model.

[0061] In AMVP mode, the inter-frame predictor 124 derives motion vector prediction candidates for the current block by utilizing neighboring blocks of the current block. As neighboring blocks used to derive motion vector prediction candidates, they can use... Figure 4 The current block in the current image shown includes all or some of the adjacent blocks: left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2. Furthermore, blocks located in a reference image (which may be the same as or different from the reference image used to predict the current block) can also be used as neighboring blocks for deriving motion vector prediction candidates, in addition to the current image containing the current block. For example, a co-located block of the current block in the reference image or a block adjacent to that co-located block can be used. If the number of motion vector candidates selected by the above method is less than a preset number, a zero vector is added to the motion vector candidates.

[0062] The inter-frame predictor 124 derives motion vector prediction candidates by utilizing motion vectors from neighboring blocks, and determines the motion vector prediction for the current block by utilizing these candidates. Furthermore, the motion vector difference is calculated by subtracting the motion vector prediction from the motion vector prediction for the current block.

[0063] Motion vector predictions can be obtained by applying predefined functions (e.g., median and average calculations) to motion vector prediction candidates. In this case, the video decoding device is also aware of the predefined functions. Furthermore, since the neighboring blocks used to derive the motion vector prediction candidates are already encoded and decoded, the video decoding device may already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information used to identify the motion vector prediction candidates. Accordingly, in this case, information about the motion vector difference and information about the reference image used to predict the current block are encoded.

[0064] Alternatively, motion vector prediction can be determined by selecting any one of the motion vector prediction candidates. In this case, the information used to identify the selected motion vector prediction candidate is further encoded together with the information about the motion vector difference used to predict the current block and the information about the reference image.

[0065] Subtractor 130 generates a residual block by subtracting the predicted block generated by intra-predictor 122 or inter-predictor 124 from the current block.

[0066] Transformer 140 transforms the residual signal in the residual block, which has pixel values ​​in the spatial domain, into transform coefficients in the frequency domain. Transformer 140 can transform the residual signal in the residual block by using the entire size of the residual block as a transform unit, or it can divide the residual block into multiple sub-blocks and perform the transformation by using the sub-blocks as transform units. Alternatively, the residual block is divided into two sub-blocks, namely a transformed region and a non-transformed region, so that the residual signal is transformed by using only the transformed region sub-block as a transform unit. Here, the transformed region sub-block can be one of two rectangular blocks with a size ratio of 1:1 based on a horizontal axis (or a vertical axis). In this case, the entropy encoder 155 encodes a flag (cu_sbt_flag) indicating that only the transformed sub-block is used, as well as orientation (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or position information (cu_sbt_pos_flag), and signals this information to the video decoding device. Additionally, the size of the transformed region sub-blocks can have a 1:3 size ratio based on the horizontal axis (or vertical axis). In this case, the entropy encoder 155 additionally encodes the flag (cu_sbt_quad_flag) for dividing the corresponding segment and signals it to the video decoding device.

[0067] On the other hand, the transformer 140 can perform the transformation of the residual block separately in the horizontal and vertical directions. Various types of transformation functions or transformation matrices can be used for this transformation. For example, the pair of transformation functions used for horizontal and vertical transformations can be defined as a multiple transform set (MTS). The transformer 140 can select the transform function pair with the highest transformation efficiency in the MTS and can transform the residual block in each of the horizontal and vertical directions. The entropy encoder 155 encodes the information (mts_idx) about the transform function pair in the MTS and signals it to the video decoding device.

[0068] Quantizer 145 quantizes the transform coefficients output from transformer 140 using quantization parameters and outputs the quantized transform coefficients to entropy encoder 155. Quantizer 145 can also quantize the relevant residual blocks immediately without transforming any block or frame. Quantizer 145 can also apply different quantization coefficients (scaling values) based on the position of the transform coefficients in the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions can be encoded and signaled to the video decoding device.

[0069] The rearrangement unit 150 can rearrange the coefficient values ​​of the quantized residual values.

[0070] The rearrangement unit 150 can transform a 2D coefficient array into a 1D coefficient sequence by utilizing coefficient scanning. For example, the rearrangement unit 150 can use a zig-zag scan or a diagonal scan to scan the DC coefficients to the high-frequency region to output a 1D coefficient sequence. Depending on the size of the transform unit and the intra-frame prediction mode, a vertical scan scanning the 2D coefficient array in the column direction and a horizontal scan scanning the 2D block-type coefficients in the row direction can also be used instead of a zig-zag scan. In other words, the scanning method to be used can be determined from zig-zag scan, diagonal scan, vertical scan, and horizontal scan, depending on the size of the transform unit and the intra-frame prediction mode.

[0071] The entropy encoder 155 encodes the sequence of 1D quantized transform coefficients output from the rearrangement unit 150 by utilizing various encoding schemes, including Context-based Adaptive Binary Arithmetic Code (CABAC) and Exponential Golomb, to generate a bitstream.

[0072] Furthermore, the entropy encoder 155 encodes block-segmentation related information (e.g., CTU size, CTU segmentation flag, QT segmentation flag, MTT segmentation type, and MTT segmentation direction) so that the video decoding device can segment blocks equivalently to the video encoding device. Additionally, the entropy encoder 155 encodes information indicating whether the current block is intra-frame predictive coding or inter-frame predictive coding. The entropy encoder 155 encodes intra-frame prediction information (i.e., information about the intra-frame prediction mode) or inter-frame prediction information (the merge index in the case of merge mode, and information about the reference picture index and motion vector difference in the case of AMVP mode) according to the prediction type. Furthermore, the entropy encoder 155 encodes quantization-related information (i.e., information about quantization parameters and information about the quantization matrix).

[0073] Inverse quantizer 160 inverse quantizes the quantized transform coefficients output from quantizer 145 to generate transform coefficients. Inverse transformer 165 transforms the transform coefficients output from inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct the residual block.

[0074] Adder 170 adds the reconstructed residual block to the prediction block generated by predictor 120 to reconstruct the current block. When performing intra-frame prediction for the next block, the pixels in the reconstructed current block are used as reference pixels.

[0075] The loop filtering unit 180 performs filtering on the reconstructed pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc., that occur due to block-based prediction and transform / quantization. The loop filtering unit 180, as an in-loop filter, may include all or some of the following: a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.

[0076] Deblocking filter 182 filters the boundaries between reconstructed blocks to remove blocking artifacts caused by block unit encoding / decoding, and SAO filter 184 and ALF 186 perform additional filtering on the deblocked video. SAO filter 184 and ALF 186 are filters used to compensate for the difference between reconstructed pixels and original pixels caused by lossy coding. SAO filter 184 applies an offset in CTU units to enhance subjective image quality and coding efficiency. On the other hand, ALF 186 performs block unit filtering and applies different filters to compensate for distortion by dividing the boundaries of the corresponding blocks and the degree of variation. Information about the filter coefficients to be used for ALF can be encoded and signaled to the video decoding device.

[0077] The reconstructed blocks filtered by deblocking filter 182, SAO filter 184, and ALF 186 are stored in memory 190. When all blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter-frame prediction of blocks in subsequent images to be encoded.

[0078] Video encoding devices can store the bitstream of encoded video data in a non-volatile storage medium or send the bitstream to a video decoding device via a communication network.

[0079] Figure 5 This is a functional block diagram of a video decoding device that can implement the technology of this invention. In the following text, reference is made to... Figure 5 It describes the video decoding device and its components.

[0080] The video decoding device may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transformer 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.

[0081] Similar to Figure 1 The video encoding device and the video decoding device each component can be implemented as hardware or software, or a combination of hardware and software. Furthermore, the function of each component can be implemented as software, and the microprocessor can also be implemented to execute the software functions corresponding to each component.

[0082] The entropy decoder 510 extracts information related to block segmentation by decoding the bitstream generated by the video encoding device to determine the current block to be decoded, and extracts the prediction information and information about the residual signal required to reconstruct the current block.

[0083] The entropy decoder 510 determines the size of the CTU by extracting information about the CTU size from the sequence parameter set (SPS) or the picture parameter set (PPS) and segments the image into CTUs of a defined size. Furthermore, the CTU is identified as the highest level (i.e., the root node) of the tree structure, and segmentation information of the CTU can be extracted to segment the CTU by utilizing the tree structure.

[0084] For example, when segmenting a CTU using a QTBTT structure, the first flag (QT_split_flag) related to the QT segmentation is first extracted to split each node into four lower-level nodes. Additionally, a second flag (mtt_split_flag), segmentation direction (vertical / horizontal), and / or segmentation type (binary / ternary) related to the MTT segmentation are extracted relative to the nodes corresponding to the leaf nodes of the QT to segment the corresponding leaf nodes using the MTT structure. As a result, each node below the leaf node of the QT is recursively segmented using either a BT or TT structure.

[0085] As another example, when splitting a CTU using a QTBTTT structure, a CU split flag (split_cu_flag) indicating whether a CU has been split is extracted. A first flag (QT_split_flag) can also be extracted when the corresponding block is split. During the splitting process, for each node, zero or more recursive MTT splits may occur after zero or more recursive QT splits. For example, for a CTU, MTT splits may occur immediately, or conversely, only multiple QT splits may occur.

[0086] As another example, when segmenting a CTU using a QTBT structure, the first flag (QT_split_flag) associated with the QT segmentation is extracted to split each node into four nodes at the lower level. Additionally, a split flag (split_flag) indicating whether the node corresponding to a leaf node of the QT will be further segmented using BT, as well as segmentation direction information, is extracted.

[0087] On the other hand, when the entropy decoder 510 determines the current block to be decoded by utilizing the segmentation of the tree structure, the entropy decoder 510 extracts information about the prediction type indicating whether the current block is predicted intra-frame or inter-frame. When the prediction type information indicates intra-frame prediction, the entropy decoder 510 extracts the syntax elements for the intra-frame prediction information (intra-frame prediction mode) for the current block. When the prediction type information indicates inter-frame prediction, the entropy decoder 510 extracts information representing the syntax elements of the inter-frame prediction information, namely, the motion vector and the reference picture of the motion vector reference.

[0088] In addition, the entropy decoder 510 extracts quantization-related information and extracts information about the quantization transform coefficients of the current block as information about the residual signal.

[0089] The rearrangement unit 515 can re-transform the sequence of 1D quantized transform coefficients entropy-decoded by the entropy decoder 510 into a 2D coefficient array (i.e., a block) in the reverse order of the coefficient scan order performed by the video encoding device.

[0090] The inverse quantizer 520 performs inverse quantization on the quantized transform coefficients, and performs inverse quantization on the quantized transform coefficients using quantization parameters. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video encoding device to a 2D array of quantized transform coefficients.

[0091] The inverse transformer 530 reconstructs the residual signal by inversely transforming the inverse-quantized transform coefficients from the frequency domain to the spatial domain to generate the residual block of the current block.

[0092] Furthermore, when the inverse transformer 530 performs an inverse transform on a portion (sub-block) of the transform block, it extracts a flag (cu_sbt_flag) indicating that only the sub-blocks of the transform block are transformed, the sub-block's orientation (vertical / horizontal) information (cu_sbt_horizontal_flag), and / or the sub-block's position information (cu_sbt_pos_flag). The inverse transformer 530 also inversely transforms the transform coefficients of the corresponding sub-blocks from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the untransformed regions with values ​​"0" as the residual signal to generate the final residual block for the current block.

[0093] Furthermore, when applying MTS, the inverse transformer 530 determines the transform function or transform matrix to be applied in each of the horizontal and vertical directions by utilizing the MTS information (mts_idx) signaled from the video encoding device. The inverse transformer 530 also performs inverse transform on the transform coefficients in the transform block in both the horizontal and vertical directions using the determined transform function.

[0094] Predictor 540 may include intra-predictor 542 and inter-predictor 544. Intra-predictor 542 is activated when the prediction type of the current block is intra-prediction, and inter-predictor 544 is activated when the prediction type of the current block is inter-prediction.

[0095] Intra-predictor 542 determines the intra-prediction mode for the current block from among multiple intra-prediction modes based on the grammatical elements of the intra-prediction mode extracted from entropy decoder 510. Intra-predictor 542 also predicts the current block by utilizing neighboring reference pixels of the current block based on the intra-prediction mode.

[0096] The inter-frame predictor 544 determines the motion vector and the reference picture of the motion vector reference for the current block by utilizing the grammatical elements of the inter-frame prediction mode extracted from the entropy decoder 510, and predicts the current block by utilizing the motion vector and the reference picture.

[0097] Adder 550 reconstructs the current block by adding the residual block output from inverse transform 530 to the predicted block output from inter-frame predictor 544 or intra-frame predictor 542. Pixels within the reconstructed current block are used as reference pixels when performing intra-frame prediction for subsequent blocks to be decoded.

[0098] The loop filtering unit 560, acting as an in-loop filter, may include a deblocking filter 562, a SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between reconstructed blocks to remove block artifacts caused by block unit decoding. The SAO filter 564 and ALF 566 perform additional filtering on the reconstructed blocks after deblocking to compensate for differences between reconstructed pixels and original pixels caused by lossy encoding. The filter coefficients of the ALF are determined by utilizing information about the filter coefficients decoded from the bitstream.

[0099] The reconstructed blocks filtered by deblocking filter 562, SAO filter 564, and ALF 566 are stored in memory 570. When all blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter-frame prediction of blocks in subsequent images to be encoded.

[0100] In some embodiments, the present invention relates to encoding and decoding video images as described above. More specifically, the present invention provides a video encoding / decoding method and apparatus that uses an extrapolation filter to perform efficient intra-frame prediction of the current sample from reconstructed reference samples and predicted samples.

[0101] The following implementation scheme can be performed by the intra-frame predictor 122 in the video encoding apparatus. The following implementation scheme can also be performed by the intra-frame predictor 542 in the video decoding apparatus.

[0102] When encoding the current block, the video encoding apparatus can generate signaling information associated with this embodiment from the perspective of optimizing rate distortion. The video encoding apparatus can encode the signaling information using the entropy encoder 155 and send the encoded signaling information to the video decoding apparatus. The video decoding apparatus can decode the signaling information associated with the decoding of the current block from the bitstream using the entropy decoder 510.

[0103] In the following description, the term "target block" may be used interchangeably with the current block or coding unit (CU), or may refer to some region of the coding unit.

[0104] In addition, a flag value of true indicates that the flag is set to 1. Conversely, a flag value of false indicates that the flag is set to 0.

[0105] I. Extrapolation-based Intra-Frame Prediction (EIP) In next-generation technology, the enhanced compression model (ECM), extrapolation-based intra prediction mode (EIP mode) performs intra prediction by utilizing an extrapolation filter based on neighboring pre-reconstructed reference samples. EIP mode uses reconstructed reference samples and prediction samples as input, performing extrapolation sequentially one pixel at a time. The reference sample region used to derive the extrapolation filter coefficients is located in... Figure 6 As shown in the figure, the shape of the extrapolation filter used for prediction is in Figure 7 As shown in the diagram. Here, the region of the reference sample and the shape of the extrapolation filter are determined by explicit signaling.

[0106] like Figure 8 As shown, intra-frame prediction can be performed based on an extrapolation filter. For example, EIP mode generates predicted samples by specifically utilizing reconstructed reference samples as input. Alternatively, EIP mode can generate predicted samples by utilizing both reconstructed reference samples and the predicted signal as input, or by utilizing only the predicted signal as input. In EIP mode, the extrapolation filter generates a single predicted sample from the input samples using a single filtering operation.

[0107] The following implementations are described with a focus on video decoding devices, but they can be implemented in the same or similar ways in video encoding devices.

[0108] II. Embodiments according to the present invention According to some implementations, a video decoding apparatus can determine prediction and transform units, and in response to the current block corresponding to the determined units, can perform prediction and inverse transform by utilizing determined prediction techniques and prediction modes to ultimately generate a reconstructed block of the current block. This can be performed by the inverse transformer 530, predictor 540, and adder 550 of the video decoding apparatus. Figure 9 The operation shown. On the other hand, the operation can be performed by the inverse transformer 165, image segmenter 110, predictor 120, and adder 170 in the video encoding apparatus. Figure 9 The same operation is shown in the diagram. In this case, the video decoding device uses the encoded information parsed from the bitstream, but the video encoding device can use encoded information set at a higher level to minimize bit rate distortion. For convenience, the following description focuses on the video decoding device.

[0109] like Figure 5 As shown, according to the prediction technique, predictor 540 includes intra-frame predictor 542 and inter-frame predictor 544, but as... Figure 9 As shown, the predictor 540 may include all or part of the prediction unit determiner 902, the prediction technique determiner 904, the prediction mode determiner 906, and the prediction executor 908.

[0110] When the input video's color format is YUV (YUV420, YUV411, YUV422, YUV444, etc.), the video decoding device can perform prediction and reconstruction of the luminance component, and then perform prediction and reconstruction of the chrominance component. In this way, the luminance and chrominance components can be determined by... Figure 9 The components shown are reconstructed sequentially. On the other hand, when the input video's color format is RGB, the video encoding device can perform a color format conversion from RGB to YUV, and then encode the converted video. Here, in YUV format, the color format represents the association between pixels in the luminance component and pixels in the chrominance component.

[0111] Prediction unit determiner 902 determines the prediction unit (PU). Prediction technique determiner 904 determines the prediction technique in response to the prediction unit, such as intra-frame prediction, inter-frame prediction, intra-block copy (IBC) mode, palette mode, etc. Prediction mode determiner 906 determines the detailed prediction mode used for the prediction technique. Prediction executor 908 generates the prediction block for the current block based on the determined prediction mode.

[0112] The inverse transformer 530 includes a transformer unit determiner 910 and an inverse transformer actuator 912. The transformer unit determiner 910 determines the transformer unit (TU) in response to the dequantized signal of the current block, and the inverse transformer actuator 912 performs an inverse transformer on the transformer unit represented by the dequantized signal to generate a residual signal.

[0113] Adder 550 sums the predicted block and the residual signal to generate a reconstructed block. The reconstructed block is stored in memory and can be used to predict other future blocks.

[0114] The prediction unit determined by prediction unit determiner 902 can be a sub-block of the current block or a sub-block segmented from the current block. In this case, depending on the color format, the prediction unit for the chroma component can correspond in size to the prediction unit for the luminance component. Alternatively, prediction units for the luminance component and chroma component can be determined separately, and prediction can be performed in response to the prediction unit for the chroma component.

[0115] Prediction technique determiner 904 determines the prediction technique used by the prediction unit. As described above, the prediction technique can be one of inter-frame prediction, intra-frame prediction, IBC mode, and palette mode. In this case, the prediction technique used for the chroma component can be determined to be the same as the prediction technique for the corresponding luma component without emitting and parsing any additional information.

[0116] In one example, when the prediction technique for the current block is not intra-frame prediction, the video decoder parses a 1-bit flag. When the parsed flag indicates a skip mode, the video decoder can determine the prediction mode for the current block as either inter-frame prediction combining mode or IBC combining mode. In skip mode, the video decoder can also use the predicted signal instead of the reconstructed signal, while skipping the inverse transform; that is, there is no need to parse the residual signal.

[0117] On the other hand, when the parsed flag does not indicate a skip mode in response to the current block, the prediction technique determiner 904 can determine one of the prediction techniques by parsing a series of 1-bit flags as the prediction technique for the current block, such as inter-frame prediction, intra-frame prediction, IBC mode, palette mode, etc.

[0118] For example, when no skip mode is applied to the current block and inter-frame prediction or IBC mode is determined as the prediction technique, the video decoder parses a 1-bit flag. The video decoder can then determine the prediction mode for the current block as either regular merging mode or advanced motion vector prediction (AMVP) mode based on the parsed flag.

[0119] The prediction pattern determiner 906 determines the detailed prediction pattern used in the prediction technique.

[0120] As an example, when the prediction technique used for the current block is intra-frame prediction, the prediction mode determiner 906 can determine a single directional prediction mode, a planar mode (horizontal plane, vertical plane, or regular plane), or a DC mode as the prediction mode for the current block.

[0121] As another example, when the prediction technique used for the current block is intra-prediction, the prediction mode determiner 906 can determine the prediction mode of the current block as a mode that performs prediction based on a predefined matrix. The prediction mode determiner 906 can compute gradients based on reconstructed reference samples and use the computed gradients to determine the prediction mode of the current block as a mode for deriving the intra-prediction mode. The prediction mode determiner 906 can perform intra-prediction based on reference samples and then compute a cost to determine the intra-prediction mode of the current block as a mode for deriving the prediction mode with the minimum cost. Here, a metric such as the Sum of Absolute Differences (SAD) or the Sum of Absolute Transformed Differences (SATD) can be used as the cost. Alternatively, the prediction mode determiner 906 can use a reference sample region to determine the prediction mode of the current block as a mode for deriving the positions and combinations of reference sample lines based on the cost (SAD or SATD) to form the prediction mode for intra-prediction. In this case, the selection of which prediction mode to use can be explicitly determined based on the flags. The signaling order of the flags for the corresponding prediction mode can vary depending on the implementation scheme. In this case, the video decoding device can additionally parse the flags indicating whether to perform extrapolation-based intra-frame prediction to determine whether to perform extrapolation-based intra-frame prediction.

[0122] For example, extrapolation-based intra-prediction may not be performed if the aspect ratio (W / H) of the current block is greater than or less than a certain threshold. Here, W and H represent the width and height of the current block, respectively. Extrapolation-based intra-prediction may not be performed when W × H is greater than or less than the certain threshold. Alternatively, extrapolation-based intra-prediction may not be performed depending on factors such as the intra-prediction mode containing some or all blocks of a reference sample or whether a reference sample region is used. In the above cases, when extrapolation-based intra-prediction is not performed, the video decoding device may omit resolving the flag indicating whether to perform extrapolation-based intra-prediction.

[0123] The prediction executor 908 generates a prediction block for the current block based on the determined prediction technique and prediction pattern.

[0124] For example, the predictive executor 908 generates a prediction block for the current block based on the prediction mode, and the adder 550 sums the prediction block of the current block and the residual signal to generate a reconstructed block.

[0125] Figure 10 This is a block diagram showing in detail a predictive actuator according to at least one embodiment of the present invention.

[0126] For cross component prediction, such as Figure 10 As shown, the prediction actuator 908 includes a reference sample region determiner 1002, a filter shape determiner 1004, a filter coefficient determiner 1006, and an extrapolation actuator 1008.

[0127] The reference sample region determiner 1002 explicitly resolves the reference sample region by deriving filter coefficients based on the reconstructed reference samples, for use in deriving filter coefficients when performing extrapolation-based intra-frame prediction. For example... Figures 11a to 11c As shown, the video decoding device uses the left and top sides (first region), the left side (second region), or the top side (third region) of the current block as a reference sample region. The video decoding device parses information indicating one of the first to third regions to determine the reference sample region. Figures 11a to 11c In the diagram, A, B, C, and D are integer values ​​greater than or equal to zero, and can be defined according to the protocol between the video encoding device and the video decoding device.

[0128] When a reference sample is not available at a specific location relative to the current block, the video decoding apparatus derives filter coefficients that exclude the relevant region. Alternatively, the video decoding apparatus can use a reference sample that includes samples generated by filling in specific locations to derive the filter coefficients.

[0129] As an example, when a specific reference sample region is used solely based on the protocol between the video encoding and decoding devices, the video decoding device can skip additional parsing and determine that specific region as the reference sample region for deriving filter coefficients.

[0130] In another example, when filter coefficients defined according to the protocol between the video encoding device and the video decoding device are used for extrapolation-based intra-frame prediction, the operation of the reference sample region determiner 1002 can be omitted.

[0131] The filter shape determiner 1004 determines the filter shape used to perform extrapolation-based intra-frame prediction. The filter shape determiner 1004 determines M, N, K, and L, which are natural numbers, such as... Figure 12 As shown. In Figure 12In the example, the rectangular shape K × L represents the size of the predicted pixel region, and M × NK × L represents the size of the input pixel region. Furthermore, M ≥ K and N ≥ L.

[0132] As an example, a video decoding device can explicitly resolve one of T filter shapes defined according to the protocol between the video encoding device and the video decoding device. Here, T represents any positive integer greater than 0. In the T filters, each filter has a different value for one or more of M, N, K, and L.

[0133] For example, the reference sample region and filter shape used for filter derivation can be defined according to the protocol between the video encoding device and the video decoding device. The video decoding device can parse an index indicating such a combination and use the parsed index to determine the reference sample region and filter shape.

[0134] For example, for each of T filter shapes, the video decoding device derives filter coefficients from a reconstructed reference sample region and then generates a prediction signal for the current block corresponding to the filter shape and filter coefficients. Here, the filter coefficients are calculated by a filter coefficient determiner 1006, and the prediction signal for the current block can be generated by an extrapolation executor 1008. The video decoding device can calculate the cost (e.g., sum of absolute differences or SAD) between the prediction signal for the current block and the reconstructed reference sample of the current block, and determine the filter shape with the minimum computational cost for predicting the current block. The region for calculating the cost between the prediction signal and the reconstructed value can be the same as the region determined by the reference sample region determiner 1002. Alternatively, the region for calculating the cost can be a specific region defined according to a protocol between the video encoding device and the video decoding device (e.g., ...). Figure 11a The first region shown in the diagram.

[0135] For example, the filter shape can be determined based on the size of the current block, the aspect ratio of the current block, the previous protocol between the video encoding and decoding devices, and one or more explicit signaling.

[0136] Extrapolation-based intra-frame prediction uses previously predicted samples as input, thus creating dependencies. For example, when K=L=1, the video decoding device uses... Figure 13 The predictions are executed in the order shown. Figure 13 In the example, pixels with the same reconstruction order represent pixels that can be predicted in parallel based on dependencies.

[0137] The number of pipelines based on the above dependencies is determined by the number of pixels (K × L) of the extrapolated predicted output. When applying the same number of predicted pixels, the number of pipelines can be determined based on the aspect ratio (W / H) of the current block and the aspect ratio (K / L) of the predicted output pixels.

[0138] To limit the maximum number of pipes based on dependencies, the values ​​of K and L can be determined. In this case, K and L can be determined based on the width, height, and aspect ratio of the current block as follows.

[0139] When the aspect ratio of the current block is greater than 1, a first filter shape is determined ((K / L)>1). When the aspect ratio of the current block is less than 1, a second filter shape is determined ((K / L)<1). When the aspect ratio of the current block is 1, a specific filter shape can be determined as a first filter shape or a second filter shape as defined by the protocol between the video encoding device and the video decoding device, or the filter shape can be explicitly determined based on flags. Figure 14 As shown, a video decoding device can predict blocks with an aspect ratio greater than 1 by utilizing the shape of a first filter. For example... Figure 15 As shown, a video decoding device can predict blocks with an aspect ratio less than 1 by utilizing the shape of a second filter. For example... Figure 16 As shown, a video decoding device can predict a block with an aspect ratio of 1 by utilizing either a first filter shape or a second filter shape.

[0140] As an example, when the aspect ratio of the current block is 1, the video decoding device can predict the current block by utilizing a single filter shape of K=L.

[0141] For example, for a block undergoing intra-frame prediction based on extrapolation, the video decoding apparatus can use a specific filter shape defined by a protocol between the video encoding and decoding apparatus, depending on the block size. Alternatively, the video decoding apparatus can define a list of candidate filter shapes based on the block size and then explicitly determine the filter shape to be used for prediction. That is, the video decoding apparatus can parse an index indicating a candidate filter shape and then use the parsed index to determine the filter shape from the list of candidate filter shapes.

[0142] The filter coefficient determiner 1006 derives the filter coefficients for prediction based on a reference sample region and the filter shape. For example, when using filter coefficients defined for extrapolation-based intra-frame prediction based on a protocol between the video encoding and decoding devices, the present invention can skip the filter coefficient derivation process of the filter coefficient determiner 1006. The video decoding device can parse the index indicating the filter coefficients, and then use the parsed index to determine the filter coefficients to be used for prediction.

[0143] Filter coefficients can be derived using a common method employed by both the video encoding and decoding devices. For example, the video decoding device derives filter coefficients that minimize costs (e.g., SAD, mean square error, etc.) between the reconstructed reference sample and the values ​​generated by multiplying the filter coefficients by the values ​​of the reconstructed reference sample, defined according to the filter shape in the reference sample region. When the filter shape is as follows... Figure 17 As shown, in order to generate predicted signals O0 to O7 at their respective positions from the reconstructed reference samples of 1 × 7 (I0 to I6), the video decoding device derives a 7 × 8 filter matrix. For example, to derive the above filter matrix, the video decoding device can derive the filter coefficients C0 to C6 at the corresponding prediction positions from the perspective of least squares (LS) optimization. Figure 9 The filter shape, and the filter coefficients used to generate the predicted value at the i-th output position, can be defined by Equation 1.

[0144] [Equation 1] In Equation 1, the filter coefficients C can be derived using methods such as Gaussian elimination, LDL decomposition, or Cholesky decomposition.

[0145] For example, when deriving filter coefficients, the video decoding device can subtract an offset from the input reference samples and the output reference samples, and then use the input reference samples and the output reference samples subtracted from the offset to derive the filter coefficients. Here, the offset can be the average value of the reference samples, the median value of the reference samples, or the median value determined based on the bit depth. As shown in Equation 2, when deriving the filter coefficients after subtracting the offset B, the video decoding device can additionally derive the filter coefficient C7 used for the offset.

[0146] [Equation 2] In Equation 1 or Equation 2, I and O represent reference samples for the reconstruction of adjacent blocks. Furthermore, in I... i,j and O i,j In this context, i represents the position of the input sample and the position of the output sample, respectively, and j represents the index of the matching pair between the input and the output, which is determined, for example, by the sliding window method.

[0147] The extrapolation executor 1008 applies the determined filter coefficients to the reference and prediction samples for the reconstruction of the current block to perform extrapolation-based intra-frame prediction, thereby generating the prediction block for the current block.

[0148] based on Figure 12Given the example and a defined filter shape, the video decoding device can multiply the defined filter matrix by the number of input samples, M × NK × L, to generate K × L predicted samples. At this point, as shown... Figures 13 to 16 As shown in the example, video decoding devices can predict predictable regions at the same point in time in parallel based on their dependencies.

[0149] As an example, when deriving filter coefficients while taking offset into account, the video decoding device can subtract the offset from the input sample and then apply the filter coefficients to the input sample minus the offset, thereby generating the final predicted signal. As mentioned above, the offset can be the average value of the reference samples, the median value of the reference samples, or the median value determined based on the bit depth. In this case, the average value can be calculated based on adjacent pre-reconstructed reference samples. For example, the video decoding device can calculate the average value of the reference samples by averaging the reconstructed reference samples within one or more reference sample lines in a neighboring reference sample region.

[0150] In some implementations, a limiting mechanism can be used to ensure that the extrapolated output O exists within a specific threshold range (i.e., satisfying TH0 ≤ O ≤ TH1). Some implementations can determine that the threshold TH0 is the minimum among the neighboring reference samples, and the threshold TH1 is the maximum among the neighboring reference samples.

[0151] Regarding the prediction of the current chroma block, when the prediction mode of the chroma block is DM mode and the corresponding luma block is predicted using extrapolation-based intra-frame prediction, the video decoding device can predict the chroma block using extrapolation-based intra-frame prediction. Alternatively, the chroma block can be predicted using any intra-frame prediction mode (e.g., DC, planar mode, etc.). When predicting the current chroma block using extrapolation-based intra-frame prediction, the video decoding device can use the filter shape and filter coefficients calculated from the corresponding luma block. Alternatively, after obtaining the filter shape from the corresponding luma block, the video decoding device can derive the filter coefficients based on the reconstructed chroma reference samples and can use the derived filter coefficients to predict the current chroma block.

[0152] As described above, the video decoding device reconstructs the residual signal by performing an inverse transform on the inverse quantization transform coefficients. In this case, when the block to be inversely transformed is predicted based on extrapolated intra-frame prediction, the video decoding device can determine the inverse transform kernel as follows when determining the primary inverse transform kernel and / or the secondary inverse transform kernel. Here, the primary inverse transform kernel can be a separable primary inverse transform kernel or a non-separable primary inverse transform kernel. The secondary inverse transform kernel can be a non-separable secondary inverse transform kernel.

[0153] For example, the video decoding device calculates the gradient values ​​of the prediction signal generated based on extrapolation-based intra-prediction, and then derives the intra-prediction mode for the current block based on the calculated gradient values. For instance, the video decoding device generates a histogram of oriented gradients (HoG) based on the prediction signal generated based on extrapolation-based intra-prediction, and derives the angle prediction mode corresponding to the direction with the most accumulated gradients based on the generated HoG. The video decoding device sets the derived angle prediction mode as the intra-prediction mode for the current block. The video decoding device can then determine the inverse transform kernel for the current block based on the set intra-prediction mode.

[0154] As another example, the video decoding device can set the intra-prediction mode of the current block to a specific intra-prediction mode (e.g., DC, planar mode, etc.) and determine the inverse transform kernel of the current block based on the set intra-prediction mode. Here, the specific intra-prediction mode can be predetermined based on the protocol between the video encoding device and the video decoding device.

[0155] The following describes how to utilize Figure 18 and Figure 19 The diagram illustrates a method for performing extrapolation-based intra-frame prediction.

[0156] Figure 18 This is a flowchart of a method for encoding the current block by a video encoding device according to at least one embodiment of the present invention.

[0157] The video coding apparatus determines a reference sample region (S1800) for deriving filter coefficients when performing extrapolation-based intra-frame prediction.

[0158] Here, the reference sample region includes the reference samples for the reconstruction of the current block. For example... Figures 11a to 11c As shown, the reference sample region is located to the left and top of the current block, or to the left or top of the current block. Figures 11a to 11c As shown, in addition to the width and height of the current block, the size of the reference sample region is defined based on preset parameter values ​​such as A, B, C, and D.

[0159] The video coding device can determine whether to perform extrapolation-based intra prediction based on the aspect ratio of the current block, the size of the current block, the intra prediction mode of the block containing the reconstructed reference sample, or whether the reference sample region is used.

[0160] The video coding apparatus determines the filter shape for extrapolation-based intra-frame prediction (S1802). Here, the filter shape is rectangular and includes a prediction pixel region and an input pixel region. The prediction pixel region is a rectangular region with width and height, and the input pixel region is located to the left, above, or both to the left and above the prediction pixel region.

[0161] As an example, the video encoding device determines an index that indicates a predefined filter shape. The video encoding device encodes the determined index.

[0162] As another example, a video encoding device can determine the filter shape based on the aspect ratio of the current block.

[0163] The video coding device derives the filter coefficients based on the reference sample region and the filter shape (S1804).

[0164] The video coding device applies the derived filter coefficients to the reference and prediction samples for the reconstruction of the current block to generate the first prediction block of the current block (S1806).

[0165] The video coding device determines the intra-prediction mode for the current block (S1808). Here, the intra-prediction mode refers to a prediction mode other than the intra-prediction mode based on extrapolation.

[0166] The video encoding device generates a second prediction block for the current block based on the intra-frame prediction mode (S1810).

[0167] The video coding device determines a flag indicating whether to perform extrapolation-based intra-frame prediction based on the first prediction block and the second prediction block (S1812).

[0168] From a rate-distortion optimization perspective, the video coding device can determine flags. For example, the flag can be set to true when the first predicted block is optimal, and vice versa when the second predicted block is optimal.

[0169] Video encoding device encoding mark (S1814).

[0170] The video encoding device uses the value of a flag to subtract either the first or second prediction block from the current block, thereby generating the residual signal for the current block. The video encoding device transforms / quantizes the residual signal to generate quantized transform coefficients and encodes the quantized transform coefficients.

[0171] On the other hand, in order to transform the residual signal derived from the first prediction block, the video coding apparatus can determine the primary transform kernel and / or the secondary transform kernel as follows. Here, the primary transform kernel can be a separable primary transform kernel or a non-separable primary transform kernel. The secondary transform kernel can be a non-separable secondary transform kernel. The video coding apparatus calculates the gradient values ​​of the predicted signal within the first prediction block and derives the intra-prediction mode of the current block based on the calculated gradient values. For example, the video coding apparatus generates a HoG based on the first prediction signal generated according to the extrapolated intra-prediction, and derives the angle prediction mode corresponding to the direction with the most accumulation based on the generated HoG. The video coding apparatus sets the derived angle prediction mode as the intra-prediction mode of the current block. The video coding apparatus can determine the transform kernel of the current block based on the set intra-prediction mode.

[0172] Figure 19 This is a flowchart of a method for reconstructing the current block by a video decoding device according to at least one embodiment of the present invention.

[0173] The video decoding device decodes a flag indicating whether to perform extrapolation-based intra-frame prediction from the bitstream (S1900).

[0174] The video decoding device decodes the aforementioned flags based on the aspect ratio of the current block, the size of the current block, the intra-frame prediction mode of the block or part of the block containing the reconstructed reference sample, or whether to use the reference sample region.

[0175] Video decoding device inspection mark (S1902).

[0176] If the flag is true (yes in S1902), the video decoding device performs the following steps.

[0177] The video decoding device determines a reference sample region for deriving the filter coefficients (S1904).

[0178] Here, the reference sample region includes the reference samples for the reconstruction of the current block. For example... Figures 11a to 11c As shown, the reference sample region is located above and to the left of the current block, or to the left of the current block, or above the current block. Figures 11a to 11c As shown, in addition to the width and height of the current block, the size of the reference sample region is defined based on preset parameter values ​​such as A, B, C, and D.

[0179] The video decoding apparatus determines the filter shape for extrapolation-based intra-frame prediction (S1906). Here, the filter shape is rectangular and includes a prediction pixel region and an input pixel region. The prediction pixel region is a rectangular region with width and height, and the input pixel region is located to the left, above, or both to the left and above the prediction pixel region.

[0180] As an example, a video decoding device decodes an index from the bitstream that indicates a predefined filter shape. The video decoding device determines the filter shape based on the decoded index.

[0181] As another example, the video decoding device can determine the filter shape based on the aspect ratio of the current block. For instance, when the aspect ratio of the current block is greater than 1, the video decoding device determines the filter shape using the aspect ratio of the predicted pixel region greater than 1. If the aspect ratio of the current block is less than 1, the video decoding device determines the filter shape using the aspect ratio of the predicted pixel region less than 1. When the aspect ratio of the current block is 1, the video decoding device determines the filter shape based on a protocol between the video encoding device and the video decoding device, thereby selecting a filter shape with an aspect ratio of the predicted pixel region greater than or less than 1. Alternatively, when the aspect ratio of the current block is 1, the video decoding device determines the filter shape using the aspect ratio of the predicted pixel region that is 1.

[0182] The video decoding device derives filter coefficients based on the reference sample region and the filter shape (S1908). The filter coefficients are derived to minimize the cost between the reconstructed reference sample and the value generated by multiplying the filter coefficients by the reconstructed reference sample defined according to the filter shape in the reference sample region.

[0183] For example, a video decoding device derives a filter matrix that depends on the size of the input pixel region and the size of the predicted pixel region.

[0184] In another example, the video decoding device subtracts an offset from the input and output reference samples defined by the filter shape in the reference sample region. The subtracted input and output reference samples can then be used to derive the filter coefficients. Here, the offset can be the average of the reconstructed reference samples, the median of the reconstructed reference samples, or a median determined based on the bit depth.

[0185] The video decoding device applies the derived filter coefficients to the reference and prediction samples for the reconstruction of the current block to generate the prediction block for the current block (S1910).

[0186] The video decoding device applies a filter matrix to the input pixel region, thereby generating prediction samples in parallel within the prediction pixel region based on dependencies.

[0187] On the other hand, if the flag is false (S1902 is no), the video decoding device performs the following steps.

[0188] The video decoding device obtains the intra-prediction mode of the current block (S1920). Here, the intra-prediction mode indicates a prediction mode other than intra-prediction based on extrapolation.

[0189] The video decoding device generates a prediction block for the current block based on the intra-frame prediction mode (S1922).

[0190] The video decoding device decodes the quantized transform coefficients from the bitstream and performs inverse quantization / inverse transform on the quantized transform coefficients to reconstruct the residual signal. The video decoding device sums the predicted block and the residual signal to generate the reconstructed signal of the current block.

[0191] On the other hand, for the inverse transform, the video decoding device can determine the primary inverse transform kernel and / or the secondary inverse transform kernel as follows. Here, the primary inverse transform kernel can be a separable primary inverse transform kernel or a non-separable primary inverse transform kernel. The secondary inverse transform kernel can be a non-separable secondary inverse transform kernel. The video decoding device calculates the gradient value of the prediction signal generated based on extrapolation-based intra-frame prediction, and derives the intra-frame prediction mode for the current block based on the calculated gradient value. For example, the video decoding device generates a HoG based on the prediction signal generated based on extrapolation-based intra-frame prediction, and derives the angle prediction mode corresponding to the direction with the most accumulation based on the generated HoG. The video decoding device sets the derived angle prediction mode as the intra-frame prediction mode for the current block. The video decoding device can determine the inverse transform kernel for the current block based on the set intra-frame prediction mode.

[0192] Although the steps in the various flowcharts are described in sequence, these steps merely exemplify the technical ideas of some embodiments of the invention. Therefore, those skilled in the art to which this invention pertains can perform the steps by changing the order described in the various figures or by performing two or more steps in parallel. Thus, the steps in the various flowcharts are not limited to the chronological order shown.

[0193] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functionality described in some embodiments can be implemented by hardware, software, firmware, and / or combinations thereof. It should also be understood that the functional components described in this invention are designated as "...units" to emphasize their potential for independent implementation.

[0194] On the other hand, the various methods or functions described in some embodiments can be implemented as instructions stored in a non-volatile recording medium, which can be read and executed by one or more processors. The non-volatile recording medium can include various types of recording devices, such as those storing data in a computer system-readable form. For example, the non-volatile recording medium can include storage media such as erasable programmable read-only memory (EPROM), flash memory drives, optical disk drives, magnetic hard disk drives, and solid-state drives (SSDs), etc.

[0195] Although exemplary embodiments of the invention have been described for illustrative purposes, those skilled in the art will understand that various modifications, additions, and substitutions can be made without departing from the spirit and scope of the invention. Therefore, embodiments of the invention have been described for the sake of brevity and clarity. The scope of the technical concept of the embodiments of the invention is not limited to the examples. Accordingly, those skilled in the art will understand that the scope of the invention should not be limited by the embodiments explicitly described above, but rather by the claims and their equivalents.

[0196] (Explanation of reference numerals in the attached image) 122: Intra-frame predictor 542: Intra-frame predictor 902: Prediction Unit Determiner 904: Predictive Technique Determiner 906: Predictive Mode Determiner 908: Predictive Actuator 1002: Reference Sample Region Determiner 1004: Filter Shape Determiner 1006: Filter Coefficient Determiner 1008: Extrapolation actuator.

[0197] Cross-reference to related applications This application claims priority and benefit to Korean Patent Application No. 10-2023-0063334, filed on May 16, 2023, and Korean Patent Application No. 10-2024-0051001, filed on April 16, 2024, the entire contents of each of which are incorporated herein by reference.

Claims

1. A method for reconstructing a current block using a video decoding device, the method comprising: Determine a reference sample region for deriving filter coefficients, the reference sample region including the reconstructed reference samples of the current block; Determine the filter shape for extrapolation-based intra-frame prediction. The filter shape is rectangular and includes a prediction pixel region and an input pixel region. The filter coefficients are derived based on the reference sample region and the filter shape. as well as The predicted block for the current block is generated by applying the derived filter coefficients to the reference and predicted samples for the reconstruction of the current block.

2. The method according to claim 1, further comprising: Based on the aspect ratio of the current block, the size of the current block, the intra-prediction mode of the block or part of the block containing the reconstructed reference sample, or whether to use the reference sample region, a flag indicating whether to perform extrapolation-based intra-prediction is generated from the bitstream decoding. Check the aforementioned mark. When the flag is true, the method continues to decode the reference sample region.

3. The method according to claim 1, wherein, The reference sample area is located above the current block, or to the left of the current block, or to both the left and top of the current block, and the reference sample area has a size defined based on the width of the current block, the height of the current block, and preset parameter values.

4. The method according to claim 1, wherein, The predicted pixel region is a rectangular region with width and height, and the input pixel region is located to the left, above, or both to the left and above the predicted pixel region.

5. The method according to claim 1, wherein, Determining the filter shape includes: The filter shape is then determined based on the index of a predefined filter shape obtained from the bitstream decoding.

6. The method according to claim 1, wherein, Determining the filter shape includes: Derive the filter coefficients from the predefined reference sample regions used for each filter shape in the predefined filter shapes; Generate a prediction signal for the current block, the prediction signal corresponding to each filter shape and filter coefficient in a predefined reference sample region; and Select a filter shape that minimizes the cost between the reconstructed sample corresponding to the prediction signal of the current block and the prediction signal of the current block.

7. The method according to claim 1, wherein, Determining the filter shape includes: When the current block has an aspect ratio greater than 1, the filter shape is determined to be the shape of the predicted pixel region with an aspect ratio greater than 1; and When the current block has an aspect ratio of less than 1, the filter shape is determined to be the shape of the predicted pixel region with an aspect ratio of less than 1.

8. The method according to claim 1, wherein, Determining the filter shape includes: When the current block has an aspect ratio of 1, the filter shape is determined to be the shape of the predicted pixel region with an aspect ratio of 1.

9. The method according to claim 1, wherein, The derivation of the filter coefficients includes: The filter coefficients are derived to minimize the cost between the reconstructed reference sample defined by the filter shape in the reference sample region and the value generated by multiplying the filter coefficients by the reconstructed reference sample.

10. The method according to claim 1, wherein, The derivation of the filter coefficients includes: The derivation depends on the size of the input pixel region and the size of the predicted pixel region.

11. The method according to claim 10, wherein, The predicted blocks that generated the current block include: By applying a filter matrix to the input pixel region, predicted samples in the predicted pixel region are generated in parallel based on dependencies.

12. The method according to claim 1, wherein, The derivation of the filter coefficients includes: The input and output reference samples, defined based on the filter shape in the reference sample region, are subtracted by an offset. Then, the filter coefficients are derived using the input and output reference samples after the offset is subtracted. The offset is the mean of the reconstructed reference samples, the median of the reconstructed reference samples, or the median determined based on the bit depth.

13. The method of claim 1, further comprising: A histogram of orientation gradients (HoG) is generated based on the prediction signal within the prediction block, and an angle prediction pattern corresponding to the direction with the most accumulation is derived based on the histogram of orientation gradients. Set the derived angle prediction mode to the intra-frame prediction mode of the current block; as well as The inverse transform kernel for the current block is determined based on the set intra-frame prediction mode.

14. A method for encoding a current block using a video encoding device, the method comprising: Determine a reference sample region for deriving filter coefficients, the reference sample region including the reconstructed reference samples of the current block; Determine the filter shape for extrapolation-based intra-frame prediction. The filter shape is rectangular and includes a prediction pixel region and an input pixel region. The filter coefficients are derived based on the reference sample region and the filter shape. as well as The first prediction block of the current block is generated by applying the derived filter coefficients to the reference and prediction samples used for the reconstruction of the current block.

15. The method according to claim 14, wherein, Based on the aspect ratio of the current block, the size of the current block, the intra-prediction mode of the block or part of the block containing the reconstructed reference sample, or whether to use the reference sample region to determine the reference sample region.

16. The method of claim 14, further comprising: Determine the intra-prediction mode of the current block, where the intra-prediction mode represents a prediction mode other than the extrapolation-based intra-prediction mode. as well as Generate a second prediction block for the current block based on the intra-frame prediction mode.

17. The method of claim 16, further comprising: Based on the first and second prediction blocks, a flag indicating whether to perform extrapolation-based intra-frame prediction is determined; The flag is encoded.

18. The method of claim 14, further comprising: A histogram of orientation gradients (HoG) is generated based on the prediction signal within the prediction block, and an angle prediction pattern corresponding to the direction with the most accumulation is derived based on the histogram of orientation gradients. Set the derived angle prediction mode to the intra-frame prediction mode of the current block; as well as The transform kernel for the current block is determined based on the set intra-frame prediction mode.

19. A method for providing video data to a video decoding device, the method comprising: Encode video data into a bitstream; as well as Send the bitstream to the video decoding device. The encoded video data includes: Determine a reference sample region for deriving filter coefficients, the reference sample region including the reconstructed reference samples of the current block; Determine the filter shape for extrapolation-based intra-frame prediction. The filter shape is rectangular and includes a prediction pixel region and an input pixel region. Deriving filter coefficients based on the reference sample region and filter shape; and The predicted block for the current block is generated by applying the derived filter coefficients to the reference and predicted samples for the reconstruction of the current block.

20. The method according to claim 19, wherein, The encoded video data further includes: A histogram of orientation gradients (HoG) is generated based on the prediction signal within the prediction block, and an angle prediction pattern corresponding to the direction with the most accumulation is derived based on the histogram of orientation gradients. Set the derived angle prediction mode to the intra-frame prediction mode of the current block; and The transform kernel for the current block is determined based on the set intra-frame prediction mode.

Citation Information

Patent Citations

  • Limiting allocation of ways in a cache based on cache maximum associativity value

    KR1020230063334A

  • Display device

    KR1020240051001A