Method and apparatus for video coding and decoding using remote reference pixels

By acquiring and utilizing reference samples position and type information of the reconstruction area far away from the current block, the replicated reference samples of the video block are generated, and the problem of low efficiency of computer-generated video encoding and decoding in the prior art is solved, and more efficient and better video quality is achieved.

CN120513622APending Publication Date: 2025-08-19HYUNDAI MOTOR CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380091255.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2023-12-28
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology is inefficient in processing computer-generated videos, making it difficult to effectively use reference pixels far from the current block for intra prediction, resulting in insufficient video quality and efficiency.

Method used

By obtaining adjacent reference sample location and type information of reconstruction areas similar to and away from the current block, the reference sample of the reconstruction area is copied to the periphery of the current block, a copy reference sample of the current block is generated, and a prediction block is generated according to the intra prediction mode.

Benefits of technology

Improves video encoding and decoding efficiency and enhances video quality, especially when processing computer-generated videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120513622A_ABST
    Figure CN120513622A_ABST
Patent Text Reader

Abstract

Methods and apparatus for video coding and decoding using remote reference pixels in accordance with the present embodiments are disclosed. In this embodiment, an image decoding device acquires an intra prediction mode of a current block. An image decoding apparatus acquires position information about surrounding reference samples of a reconstruction region. Here, the reconstruction region is similar to and located away from the current block. An image decoding apparatus acquires type information on surrounding reference samples. The image decoding device generates copied reference samples of the current block by copying the surrounding reference samples of the reconstruction region to the periphery of the current block using the position information and the type information on the surrounding reference samples. The image decoding device generates a prediction block of the current block by using the copied reference sample according to the intra prediction mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video coding and decoding method and device using non-adjacent or distant reference pixels. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and may not constitute prior art.

[0003] Since video data has a large amount of data compared to audio data or still image data, video data requires a large amount of hardware resources (including memory) to store or transmit the uncompressed video data.

[0004] Accordingly, an encoder is typically used to compress and store or transmit video data. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression technologies include H.264 / Advanced Video Codec (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Codec (VVC), which improves codec efficiency by approximately 30% or more compared to HEVC.

[0005] However, as image size, resolution, and frame rate continue to increase, the amount of data to be encoded will also increase. Accordingly, new compression technologies are needed that provide higher encoding and decoding efficiency and improved image enhancement effects than existing compression technologies.

[0006] Intra-frame prediction uses information about pixels within a common picture to predict the pixel values of the current block to be encoded. Intra-frame prediction can select the best one of multiple intra-frame prediction modes to suit the characteristics of the picture and use the selected mode to predict the current block. The encoder selects one of the multiple intra-frame prediction modes and encodes the current block using the selected mode. The encoder can then communicate information about the mode to the decoder.

[0007] HEVC technology uses a total of 35 intra-frame prediction modes for intra-frame prediction, including 33 angular modes with directionality and 2 non-angular modes without directionality. However, as the spatial resolution of the video increases from 720×480 to 2048×1024 or 8192×4096, the size of the prediction block unit also increases, which further requires the addition of more intra-frame prediction modes. Figure 3a As shown, the VVC technology utilizes 65 prediction modes that are further subdivided for intra prediction, thereby allowing more diverse prediction directions than previous technologies.

[0008] On the other hand, when performing intra prediction, the prediction block is generated by utilizing pixels around the current block, so the performance of intra prediction depends on the selection of appropriate reference pixels. As a method of selecting reference pixels, a method of obtaining reference pixels from a more accurate direction by ensuring the diversity of prediction modes or a method of increasing the number of available reference pixel candidates can be used. The existing technology corresponding to the latter is called multiple reference lines (MRL) or multiple reference line prediction (MRLP). For example, when MRL is applied to intra prediction of the current block, in addition to using a reference line at a pixel interval immediately adjacent to the current block for prediction, pixels further away from the reference line can also be used as reference pixels.

[0009] Compared to images typically captured by using an image sensor (i.e., natural video), computer-generated video (i.e., screen content) does not contain repetitive patterns, strong edges, noise, etc., and therefore exhibits characteristics different from natural video. When screen content is encoded by a traditional encoder (HEVC, VVC, etc.) that focuses on encoding natural video, the encoding and decoding efficiency is reduced. In order to combat the reduction in encoding and decoding efficiency, some technologies specifically encode screen content. The intra block copy (IBC) technology uses pixels from an area (i.e., a reference block) that is most similar to the encoding or decoding target area (i.e., the current block in the current picture). Here, the position of the reference block can be indicated by a block vector.

[0010] As mentioned above, existing MRL techniques utilize reference pixels adjacent to the current block. Furthermore, existing IBC techniques utilize the reconstructed region as a reference block, but do not utilize reference pixels surrounding the reconstructed region. Therefore, a method utilizing reference pixels distant from the current block is needed to improve video encoding and decoding efficiency and enhance video quality. Summary of the Invention

[0011] Technical issues

[0012] The present invention is directed to providing a video coding method and apparatus, which are used to generate a prediction block of the current block by using reconstructed samples far away from the current block as reference samples when performing intra-frame prediction of the current block.

[0013] Technical Solution

[0014] At least one aspect of the present invention provides a method for reconstructing a current block by a video decoding device. The method includes obtaining an intra-frame prediction mode for the current block. The method also includes obtaining position information of neighboring reference samples of a reconstructed region that is similar to and distant from the current block. The method also includes obtaining type information of the neighboring reference samples. The method also includes copying the neighboring reference samples of the reconstructed region to a periphery of the current block using the position information and the type information of the neighboring reference samples, thereby generating copied reference samples of the current block. The method also includes generating a prediction block for the current block using the copied reference samples according to the intra-frame prediction mode.

[0015] Another aspect of the present invention provides a method for encoding a current block using a video encoding device. The method includes obtaining position information of neighboring reference samples of a reconstructed region that is similar to the current block and distant from the current block. The method also includes obtaining type information of the neighboring reference samples. The method also includes obtaining an intra-frame prediction mode for the current block. The method also includes copying the neighboring reference samples of the reconstructed region to a periphery of the current block using the position information and type information of the neighboring reference samples, thereby generating copied reference samples for the current block. The method also includes generating a first prediction block for the current block using the copied reference samples according to the intra-frame prediction mode.

[0016] Another aspect of the present invention provides a computer-readable recording medium storing a bitstream generated by a video encoding method. The video encoding method includes obtaining position information of adjacent reference samples of a reconstructed area that is similar to a current block and distant from the current block. The video encoding method also includes obtaining type information of the adjacent reference samples. The video encoding method also includes obtaining an intra-frame prediction mode for the current block. The video encoding method also includes generating replicated reference samples of the current block by copying the adjacent reference samples of the reconstructed area to the periphery of the current block using the position information of the adjacent reference samples and the type information of the adjacent reference samples. The video encoding method also includes generating a prediction block of the current block according to the intra-frame prediction mode by using the copied reference samples.

[0017] Beneficial effects

[0018] As described above, the present invention provides a video coding method and apparatus for generating a prediction block for a current block by utilizing reconstructed samples located away from the current block as reference samples when performing intra-frame prediction of the current block. Thus, the video coding method and apparatus improve video coding efficiency and enhance video quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a block diagram of a video encoding device that can implement the technology of the present invention.

[0020] Figure 2A method of partitioning a block using a quadtree plus binary tree ternary tree (QTBTTT) structure is shown.

[0021] Figure 3a and Figure 3b A plurality of intra prediction modes including a wide-angle intra prediction mode are shown.

[0022] Figure 4 Shows the neighboring blocks of the current block.

[0023] Figure 5 is a block diagram of a video decoding device that can implement the technology of the present invention.

[0024] Figure 6 is a schematic diagram illustrating reference lines utilized by multiple reference lines (MRL).

[0025] Figure 7 is another schematic diagram illustrating reference lines utilized by MRL.

[0026] Figure 8 is a schematic diagram illustrating an intra block copy (IBC) technique.

[0027] Figure 9 Flowchart of a block vector transmission method classified by IBC transmission scheme.

[0028] Figure 10 is a schematic diagram showing reconstructed samples according to some prediction methods.

[0029] Figure 11 is a schematic diagram illustrating the locations of distant reference pixels according to at least one embodiment of the present invention.

[0030] Figure 12 is a schematic diagram illustrating the shape of a template according to at least one embodiment of the present invention.

[0031] Figure 13 is a schematic diagram showing the shape of a template according to another embodiment of the present invention.

[0032] Figure 14 is a schematic diagram illustrating types of reference samples according to at least one embodiment of the present invention.

[0033] Figure 15 is a diagram illustrating unidirectional intra prediction according to at least one embodiment of the present invention.

[0034] Figure 16 is a diagram illustrating symmetric bidirectional intra prediction according to at least one embodiment of the present invention.

[0035] Figure 17is a schematic diagram illustrating symmetric bidirectional intra prediction according to another embodiment of the present invention.

[0036] Figure 18 is a diagram illustrating asymmetric bidirectional intra prediction according to at least one embodiment of the present invention.

[0037] Figure 19 is a schematic diagram illustrating asymmetric bidirectional intra prediction according to another embodiment of the present invention.

[0038] Figure 20 is a flowchart of a method of encoding a current block by a video encoding apparatus according to at least one embodiment of the present invention.

[0039] Figure 21 is a flowchart of a method for reconstructing a current block by a video decoding device according to at least one embodiment of the present invention. DETAILED DESCRIPTION

[0040] Hereinafter, some embodiments of the present invention will be described in detail with reference to the accompanying illustrative drawings. In the following description, the same reference numerals represent the same elements, even though the elements are shown in different drawings. In addition, in the following description of some embodiments, when it is believed that the detailed description of related known components and functions obscures the subject matter of the present invention, the detailed description of the related known components and functions may be omitted for the sake of clarity and brevity.

[0041] Figure 1 FIG. 1 is a block diagram of a video encoding device that can implement the technology of the present invention. Figure 1 , a video encoding device and components of the device are described.

[0042] The encoding device may include: an image splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filtering unit 180 and a memory 190.

[0043] Each component of the encoding device can be implemented as hardware or software, or as a combination of hardware and software. In addition, the function of each component can be implemented as software, and the microprocessor can also be implemented to execute the function of the software corresponding to each component.

[0044] A video consists of one or more sequences including multiple images. Each image is divided into multiple regions, and encoding is performed on each region. For example, an image is divided into one or more tiles or / and slices. Here, one or more tiles can be defined as a tile group. Each tile or / and slice is divided into one or more coding tree units (CTUs). In addition, each CTU is divided into one or more coding units (CUs) through a tree structure. The information applied to each coding unit (CU) is encoded as the syntax of the CU, and the information commonly applied to the CUs included in a CTU is encoded as the syntax of the CTU. In addition, the information commonly applied to all blocks in a slice is encoded as the syntax of the slice header, and the information applied to all blocks constituting one or more images is encoded as a picture parameter set (PPS) or a picture header. In addition, information commonly referenced by multiple images is encoded as a sequence parameter set (SPS). In addition, information commonly referenced by one or more SPSs is encoded as a video parameter set (VPS). In addition, information commonly applied to one tile or tile group may also be encoded as syntax of the tile or tile group header.The syntax included in the SPS, PPS, slice header, tile or tile group header may be referred to as a high-level syntax.

[0045] The image splitter 110 determines the size of a codec tree unit (CTU). Information on the size of the CTU (CTU size) is encoded as a syntax of an SPS or PPS and transmitted to a video decoding apparatus.

[0046] The image splitter 110 splits each image constituting a video into a plurality of codec tree units (CTUs) of a predetermined size, and then recursively splits the CTUs using a tree structure. Leaf nodes in the tree structure become codec units (CUs), which are basic units of coding.

[0047] The tree structure can be a quadtree (QT), in which a higher node (or parent node) is split into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), in which a higher node is split into two lower nodes. The tree structure can also be a ternary tree (TT), in which a higher node is split into three lower nodes at a ratio of 1:2:1. The tree structure can also be a structure in which two or more structures of the QT structure, the BT structure, and the TT structure are mixed. For example, a quadtree plus binary tree (QTBT) structure can be used, or a quadtree plus binary tree ternary tree (QTBTTT) structure can be used. Here, a binary tree ternary tree (BTTT) is added to the tree structure to be called a multiple-type tree (MTT).

[0048] Figure 2 It is a schematic diagram for describing a method of dividing a block by using a QTBTTT structure.

[0049] like Figure 2 As shown, the CTU can first be split into a QT structure. The quadtree splitting can be recursive until the size of the split block reaches the minimum block size (MinQTSize) of the leaf node allowed in QT. The first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoder 155 and notified to the video decoding device with a signal. When the leaf node of QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in BT, the leaf node can be further split into at least one of a BT structure or a TT structure. There can be multiple splitting directions in the BT structure and / or the TT structure. For example, there can be two directions, namely, the direction of splitting the blocks of the corresponding node horizontally and the direction of splitting the blocks of the corresponding node vertically. As shown in FIG. Figure 2 As shown, when MTT splitting starts, the entropy encoder 155 encodes a second flag (mtt_split_flag) indicating whether the node is split, and a flag indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (binary or trifurcated) when the node is split, and notifies the video decoding device of the same with a signal.

[0050] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four nodes in the lower layer, a CU split flag (split_cu_flag) indicating whether the node is split may also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the partition tree structure and becomes a CU, which is the basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding device first starts encoding the first flag in the above-mentioned scheme.

[0051] When QTBT is used as another example of a tree structure, there may be two types, namely, a type in which the block of the corresponding node is horizontally split into two blocks of the same size (i.e., symmetrical horizontal splitting) and a type in which the block of the corresponding node is vertically split into two blocks of the same size (i.e., symmetrical vertical splitting). The entropy encoder 155 encodes a split flag (split_flag) indicating whether each node of the BT structure is split into blocks of the lower layer and split type information indicating the split type, and transmits them to the video decoding device. On the other hand, there may also be a type in which the block of the corresponding node is split into two blocks that are asymmetric to each other. The asymmetric form may include a form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or may also include a form in which the block of the corresponding node is split in a diagonal direction.

[0052] The CU can have various sizes depending on the QTBT or QTBTTT partitioned from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block." When QTBTTT partitioning is adopted, the shape of the current block can also be a rectangular shape in addition to a square shape.

[0053] The predictor 120 predicts the current block to generate a predicted block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.

[0054] Typically, each current block in an image can be predictively encoded. Typically, prediction of the current block can be performed using intra-frame prediction techniques (which utilize data from the image that includes the current block) or inter-frame prediction techniques (which utilize data from an image that was encoded before the image that includes the current block). Inter-frame prediction includes both unidirectional prediction and bidirectional prediction.

[0055] The intra-frame predictor 122 predicts pixels in the current block by using pixels (reference pixels) located adjacent to the current block in the current image including the current block. There are multiple intra-frame prediction modes according to the prediction direction. For example, Figure 3aAs shown, the plurality of intra prediction modes may include two non-directional modes including a planar mode and a DC mode, and may include 65 directional modes. Neighboring pixels to be used and algorithm equations are defined differently according to each prediction mode.

[0056] In order to perform efficient directional prediction for a current block with a rectangular shape, we can additionally use Figure 3b The directional modes are shown in the figure with dotted arrows (intra-frame prediction modes #67 to #80, #-1 to #-14). The directional modes can be called "wide angle intra-prediction modes". Figure 3b , the arrows indicate the corresponding reference samples used for prediction, rather than indicating the prediction direction. The prediction direction is opposite to the direction indicated by the arrow. When the current block has a rectangular shape, the wide-angle intra prediction mode is a mode in which prediction is performed in a direction opposite to the specific direction mode without additional bit transmission. In this case, in the wide-angle intra prediction mode, some wide-angle intra prediction modes available for the current block can be determined by the ratio of the width to the height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height less than the width, a wide-angle intra prediction mode (intra prediction modes #67 to #80) with an angle less than 45 degrees is available. When the current block has a rectangular shape with a width greater than the height, a wide-angle intra prediction mode with an angle greater than -135 degrees is available.

[0057] The intra-frame predictor 122 can determine the intra-frame prediction to be used for encoding the current block. In some examples, the intra-frame predictor 122 can encode the current block by utilizing multiple intra-frame prediction modes, and can also select an appropriate intra-frame prediction mode to use from a test mode. For example, the intra-frame predictor 122 can calculate the rate-distortion value by utilizing a rate-distortion analysis of multiple test intra-frame prediction modes, and can also select the intra-frame prediction mode with the best rate-distortion characteristics from the test mode.

[0058] The intra-frame predictor 122 selects one intra-frame prediction mode from a plurality of intra-frame prediction modes and predicts the current block by using adjacent pixels (reference pixels) and an algorithm equation determined according to the selected intra-frame prediction mode. The entropy encoder 155 encodes information about the selected intra-frame prediction mode and transmits it to the video decoding device.

[0059] The inter-frame predictor 124 generates a prediction block for the current block by utilizing motion compensation processing. The inter-frame predictor 124 searches for a block that is most similar to the current block in a reference image that has been encoded and decoded earlier than the current image, and generates a prediction block for the current block by utilizing the searched block. In addition, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current image and the prediction block in the reference image. Typically, motion estimation is performed on the luminance (luma) component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chrominance component. The motion information including the information of the reference image and the information about the motion vector used to predict the current block is encoded by the entropy encoder 155 and transmitted to the video decoding device.

[0060] The inter-frame predictor 124 may also interpolate a reference image or reference block to increase prediction accuracy. In other words, subsamples are interpolated between two consecutive integer samples by applying filter coefficients to a plurality of consecutive integer samples. When searching for a block most similar to the current block in the interpolated reference image, fractional precision rather than integer sample precision may be used for the motion vector. The precision or resolution of the motion vector may be set differently for each target region to be encoded, such as a unit such as a slice, tile, CTU, or CU. When such adaptive motion vector resolution (AMVR) is applied, information regarding the motion vector resolution to be applied to each target region should be signaled for each target region. For example, when the target region is a CU, information regarding the motion vector resolution to be applied to each CU is signaled. The information regarding the motion vector resolution may be information indicating the precision of the motion vector difference, as described below.

[0061] On the other hand, the inter-frame predictor 124 can perform inter-frame prediction using bidirectional prediction. Bidirectional prediction uses two reference images and two motion vectors representing the block positions most similar to the current block in each reference image. The inter-frame predictor 124 selects a first reference image and a second reference image from reference image list 0 (RefPicList0) and reference image list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for blocks most similar to the current block in the respective reference images to generate first and second reference blocks. Furthermore, a prediction block for the current block is generated by averaging or weighted averaging the first and second reference blocks. Motion information, including information about the two reference images used to predict the current block and information about the two motion vectors, is transmitted to the entropy encoder 155. Reference image list 0 may consist of images preceding the current image in display order among the pre-reconstructed images, and reference image list 1 may consist of images following the current image in display order among the pre-reconstructed images. However, while not particularly limited to this, pre-reconstructed images following the current image in display order may also be included in reference image list 0. Conversely, a pre-reconstructed image preceding the current image may also be additionally included in the reference image list 1 .

[0062] To minimize the amount of bits consumed for encoding motion information, various methods may be used.

[0063] For example, when the reference image and motion vector of the current block are the same as those of a neighboring block, information identifying the neighboring block is encoded to transmit the motion information of the current block to the video decoding device. This method is called merge mode.

[0064] In the merge mode, the inter predictor 124 selects a predetermined number of merge candidate blocks (hereinafter, referred to as “merge candidates”) from neighboring blocks of the current block.

[0065] As the adjacent blocks for deriving the merge candidate, all or some of the left block A0, the lower left block A1, the upper block B0, the upper right block B1 and the upper left block B2 adjacent to the current block in the current image may be used, such as Figure 4 As shown. In addition, in addition to the current image where the current block is located, blocks located in the reference image (which may be the same as or different from the reference image used to predict the current block) can also be used as merge candidates. For example, the co-located block of the current block in the reference image or a block adjacent to the co-located block can be additionally used as a merge candidate. If the number of merge candidates selected by the above method is less than the preset number, a zero vector is added to the merge candidates.

[0066] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates using neighboring blocks. From the merge candidates included in the merge list, a merge candidate to be used as motion information for the current block is selected, and merge index information is generated to identify the selected candidate. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding device.

[0067] Merge skip mode is a special case of merge mode. After quantization, when all transform coefficients used for entropy coding are close to zero, only the neighboring block selection information is transmitted, without the residual signal. By utilizing merge skip mode, relatively high coding efficiency can be achieved for images with minimal motion, still images, and images with on-screen content.

[0068] Hereinafter, merge mode and merge skip mode are collectively referred to as merge / skip mode.

[0069] Another method for encoding motion information is the advanced motion vector prediction (AMVP) mode.

[0070] In the AMVP mode, the inter-frame predictor 124 derives a motion vector prediction candidate for the motion vector of the current block by using the neighboring blocks of the current block. As the neighboring blocks for deriving the motion vector prediction candidate, the neighboring blocks may be used. Figure 4 All or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image shown. Furthermore, in addition to the current image in which the current block is located, blocks within a reference image (which may be the same as or different from the reference image used to predict the current block) may also be used as adjacent blocks for deriving motion vector prediction candidates. For example, a co-located block of the current block within the reference image or a block adjacent to the co-located block may be used. If the number of motion vector candidates selected by the above method is less than a preset number, a zero vector is added to the motion vector candidates.

[0071] The inter-frame predictor 124 derives motion vector prediction candidates by using the motion vectors of the neighboring blocks, and determines the motion vector prediction of the motion vector of the current block by using the motion vector prediction candidates. In addition, the motion vector difference is calculated by subtracting the motion vector prediction from the motion vector of the current block.

[0072] Motion vector predictions can be obtained by applying a predefined function (e.g., median and mean calculations) to motion vector prediction candidates. In this case, the video decoding device also knows the predefined function. Furthermore, since the neighboring blocks used to derive motion vector prediction candidates are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode information for identifying motion vector prediction candidates. Accordingly, in this case, information about the motion vector difference and information about the reference image used to predict the current block are encoded.

[0073] Alternatively, the motion vector prediction may be determined by selecting any one of the motion vector prediction candidates. In this case, information identifying the selected motion vector prediction candidate is additionally encoded together with information about the motion vector difference used to predict the current block and information about the reference image.

[0074] The subtractor 130 generates a residual block by subtracting a prediction block generated by the intra predictor 122 or the inter predictor 124 from the current block.

[0075] The transformer 140 transforms the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 may transform the residual signal in the residual block by using the entire size of the residual block as a transform unit, or may divide the residual block into multiple sub-blocks and perform the transform using the sub-blocks as transform units. Alternatively, the residual block is divided into two sub-blocks, namely a transform region and a non-transform region, to transform the residual signal using only the transform region sub-block as a transform unit. Here, the transform region sub-block may be one of two rectangular blocks with a size ratio of 1:1 based on the horizontal axis (or vertical axis). In this case, the entropy encoder 155 encodes a flag (cu_sbt_flag) indicating that only the sub-block is transformed, as well as direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or position information (cu_sbt_pos_flag), and signals these to the video decoding device. In addition, the size of the transform region subblock may have a size ratio of 1:3 based on the horizontal axis (or vertical axis). In this case, the entropy encoder 155 additionally encodes a flag (cu_sbt_quad_flag) for dividing the corresponding partition and signals it to the video decoding apparatus.

[0076] On the other hand, the transformer 140 can perform transformations of the residual block separately in the horizontal direction and the vertical direction. For this transformation, various types of transformation functions or transformation matrices can be used. For example, paired transformation functions for horizontal transformation and vertical transformation can be defined as a multiple transform set (MTS). The transformer 140 can select a transformation function pair with the highest transformation efficiency in the MTS and can transform the residual block in each of the horizontal and vertical directions. The information (mts_idx) about the transformation function pair in the MTS is encoded by the entropy encoder 155 and notified to the video decoding device with a signal.

[0077] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using a quantization parameter and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also immediately quantize the relevant residual block without transforming any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) depending on the position of the transform coefficient in the transform block. The quantization matrix applied to the quantized transform coefficients arranged in two dimensions can be encoded and signaled to the video decoding device.

[0078] The rearrangement unit 150 may perform rearrangement of coefficient values on the quantized residual value.

[0079] The rearrangement unit 150 can change the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can use a zigzag scan or a diagonal scan to scan the DC coefficient to the coefficient of the high-frequency region to output a 1D coefficient sequence. Depending on the size of the transform unit and the intra-frame prediction mode, the zigzag scan can also be replaced by a vertical scan that scans the 2D coefficient array in the column direction and a horizontal scan that scans the 2D block type coefficients in the row direction. In other words, depending on the size of the transform unit and the intra-frame prediction mode, the scanning method to be used can be determined from zigzag scanning, diagonal scanning, vertical scanning, and horizontal scanning.

[0080] The entropy encoder 155 encodes the sequence of 1D quantized transform coefficients output from the rearrangement unit 150 by using various encoding schemes including Context-based Adaptive Binary Arithmetic Code (CABAC), Exponential Golomb, etc. to generate a bitstream.

[0081] In addition, the entropy encoder 155 encodes information related to block partitioning (e.g., CTU size, CTU partition flag, QT partition flag, MTT partition type, and MTT partition direction, etc.) so that the video decoding device can partition the block in the same manner as the video encoding device. In addition, the entropy encoder 155 encodes information about the prediction type indicating whether the current block is encoded by intra-frame prediction or inter-frame prediction. The entropy encoder 155 encodes intra-frame prediction information (i.e., information about the intra-frame prediction mode) or inter-frame prediction information (merge index in the case of merge mode, and information about reference image index and motion vector difference in the case of AMVP mode) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information about quantization parameters and information about quantization matrices).

[0082] The inverse quantizer 160 inversely quantizes the quantized transform coefficient output from the quantizer 145 to generate a transform coefficient. The inverse transformer 165 transforms the transform coefficient output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct a residual block.

[0083] The adder 170 reconstructs the current block by adding the reconstructed residual block and the prediction block generated by the predictor 120. When performing intra prediction on the next block, the pixels in the reconstructed current block are used as reference pixels.

[0084] The loop filtering unit 180 performs filtering on the reconstructed pixels to reduce blocking artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transformation / quantization. The loop filtering unit 180, as an in-loop filter, may include all or some of the deblocking filter 182, the sample adaptive offset (SAO) filter 184, and the adaptive loop filter (ALF) 186.

[0085] The deblocking filter 182 filters the boundaries between reconstructed blocks to remove blocking artifacts caused by block-based encoding / decoding, and the SAO filter 184 and ALF 186 perform additional filtering on the deblocked video. The SAO filter 184 and ALF 186 are filters used to compensate for the difference between reconstructed and original pixels caused by lossy coding. The SAO filter 184 applies an offset per CTU to enhance subjective image quality and coding efficiency. On the other hand, the ALF 186 performs block-based filtering and applies different filters to compensate for distortion based on the boundaries of the corresponding blocks and the degree of change. Information about the filter coefficients to be used for the ALF can be encoded and signaled to the video decoding device.

[0086] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all blocks in one image are reconstructed, the reconstructed image can be used as a reference image for inter-frame prediction of blocks within an image to be subsequently encoded.

[0087] The video encoding device may store the bit stream of the encoded video data in a non-volatile storage medium or transmit the bit stream to the video decoding device through a communication network.

[0088] Figure 5 is a functional block diagram of a video decoding device that can implement the technology of the present invention. Figure 5 , describes a video decoding device and components of the device.

[0089] The video decoding apparatus may include an entropy decoder 510 , a rearrangement unit 515 , an inverse quantizer 520 , an inverse transformer 530 , a predictor 540 , an adder 550 , a loop filtering unit 560 , and a memory 570 .

[0090] Similar to Figure 1 Each component of the video encoding device and the video decoding device can be implemented as hardware or software, or as a combination of hardware and software. In addition, the function of each component can be implemented as software, and the microprocessor can also be implemented to execute the function of the software corresponding to each component.

[0091] The entropy decoder 510 extracts information related to block partitioning by decoding a bitstream generated by a video encoding apparatus to determine a current block to be decoded, and extracts prediction information required to reconstruct the current block and information about a residual signal.

[0092] The entropy decoder 510 extracts information about the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS) to determine the size of the CTU and partitions the image into CTUs of the determined size. Furthermore, the CTU is determined as the highest level (i.e., the root node) of the tree structure, and partition information of the CTU is extracted to partition the CTU using the tree structure.

[0093] For example, when a CTU is segmented using a QTBTTT structure, a first flag (QT_split_flag) related to QT segmentation is first extracted to segment each node into four nodes in the lower layer. Furthermore, a second flag (mtt_split_flag) related to MTT segmentation, a segmentation direction (vertical / horizontal), and / or a segmentation type (binary / trifurcated) are extracted with respect to nodes corresponding to QT leaf nodes to segment the corresponding leaf nodes into an MTT structure. As a result, each node below the QT leaf node is recursively segmented into a BT or TT structure.

[0094] As another example, when a CTU is split using the QTBTTT structure, a CU split flag (split_cu_flag) indicating whether the CU is split is extracted. When the corresponding block is split, a first flag (QT_split_flag) may also be extracted. During the splitting process, for each node, zero or more recursive MTT splits may occur after zero or more recursive QT splits. For example, for a CTU, MTT splits may occur immediately, or, conversely, multiple QT splits may occur.

[0095] As another example, when a CTU is split using a QTBT structure, a first flag (QT_split_flag) related to the splitting of the QT is extracted to split each node into four nodes in the lower layer. In addition, a split flag (split_flag) indicating whether a node corresponding to a leaf node of the QT is further split into a BT and split direction information are extracted.

[0096] On the other hand, when the entropy decoder 510 determines the current block to be decoded by partitioning using a tree structure, the entropy decoder 510 extracts information about the prediction type indicating whether the current block is intra-predicted or inter-predicted. When the prediction type information indicates intra-prediction, the entropy decoder 510 extracts syntax elements of intra-prediction information (intra-prediction mode) for the current block. When the prediction type information indicates inter-prediction, the entropy decoder 510 extracts information of syntax elements representing inter-prediction information, that is, a motion vector and a reference image referenced by the motion vector.

[0097] In addition, the entropy decoder 510 extracts quantization-related information and extracts information on a quantized transform coefficient of the current block as information on a residual signal.

[0098] The rearrangement unit 515 may change the sequence of 1D quantized transform coefficients entropy-decoded by the entropy decoder 510 into a 2D coefficient array (ie, block) again in the reverse order of the coefficient scanning order performed by the video encoding device.

[0099] The inverse quantizer 520 inversely quantizes the quantized transform coefficients and inversely quantizes the quantized transform coefficients by using a quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video encoding device to the 2D array of quantized transform coefficients.

[0100] The inverse transformer 530 reconstructs a residual signal by inversely transforming the inversely quantized transform coefficient from the frequency domain to the spatial domain to generate a residual block of the current block.

[0101] In addition, when the inverse transformer 530 inverse-transforms a partial region (sub-block) of a transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transform block is transformed, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the uninverse-transformed region with a value of "0" as the residual signal to generate the final residual block of the current block.

[0102] In addition, when MTS is applied, the inverse transformer 530 determines the transformation function or transformation matrix to be applied in each of the horizontal and vertical directions by using the MTS information (mts_idx) signaled from the video encoding device. The inverse transformer 530 also performs inverse transformation on the transformation coefficients in the transformation block in the horizontal and vertical directions by using the determined transformation function.

[0103] The predictor 540 may include an intra predictor 542 and an inter predictor 544. The intra predictor 542 is activated when the prediction type of the current block is intra prediction, and the inter predictor 544 is activated when the prediction type of the current block is inter prediction.

[0104] The intra predictor 542 determines an intra prediction mode of a current block from among a plurality of intra prediction modes according to syntax elements of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block by using neighboring reference pixels of the current block according to the intra prediction mode.

[0105] The inter predictor 544 determines a motion vector of a current block and a reference image to which the motion vector refers by using a syntax element of the inter prediction mode extracted from the entropy decoder 510 , and predicts the current block by using the motion vector and the reference image.

[0106] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 to the prediction block output from the inter predictor 544 or the intra predictor 542. When intra-predicting a subsequent block to be decoded, pixels within the reconstructed current block are used as reference pixels.

[0107] The loop filter unit 560, which serves as an in-loop filter, may include a deblocking filter 562, an SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between reconstructed blocks to remove blocking artifacts caused by block-by-block decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after deblocking filtering to compensate for differences between reconstructed and original pixels caused by lossy encoding. The filter coefficients of the ALF are determined by utilizing information about the filter coefficients decoded from the bitstream.

[0108] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all blocks in one image are reconstructed, the reconstructed image can be used as a reference image for inter-frame prediction of blocks within an image to be subsequently encoded.

[0109] In some embodiments, the present invention relates to encoding and decoding video images as described above. More specifically, the present invention provides a video encoding and decoding method and apparatus for generating a prediction block of a current block by using reconstructed samples far from the current block as reference samples when performing intra-frame prediction of the current block.

[0110] The following embodiments may be performed by the intra-frame predictor 122 in the video encoding apparatus. The following embodiments may also be performed by the intra-frame predictor 542 in the video decoding apparatus.

[0111] The video encoding device may generate signaling information related to this embodiment in terms of optimizing rate-distortion when encoding the current block. The video encoding device may encode the signaling information using the entropy encoder 155 and transmit the encoded signaling information to the video decoding device. The video decoding device may decode the signaling information associated with decoding the current block from the bitstream using the entropy decoder 510.

[0112] In the following description, the term “target block” may be used interchangeably with a current block or a coding unit (CU), or may refer to some areas of a coding unit.

[0113] Furthermore, a flag value of true indicates a case where the flag is set to 1. Furthermore, a flag value of false indicates a case where the flag is set to 0.

[0114] I. Reference Pixels for Intra Prediction

[0115] This article introduces several techniques for improving codec efficiency based on intra-frame prediction. When predicting the current block using intra-frame prediction, the MRL technique uses neighboring pixels separated by one pixel from the current block, as well as pixels further away, as reference pixels for prediction. Pixels at the same distance from the current block are grouped and named reference lines. MRL uses pixels on selected reference lines to perform intra-frame prediction for the current block.

[0116] In order to indicate the reference line to be used when performing intra prediction, the video encoding device signals the reference line index intra_luma_ref_idx to the video decoding device. In the existing Versatile Video Coding (VVC), the reference line represented by each intra_luma_ref_idx is Figure 6 The existing VVC uses intra_luma_ref_idx to indicate one of the three reference lines closest to the current block. The bit allocation of the corresponding reference line index value is shown in Table 1.

[0117] [Table 1]

[0118] intra_luma_ref_idx Bit allocation 0 0 1 10 2 11

[0119] In the Enhanced Compression Model (ECM) (a technique that goes beyond VVC), the number of reference lines that can be referenced in MRL is extended to six, so that reference lines with intra_luma_ref_idx values of {0, 1, 3, 5, 7, 12} can be used. In ECM, the reference line represented by the corresponding index intra_luma_ref_idx is in Figure 7 In addition, the bit allocation of the corresponding reference line index value is shown in Table 2.

[0120] [Table 2]

[0121] intra_luma_ref_idx Bit allocation 0 0 1 10 3 110 5 1110 7 11110 12 11111

[0122] In VVC, since MRL cannot be applied to a block located at the first line in a CTU, the block at that position is always predicted by utilizing intra_luma_ref_idx0 without parsing information about the reference line. Similarly, in ECM, since MRL cannot be applied to a block located at the first line in a CTU, the block at that position is always predicted by utilizing intra_luma_ref_idx0 without parsing information about the reference line. In addition, in ECM, in order to predict a block other than the first line in the current CTU, the video encoding device does not test whether to use a reference line included in the upper CTU among the reference lines available for MRL. The video encoding device can signal to the video decoding device one of the reference lines that has been tested for use based on Table 2.

[0123] The reference line index intra_luma_ref_idx used by VVC for intra prediction and the syntax for signaling the prediction mode of the current block are shown in Table 3.

[0124] [Table 3]

[0125]

[0126] The video decoding device parses intra_luma_ref_idx to determine the reference line index used for prediction. When the reference line index is 0, the intra sub-partition (ISP) technology is applied. Therefore, if the reference line index is non-zero, the information related to the ISP is not parsed. In addition, when the prediction mode determined by the MPM is not planar mode, MRL is applied. Therefore, since a non-zero reference line index indicates the application of MRL, both intra_luma_mpm_flag and intra_luma_not_planar_flag are inferred to be 1.

[0127] In order to improve the encoding and decoding efficiency of screen content with characteristics different from natural images, there are several technologies specifically used to encode screen content. When used for encoding, the intra block copy (IBC) technology searches for a reconstructed area in the current picture for a reference block (i.e., a prediction value) that is the most similar area to the current block, such as Figure 8As shown in the figure. The position of the reference block is expressed by a block vector (BV), which is the displacement between the current block and the reference block. To improve encoding and decoding efficiency, the video encoding device does not transmit the block vector as is, but instead separates it into a block vector predictor (BVP) and a block vector difference (BVD). The video encoding device can encode the BVP and BVD and then send the BVP and BVD to the video decoding device.

[0128] On the other hand, the video encoding apparatus signals pred_mode_ibc_flag indicating whether the IBC mode is enabled or disabled to the video decoding apparatus. If pred_mode_ibc_flag is true, the current block is encoded / decoded according to the IBC mode.

[0129] In the following, the spatial resolution of the BVD and the spatial resolution of the block vector are considered to be the same. In addition, it is possible to determine that the spatial resolution values of the horizontal component and the vertical component of the block vector are the same by using a single flag.

[0130] Figure 9 Flowchart of a block vector transmission method classified by an intra block copy (IBC) transmission scheme.

[0131] like Figure 9 As shown in FIG, based on the block vector transmission method, IBC technology can be classified into IBC skip mode, IBC merge mode, and IBC AMVP mode. In IBC skip mode, the video encoding device uses the same block vector transmission method as in IBC merge mode, but does not transmit the residual block equivalent to the difference between the current block and the predicted block. Figure 9 The example shown in can be similarly applied to a video decoding apparatus. In this case, the video decoding apparatus can parse the flag required for decoding the block vector from the bitstream.

[0132] The video encoding apparatus determines whether it uses the IBC skip mode (S900). When the IBC skip mode is not used (No in S900), the video encoding apparatus checks to see whether it uses the IBC merge mode (S904).

[0133] When using IBC skip mode or IBC merge mode, the video encoding device determines a merge index merge_idx indicating one of the block vectors included in the IBC merge list (S902, S906). The video encoding device can determine the merge index in terms of rate-distortion optimization. However, the video encoding device does not calculate the block vector difference (BVD). The IBC merge list can be constructed in a manner shared by the video encoding device and the video decoding device. The video encoding device can select a block vector predictor (BVP) indicated by the merge index, and then use the BVP as a block vector. On the other hand, the video encoding device sends the merge index to the video decoding device, but does not send the BVD.

[0134] When using the IBC AMVP mode, the video encoding device sequentially determines mvp_l0_flag, BVD, and amvr_precision_idx (S908 to S912). Here, mvp_l0_flag is an index for indicating the predicted value of the motion vector, which is also used as an index for indicating the BVP of the block vector. In addition, amvr_precision_idx is an index for indicating the spatial resolution of the motion vector according to the application of adaptive motion vector resolution (AMVR), which is also used as an index for indicating the spatial resolution of the block vector. The video encoding device can select the block vector indicated by mvp_l0_flag as the BVP, and can then generate a block vector by summing the BVP with the BVD. On the other hand, the video encoding device sends mvp_l0_flag, BVD, and amvr_precision_idx to the video decoding device.

[0135] When AMVR is used in IBC AMVP mode, the video encoding device can adaptively determine the spatial resolution of the BVD in terms of rate-distortion optimization. Depending on the prediction mode of the current block, the type of spatial resolution of the motion vector that can be adaptively selected by AMVR can vary. AMVR can be used in prediction modes using vector differences (such as conventional inter-frame prediction mode, affine model-based inter-frame prediction mode, IBC mode, etc.).

[0136] Hereinafter, mvp_l0_flag is represented by a block vector prediction value index, and amvr_precision_idx is represented by a block vector spatial resolution precision index.

[0137] As described above, when using the IBC AMVP mode, the video encoding device can determine amvr_precision_idx, which indicates the spatial resolution of the BVD in terms of rate-distortion optimization. The video encoding device can signal amvr_flag and amvr_precision_idx to the video decoding device to convey the spatial resolution of the block vector. That is, the video encoding device can send amvr_flag to signal whether the AMVR technology is to be applied to the block vector. The video encoding device can also send amvr_precision_idx, which indicates one of a list of candidate spatial resolutions for the block vector, to signal the spatial resolution used for prediction. The video encoding device and the video decoding device share the same list of candidate spatial resolutions for the block vector. On the other hand, when the AMVR technology is adopted in the traditional IBC AMVP mode, amvr_flag is assumed to be 1, which can save the video encoding device from sending amvr_flag.

[0138] In the conventional IBC AMVP mode, the candidate spatial resolution list for block vectors shared between the video encoding and decoding devices is {1-pel, 4-pel}. When applying the AMVR technique based on the candidate spatial resolution list, the video encoding and decoding devices can determine the spatial resolution of the block vector, as shown in Table 4.

[0139] [Table 4]

[0140]

[0141] As described above, according to Table 4, one of 1-pel and 4-pel is selected as the spatial resolution of the block vector, and according to the selected resolution, the spatial resolution of the block vector and the spatial resolution of the block vector difference can be determined.

[0142] However, conventional intra prediction utilizes reconstructed samples located around the left and top sides of the current block as reference samples. Since encoding / decoding of the current picture is performed in z-scan order, there are no reference pixels on the right and bottom sides of the current block. Even when using MRL (which was introduced to use reference samples farther away), the convention is still to use reference samples located on the left and top sides of the current block. For example, when using the commonly used intra prediction modes of angular mode, planar mode or DC mode, reference samples located near the top and left sides of the current block are used. Since the IBC mode uses reference blocks searched from the reconstructed area, it can use reconstructed samples that are not adjacent to the current block for prediction. However, the IBC mode differs from the usual intra prediction mode in that it directly uses reconstructed samples as prediction values. There is currently no technology for performing intra prediction by utilizing reconstructed samples far away from the current block (i.e., those samples that are not adjacent to the current block) as reference samples.

[0143] Figure 10 is a schematic diagram showing reconstructed samples according to some prediction methods.

[0144] like Figure 10 As shown, the reconstructed samples used can be represented according to the prediction method. Figure 10 In the example of , reconstructed samples included in a shadow area around a reference block are not used as reference samples for the current block in the prior art prediction technique.

[0145] Hereinafter, intra prediction mode, intra mode, or prediction mode may be used interchangeably.

[0146] The following embodiments are described with respect to a video decoding apparatus, but they can also be implemented by a video encoding apparatus in the same or similar manner.

[0147] II. Embodiments according to the present invention

[0148] The above-mentioned problems of the prior art can be solved by using reconstructed samples that are not located around the current block as reference samples to perform intra-frame prediction. Therefore, according to the embodiments of the present invention, it is possible to improve encoding and decoding efficiency and / or enhance the image clarity and quality of the decoded video. For a reconstructed area similar to the current block of the original image, the adjacent samples of the relevant reconstructed area are likely to be similar to the adjacent samples of the current block. According to the present invention, various reference samples that are expected to produce prediction values similar to the current block can be used, and the performance of intra-frame prediction can be improved under the embodiment of intra-frame prediction.

[0149] In an embodiment according to the present invention, the video decoding apparatus may determine the reconstructed samples not adjacent to the current block as reference samples based on their positions and types. The following describes methods for determining the positions and types of reference samples.

[0150] The position indicates the location of the reference sample in the current picture. The video decoding device may use a reconstructed sample adjacent to a reconstructed region similar to the current block as a reference sample, or may use a reconstructed sample offset from the reconstructed region as a reference sample. The position may be determined in one of the following ways.

[0151] ① Determine the position of the reference sample by signaling its location

[0152] The following describes a case where the positions of reference samples are signaled and determined using a reconstructed area similar to the current block.

[0153] Figure 11 is a schematic diagram illustrating the location of distant reference pixels according to at least one embodiment of the present invention.

[0154] The position of the reference sample can be determined by signaling the coordinates and offset. Figure 11 In the example of , the coordinates represent the coordinates of one of pixels 1 to 4 on the upper left, upper right, lower left, and lower right sides within the reconstruction area, or the coordinates of one of pixels 5 to 8 within the reference line (along which the reference sample is positioned). The offset represents the distance between the reconstruction area and the reference line. For example, the offset between the reconstruction area and the reference line adjacent to the reconstruction area is defined as 0.

[0155] The video decoding device can determine which pixel at the above-mentioned position to use according to a predetermined method. The predetermined method may include using a pixel at a preset position, using a pixel at a position sent by a signal, etc. The coordinates of the pixel can be expressed as a vector representing the displacement between the pixel and the current block. The video decoding device can directly parse the value of the vector. The video decoding device can parse the index of one of the vectors in the list indicating the vector generated according to the merging method. Alternatively, the video decoding device can parse the index of one of the vectors in the list indicating the vector generated according to the AMVP method and the difference between the actual vector and the indicated vector. The video decoding device can directly parse the offset, or can set the offset to a preset value. Alternatively, the video decoding device can parse the index of one of the offsets in the list indicating the offset.

[0156] Hereinafter, the position information of a reference sample includes the position of a pixel, a vector indicating the pixel, and an offset.

[0157] Figure 11 The example in the illustration assumes that the coordinates of the upper left pixel of the reconstructed area are (12, 12) and the width and height of the current block are 4 and 4, respectively. When using the adjacent reference sample and determining the position based on pixel 1, the video decoding device can determine the position by using the coordinates (12, 12) of pixel 1 and offset = 0. As another example, when using the adjacent reference sample and determining the position based on pixel 5, the video decoding device can determine the position by using the coordinates (11, 11) of pixel 5 and offset = 0. As yet another example, when using a reference sample 2 pixels away from the reconstructed area and determining the position based on pixel 1, the position can be determined by using the coordinates (12, 12) of pixel 1 and offset = 2. As yet another example, when using a reference sample 2 pixels away from the reconstructed area and determining the position based on pixel 5, the video decoding device can determine the position by using the coordinates (9, 9) of pixel 5 and offset = 2.

[0158] ②Determine the location of the reference sample by inference

[0159] When determining the position by inference, the video decoding device can search for a template that is most similar to a template composed of adjacent samples of the current block to find a reconstructed area similar to the current block, and then can use the reconstructed area to determine the position of the reference sample. The shape of the template is Figure 12 . The template may be adjacent to the current block, or may be at a certain distance from the current block. The video decoding device may directly resolve the distance to the template, or may set the distance to the template to a preset value. Alternatively, the video decoding device may resolve an index of an offset in a list indicating the distance. In some embodiments, all or some pixels of the template may be used for the search. For example, when utilizing a template adjacent to the current block, the video decoding device may search for a template that is most similar to the template of the current block to determine a matching block, and may then use surrounding reconstructed samples adjacent to the matching block as reference samples.

[0160] When searching for similar templates, the similarity between templates can be calculated based on the sum of absolute differences (SAD), mean squared error (MSE), sum of squared differences (SSD), sum of absolute transformed differences (SATD), etc. The shape of the template can be determined based on its width and height, such as Figure 13 As shown, the width and height of the template can be determined as predetermined values.

[0161] Figure 14 is a schematic diagram illustrating types of reference samples according to at least one embodiment of the present invention.

[0162] The type is determined by the number of directions in which the reference sample exists, such as the upper side, lower side, left side and / or right side of the reconstruction area. The video decoding device can determine the type of the reference sample as Figure 14 One of the types shown. Figure 14 In the example shown in FIG, both type 2 and type 3 utilize reference samples in two directions. However, type 2 uses reference samples in directions that intersect each other, and type 3 uses reference samples in continuous directions. The video decoding device can directly parse the type, or can set the type to a preset value. The video decoding device can also determine the predetermined direction of the reference sample to be utilized by the selected type. The video decoding device can parse the direction of the reference sample to be utilized by the selected type. Based on the determined type of reference sample, the video decoding device can determine the range of available angle modes.

[0163] Hereinafter, the type information of the reference sample includes the above-mentioned type and direction of the reference sample used.

[0164] In order to indicate whether the embodiment according to the present invention is applicable, the video encoding device can signal sps_intra_reference_copy_enabled_flag or pps_intra_reference_copy_enabled_flag to the video decoding device at a higher level (such as Sequence Parameter Set (SPS), Picture Parameter Set (PPS), etc.).

[0165] <Implementation 1> Using unidirectional intra-frame prediction

[0166] In this embodiment, the video decoding device performs unidirectional intra prediction by using distant reconstructed samples that are not adjacent to the current block as reference samples. Unidirectional intra prediction requires determining an intra prediction mode and a reference sample, the latter of which can be determined by the video decoding device based on the above-mentioned position and type. The intra prediction mode is represented by irc_pred_mode. The video decoding device can parse the prediction mode or set the prediction mode to a preset value. Alternatively, the video decoding device can infer the prediction mode according to a predetermined method. When signaled, the prediction mode irc_pred_mode can be a value within a range of available values depending on the type of reference sample, or it can be a non-angular mode such as plane, DC mode, etc.

[0167] Figure 15 is a diagram illustrating unidirectional intra prediction according to at least one embodiment of the present invention.

[0168] For example, Figure 15 As shown, when a reconstructed area similar to the current block is determined and the reference sample is type 4, the video decoding apparatus copies adjacent samples of the similar reconstructed area to the periphery of the current block. Then, the video decoding apparatus performs prediction of the current block according to the prediction mode indicated by irc_pred_mode.

[0169] The video decoding device can determine the applicability of this embodiment based on irc_flag (which is a flag indicating whether to generate a duplicate reference sample of the current block by using the adjacent samples of the reconstructed area). The video encoding device conveys irc_flag to the video decoding device, and the video decoding device can then parse irc_flag. If irc_flag is parsed as 0, the video decoding device does not use the method according to this embodiment. If irc_flag is parsed as 1, the video decoding device can use the method according to this embodiment. If irc_flag is parsed as 1 and the prediction mode is determined by a signal, the video decoding device can parse irc_pred_mode to determine which prediction mode is used for unidirectional intra prediction.

[0170] <Implementation 2> Using symmetric bidirectional intra-frame prediction

[0171] In this embodiment, the video decoding device performs symmetric bidirectional intra-frame prediction by utilizing distant reconstructed samples that are not adjacent to the current block as reference samples. In symmetric bidirectional intra-frame prediction, when given an angular prediction mode, the video decoding device uses the angular prediction mode and a prediction mode that is 180 degrees symmetrical with the angular prediction mode to generate an intra-frame prediction value for the current block. To generate a final prediction value for a sample in the current block, the video decoding device may weightedly sum the reference samples indicated by the angular prediction mode and the reference samples indicated by the 180-degree symmetric prediction mode to form the angular prediction mode. Alternatively, the video decoding device may weightedly sum the prediction value generated according to the angular prediction mode and the prediction value generated according to the 180-degree symmetric prediction mode to form the angular prediction mode. For the weighted summation, the video decoding device may determine the weight based on the distance between the reference sample used and the sample in the current block. Alternatively, the video decoding device may determine the weight based on a fixed predetermined ratio.

[0172] For symmetric bidirectional intra prediction, it is necessary to determine the intra prediction mode and the reference sample, wherein the reference sample can be determined by the video decoding device according to the above-mentioned position and type. The intra prediction mode is represented by irc_pred_mode. The video decoding device can parse the prediction mode, or set the prediction mode to a preset value. Alternatively, the video decoding device can infer the prediction mode according to a predetermined method. When signaled, the prediction mode irc_pred_mode can be a value within a range of available values depending on the type of the reference sample. Two prediction modes are required to perform symmetric bidirectional intra prediction. However, in this embodiment, irc_pred_mode can specify only one prediction mode, because the determination of one prediction mode automatically determines the other prediction mode.

[0173] Figure 16is a diagram illustrating symmetric bidirectional intra prediction according to at least one embodiment of the present invention.

[0174] For example, Figure 16 As shown, when a reconstructed area similar to the current block is determined and the reference sample is type 4, the video decoding apparatus copies adjacent samples of the similar reconstructed area to the periphery of the current block. The video decoding apparatus then performs prediction of the current block according to the prediction mode indicated by irc_pred_mode and its inverse prediction mode.

[0175] Figure 17 is a schematic diagram illustrating symmetric bidirectional intra prediction according to another embodiment of the present invention.

[0176] Within the range of available intra prediction modes determined according to the type of reference sample, there may not be a prediction mode and a 180-degree symmetric prediction mode previously determined according to signaling and a predetermined method. If so, the video decoding apparatus may perform a weighted summation of reference samples guided by the previously determined prediction mode and the 180-degree symmetric prediction mode in adjacent reconstructed samples of the current block. For example, Figure 17 As shown, when a reconstructed area similar to the current block is determined and the reference sample is type 1, the video decoding apparatus copies adjacent samples of the similar reconstructed area to the periphery of the current block. Assume that the prediction mode irc_pred_mode is determined by parsing. Figure 17 In the example, the prediction mode in the direction opposite to irc_pred_mode is not included in the range of available intra prediction modes, and the reconstructed samples in this direction exist outside the current block. Therefore, the video decoding device can select the sample indicated by irc_pred_mode from the copied samples and select the reference sample indicated by the 180-degree symmetric prediction mode from the adjacent reconstructed samples of the current block.

[0177] The video decoding device can determine the applicability of this embodiment based on irc_flag (which is a flag indicating whether to generate a replicated reference sample of the current block by using adjacent samples of the reconstruction area). If irc_flag is parsed as 0, the video decoding device does not use the method according to this embodiment. If irc_flag is parsed as 1, the video decoding device can continue to apply the method according to this embodiment. If irc_flag is parsed as 1 and an angular prediction mode is determined by signaling, the video decoding device can parse irc_pred_mode to determine an angular prediction mode for symmetric bidirectional intra prediction.

[0178] <Implementation 3> Using asymmetric bidirectional intra-frame prediction

[0179] In this embodiment, a video decoding device performs asymmetric bidirectional intra prediction by utilizing distant reconstructed samples that are not adjacent to the current block as reference samples. In asymmetric bidirectional intra prediction, when given an angular prediction mode, the video decoding device utilizes the angular prediction mode and a prediction mode with a direction that is not 180 degrees symmetric with the angular prediction mode to generate an intra prediction value for the current block. To generate a final prediction value for a sample in the current block, the video decoding device may weightedly sum the reference samples indicated by the angular prediction mode and the reference samples indicated by the prediction mode with a direction that is not 180 degrees symmetric with the angular prediction mode, forming the angular prediction mode. Alternatively, the video decoding device may weightedly sum the prediction value generated based on the angular prediction mode and the prediction value generated based on the prediction mode with a direction that is not 180 degrees symmetric with the angular prediction mode, forming the angular prediction mode. For this weighted summation, the video decoding device may determine the weight based on the distance between the reference samples used and the samples in the current block. Alternatively, the video decoding device may determine the weight based on a fixed predetermined ratio.

[0180] For asymmetric bidirectional intra prediction, it is necessary to determine the intra prediction mode and reference samples, where the reference samples can be determined by the video decoding device based on the above-mentioned position and type. The intra prediction mode is represented by irc_pred_mode. The video decoding device can parse the prediction mode or set the prediction mode to a preset value. Alternatively, the video decoding device can infer the prediction mode according to a predetermined method. When signaled, the prediction mode irc_pred_mode can be a value within a range of available values depending on the type of reference sample, or a non-angular prediction mode such as plane, DC mode, etc. Two prediction modes are required to perform asymmetric bidirectional intra prediction. The video decoding device can determine the two prediction modes by parsing the two prediction modes. Alternatively, the video decoding device can parse one prediction mode and derive the other prediction mode on the decoder side. Alternatively, if one prediction mode is determined according to a predetermined method without signaling, the video decoding device can determine the other prediction mode by parsing, or can derive the other prediction mode on the decoder side. When two prediction modes are signaled, irc_pred_mode may be in the form of a sequence of values of the two prediction modes, or may be in the form of an index indicating a pair containing values of the two prediction modes.

[0181] When one prediction mode has been determined as described above, another prediction mode can be derived on the decoder side according to any of the following methods.

[0182] ① Using intra-frame mode derivation technology on the decoder side

[0183] The template-based intra mode derivation (TIMD) technique derives an intra prediction mode from a neighboring template of a current block at the decoder side, and then uses the derived prediction mode to generate a prediction block for the current block. The decoder-side intra mode derivation (DIMD) technique calculates the gradient of each of the neighboring samples of the current block, and then uses the calculated gradient to derive the intra prediction mode for predicting the current block. As described above, asymmetric bidirectional intra prediction requires two different prediction modes. When a prediction mode is determined as described above, that is, determined according to signaling, determined as a preset value, or determined according to a predetermined method, the video decoding device can derive another prediction mode by utilizing the decoder-side intra mode derivation technique as described above. The derived prediction mode may not be 180 degrees symmetrical with the prediction mode sent by the signal, and may be within a range of available values depending on the type of reference sample.

[0184] ② Intra-mode derivation based on the signaled prediction mode

[0185] The video decoding device generates a prediction value based on one of the prediction modes determined as described above (i.e., determined by signaling, determined as a preset value, or determined according to a predetermined method), and then uses the prediction mode that generates a prediction value most similar to the generated prediction value as the remaining prediction mode. In this case, the two prediction modes of asymmetric bidirectional intra-frame prediction are asymmetric to each other, so the prediction mode that is 180 degrees symmetric to the signaled prediction mode can be excluded from the search range. The similarity between the two prediction values can be calculated based on SAD, MSE, SSD, SATD, etc.

[0186] Figure 18 is a diagram illustrating asymmetric bidirectional intra prediction according to at least one embodiment of the present invention.

[0187] For example, Figure 18 As shown in FIG, when a reconstructed region similar to the current block is determined and the reference sample is type 4, the video decoding apparatus copies adjacent samples of the similar reconstructed region to the periphery of the current block. Then, the video decoding apparatus performs prediction of the current block by using the reconstructed samples indicated by intra mode 1 (prediction mode indicated by irc_pred_mode) and intra mode 2 (prediction mode derived from DIMD), respectively.

[0188] Figure 19 is a schematic diagram illustrating asymmetric bidirectional intra prediction according to another embodiment of the present invention.

[0189] Together with one intra-frame prediction mode first determined according to signaling and a predetermined method, this prediction mode can be used even if the prediction mode does not exist in the range of available intra-frame prediction modes determined by the type of reference sample. This means that when the remaining one prediction mode is determined by signaling or derivation, the range of available intra-frame prediction modes does not limit the value of the remaining one prediction mode. The video decoding device can select a reconstructed sample indicated by the remaining one prediction mode from the neighboring reconstructed samples of the current block. The video decoding device can weightedly sum the reference samples in the direction indicated by one intra-frame prediction mode and the reconstructed samples in the direction indicated by another prediction mode.

[0190] For example, Figure 19 As shown, when a reconstructed area similar to the current block is determined and the reference sample is type 1, the video decoding device copies adjacent samples of the similar reconstructed area to the current block. It is assumed that the two prediction modes used in the asymmetric bidirectional intra prediction are determined by parsing. Figure 19 In the example of , the remaining one prediction mode is not included in the range of available intra prediction modes, and the reconstructed samples in the corresponding direction exist outside the current block. Therefore, the video decoding device can select samples indicated by one of the two prediction modes indicated by irc_pred_mode from the copied samples, and select reference samples indicated by the other prediction mode from the reconstructed samples adjacent to the current block.

[0191] The video decoding device can determine the applicability of this embodiment based on irc_flag (which is a flag indicating whether to generate a replicated reference sample of the current block by utilizing adjacent samples of the reconstructed area). If irc_flag is parsed as 0, the video decoding device does not use the method according to this embodiment. If irc_flag is parsed as 1, the video decoding device can continue to apply the method according to this embodiment. If irc_flag is parsed as 1 and a prediction mode is determined with the signal, the video decoding device can parse irc_pred_mode to determine which prediction mode is used for asymmetric bidirectional intra-frame prediction. When two prediction modes are signaled, irc_pred_mode can indicate both prediction modes. When only one prediction mode is signaled and the other prediction mode is derived, irc_pred_mode can represent one prediction mode. When the first determined prediction mode is also determined as a preset value or determined by a predetermined method without being signaled, the video decoding device does not perform separate parsing of the prediction mode.

[0192] <Implementation 4> Using multiple prediction values

[0193] In this embodiment, the video decoding device generates a final prediction value of the current block using multiple prediction values. The prediction value may include (1) a prediction value using a distant reconstructed sample that is not adjacent to the current block as a reference sample, (2) a reconstructed area similar to the current block, (3) a prediction value generated based on a reconstructed sample adjacent to the current block, and the like. The video decoding device may generate a final prediction value using at least two prediction values including the (1) prediction value. For example, the final prediction value may be generated by relying on a total of three cases: "(1) prediction value and (2) prediction value", "(1) prediction value and (3) prediction value", and "(1) prediction value, (2) prediction value, and (3) prediction value". The (1) prediction value may be generated according to any one of the above-mentioned embodiments 1 to 3. The weight corresponding to each prediction value may be determined according to a predetermined method and used in the weighted summation of the prediction values, as shown in Equation 1.

[0194] [Equation 1]

[0195] pred cur =w1×pred1+w2×pred2+w3×pred3

[0196] Unlike the (2) prediction value generated by using the reconstructed region, the video decoding device determines the prediction mode. In order to minimize the signaling cost of determining the prediction mode, the video decoding device can generate the (3) prediction value by using the same prediction mode as the irc_pred_mode used for the (1) prediction value. The video decoding device can generate the (3) prediction value by using the prediction mode determined by the prediction mode derivation method described in Embodiment 3. Alternatively, the prediction mode irc_cur_pred_mode used for the (3) prediction value can be further signaled.

[0197] The video decoding device can determine the applicability of this embodiment based on the flag irc_flag (which indicates whether to generate a copied reference sample of the current block by utilizing the adjacent samples of the reconstructed area). If irc_flag is parsed as 0, the video decoding device does not use the method according to this embodiment. If irc_flag is parsed as 1, the video decoding device can continue to apply the method according to this embodiment. If irc_flag is parsed as 1 and irc_pred_mode is parsed to determine the prediction mode for (1) prediction value, the video decoding device parses irc_pred_mode to determine the prediction mode for one of embodiments 1 to 3. For example, when it is determined to use (3) prediction value and further transmit the prediction mode to generate (3) prediction value, the video decoding device can parse irc_pred_mode to determine the prediction mode for (1) prediction value, and then parse irc_cur_pred_mode to determine the prediction mode for (3) prediction value.

[0198] Now refer to Figure 20 and Figure 21 , a method for predicting a current block by utilizing distant reference samples is described.

[0199] Figure 20 is a flowchart of a method of encoding a current block by a video encoding apparatus according to at least one embodiment of the present invention.

[0200] The video encoding apparatus obtains position information of adjacent reference samples of a reconstructed region (S2000). Here, the reconstructed region is similar to the current block and is far away from the current block.

[0201] The position information of the adjacent reference samples includes the position of the pixel, the vector indicating the pixel, and the offset. Figure 11 As shown, the pixels represent those located on the upper left, upper right, lower left, and lower right sides of a reconstructed region similar to the current block. Alternatively, the pixels represent those located at preset positions within a reference line where adjacent reference samples exist. The vector represents the displacement between the current block and the pixel. The offset represents the distance between the reconstructed region and the reference line.

[0202] The video encoding apparatus may obtain the position information of the adjacent reference samples from a higher level. The video encoding apparatus may determine the position information of the adjacent reference samples in terms of rate-distortion optimization.

[0203] Alternatively, the video encoding apparatus may find the reconstructed region by searching for a template that is most similar to a template composed of adjacent samples of the current block. The video encoding apparatus may then use the found reconstructed region to determine the positions of the adjacent reference samples.

[0204] The video encoding apparatus obtains type information of adjacent reference samples ( S2002 ).

[0205] The type information of the adjacent reference samples includes the type of the adjacent reference samples and the direction of the adjacent reference samples used. Figure 14 As shown, the type of the adjacent reference samples can be determined according to the number of directions in which the adjacent reference samples exist around the reconstruction area.

[0206] The video encoding apparatus may obtain the type information of the adjacent reference samples from a higher level. The video encoding apparatus may determine the type information of the adjacent reference samples in terms of rate-distortion optimization. Alternatively, the video encoding apparatus may set the type information of the adjacent reference samples to a preset value.

[0207] The video encoding apparatus obtains an intra prediction mode of a current block ( S2004 ).

[0208] The video encoding device may obtain the intra-frame prediction mode from a higher level. The video encoding device may determine the intra-frame prediction mode in terms of rate-distortion optimization. The video encoding device may set the intra-frame prediction mode to a preset value. Alternatively, the video encoding device may infer the intra-frame prediction mode based on a preset method.

[0209] The video encoding apparatus copies the adjacent reference samples of the reconstructed area to the periphery of the current block using the position information of the adjacent reference samples and their type information, thereby generating copied reference samples of the current block ( S2006 ).

[0210] The video encoding apparatus generates a first prediction block of the current block using the copied reference samples according to the intra prediction mode ( S2008 ).

[0211] The video encoding apparatus generates a second prediction block for the current block using adjacent samples of the current block according to an intra-frame prediction mode (S2010). The video encoding apparatus may generate the second prediction block by using reference samples adjacent to the current block. The video encoding apparatus may generate the second prediction block by using reference samples according to MRL. Alternatively, the video encoding apparatus may generate the second prediction block for the current block by using IBC technology.

[0212] The video encoding apparatus determines a flag based on the first prediction block and the second prediction block (S2012). Here, the flag indicates whether to generate a copied reference sample of the current block by using adjacent reference samples of the reconstruction area.

[0213] The video encoding device may determine the value of the flag in terms of rate-distortion optimization. For example, if the first prediction block is optimal, the flag may be set to true. On the other hand, if the second prediction block is optimal, the flag may be set to false. The video encoding device may subtract the prediction block depending on the flag value from the current block to generate a residual block.

[0214] The video encoding apparatus encodes the flag (S2014). In addition, the video encoding apparatus may encode the residual block.

[0215] Figure 21 is a flowchart of a method for reconstructing a current block by a video decoding device according to at least one embodiment of the present invention.

[0216] The video decoding apparatus obtains an intra prediction mode of a current block ( S2100 ).

[0217] The video decoding apparatus may decode the intra prediction mode from the bitstream. The video decoding apparatus may set the intra prediction mode to a preset value. Alternatively, the video decoding apparatus may infer the intra prediction mode according to a preset method.

[0218] The video decoding apparatus decodes a flag indicating whether to generate a copied reference sample of the current block by using an adjacent reference sample of a reconstructed region ( S2102 ). Here, the reconstructed region is similar to the current block and is distant from the current block.

[0219] The video decoding device checks the flag (S2104).

[0220] If the flag is true (Yes in S2104 ), the video decoding apparatus performs the following steps.

[0221] The video decoding apparatus obtains position information of adjacent reference samples of the reconstructed region ( S2106 ).

[0222] The video decoding apparatus may decode location information of adjacent reference samples from a bitstream.

[0223] Alternatively, the video decoding apparatus finds the reconstructed region by searching for a template that is most similar to a template composed of adjacent samples of the current block. The video decoding apparatus can then use the found reconstructed region to determine the positions of adjacent reference samples.

[0224] The video decoding apparatus obtains type information of adjacent reference samples ( S2108 ).

[0225] The video decoding apparatus may obtain the type information of the neighboring reference samples from the bitstream. Alternatively, the video decoding apparatus may set the type information of the neighboring reference samples to a preset value.

[0226] The video decoding apparatus copies the adjacent reference samples of the reconstructed area to the periphery of the current block using the position information of the adjacent reference samples and their type information, thereby generating copied reference samples of the current block ( S2110 ).

[0227] The video decoding apparatus generates a prediction block of the current block using the copied reference samples according to the intra prediction mode ( S2112 ).

[0228] On the other hand, if the flag is false (No at S2104), the video decoding apparatus generates a prediction block for the current block by utilizing adjacent samples of the current block according to the intra prediction mode (S2120). The video decoding apparatus may generate the prediction block by utilizing reference samples adjacent to the current block. The video decoding apparatus may generate the prediction block by utilizing reference samples according to MRL. Alternatively, the video decoding apparatus may generate the prediction block for the current block by utilizing IBC technology.

[0229] Then, the video decoding apparatus decodes the residual block of the current block and may sum the residual block and the prediction block to generate a reconstructed block of the current block.

[0230] Although the steps in the various flowcharts are described as being performed sequentially, these steps merely illustrate the technical concepts of some embodiments of the present invention. Therefore, a person skilled in the art to which the present invention pertains can implement the steps by changing the order in which they are described in the various figures or by performing two or more steps in parallel. Therefore, the steps in the various flowcharts are not limited to the chronological order shown.

[0231] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or a combination thereof. It should also be understood that the functional components described in the present invention are labeled as "units" to highlight their ability to be implemented independently.

[0232] On the other hand, the various methods or functions described in some embodiments can be implemented as instructions stored in a non-volatile recording medium, which can be read and executed by one or more processors. The non-volatile recording medium can include, for example, various types of recording devices that store data in a form readable by a computer system. For example, the non-volatile recording medium can include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical disk drives, magnetic hard drives, and solid-state drives (SSDs), etc.

[0233] Although exemplary embodiments of the present invention have been described for illustrative purposes, it will be understood by those skilled in the art that various modifications, additions, and substitutions may be made without departing from the spirit and scope of the present invention. Therefore, embodiments of the present invention have been described for brevity and clarity. The scope of the technical ideas of the embodiments of the present invention is not limited by the examples. Accordingly, it will be understood by those skilled in the art that the scope of the present invention should not be limited by the embodiments clearly described above, but by the claims and their equivalents.

[0234] Reference numerals

[0235] 122: Intra-frame predictor

[0236] 155: Entropy Encoder

[0237] 510: Entropy Decoder

[0238] 542: Intra-frame predictor.

[0239] CROSS-REFERENCE TO RELATED APPLICATIONS

[0240] This application claims priority to and the benefit of Korean Patent Application No. 10-2023-0005209, filed on January 13, 2023, and Korean Patent Application No. 10-2023-0192850, filed on December 27, 2023, the entire contents of each of which are incorporated herein by reference.

Claims

1. A method for reconstructing a current block by a video decoding device, the method comprising: Get the intra prediction mode of the current block; Obtaining position information of adjacent reference samples of a reconstructed area that is similar to the current block and far away from the current block; Obtaining type information of adjacent reference samples; By using the position information of the adjacent reference samples and the type information of the adjacent reference samples, the adjacent reference samples of the reconstructed area are copied to the periphery of the current block, thereby generating copied reference samples of the current block; as well as By using the copied reference samples, a prediction block of the current block is generated according to the intra prediction mode.

2. The method according to claim 1, wherein The location information includes: The position of the pixel, the vector indicating the pixel, and the offset.

3. The method according to claim 2, wherein: The pixels represent pixels located at the upper left part, the upper right part, the lower left part, and the lower right part within the reconstruction area, or represent pixels at preset positions within a reference line where an adjacent reference sample exists, the vector represents a displacement between the current block and the pixel, and the offset represents a distance between the reconstruction area and the reference line.

4. The method according to claim 2, wherein: Obtaining location information includes: The location information of neighboring reference samples is decoded from the bitstream.

5. The method according to claim 1, wherein Obtaining location information includes: Locating a reconstruction area by searching for a template that is most similar to a template including neighboring samples of the current block; and The locations of neighboring reference samples are determined by utilizing the reconstructed region.

6. The method according to claim 1, wherein The type information includes the type of the used adjacent reference sample and the direction of the used adjacent reference sample. The type of the adjacent reference sample is determined by the number of directions in which the adjacent reference samples exist around the reconstruction area.

7. The method according to claim 6, wherein: Obtaining type information includes: The type information of the adjacent reference samples is decoded from the bitstream, or the type information is set to a preset value.

8. The method according to claim 1, further comprising: decoding a flag indicating whether to generate copied reference samples of the current block by utilizing neighboring reference samples of the reconstruction region; as well as Check the flags, Wherein, the method further comprises, when the flag is true: The position information is then obtained by generating a prediction block.

9. The method according to claim 1, wherein Obtaining intra prediction modes includes: The intra prediction mode is decoded from the bitstream, and the intra prediction mode is set to a preset value, or the intra prediction mode is inferred according to a preset method.

10. The method according to claim 1, wherein Generating the prediction block includes, when the intra prediction mode is an angular prediction mode: generating, according to an intra prediction mode, a first prediction block of the current block by utilizing the copied reference samples; Generate a second prediction block of the current block by using a copied reference sample or a neighboring sample of the current block according to a prediction mode in a direction 180 degrees opposite to the intra prediction mode; as well as A weighted sum is performed on the first prediction block and the second prediction block.

11. A method for encoding a current block by a video encoding apparatus, the method comprising: Obtaining position information of adjacent reference samples of a reconstructed area that is similar to the current block and far away from the current block; Obtaining type information of adjacent reference samples; Get the intra prediction mode of the current block; By using the position information of the adjacent reference samples and the type information of the adjacent reference samples, the adjacent reference samples of the reconstructed area are copied to the periphery of the current block, thereby generating copied reference samples of the current block; as well as According to the intra prediction mode, a first prediction block of the current block is generated by using the copied reference samples.

12. The method according to claim 11, further comprising: According to the intra prediction mode, a second prediction block of the current block is generated by using neighboring samples of the current block.

13. The method according to claim 12, further comprising: determining, based on the first prediction block and the second prediction block, a flag indicating whether to generate a copied reference sample of the current block by using a neighboring reference sample of the reconstruction area; as well as The flag is encoded.

14. The method according to claim 11, wherein Obtaining location information includes: The position information of adjacent reference samples is obtained from a high level.

15. The method according to claim 11, wherein Obtaining location information includes: Locating a reconstruction area by searching for a template that is most similar to a template including neighboring samples of the current block; and The locations of neighboring reference samples are determined by utilizing the reconstructed region.

16. The method according to claim 11, wherein Obtaining type information includes: The type information of the adjacent reference samples is obtained from a high level, or the type information is set to a preset value.

17. A computer-readable recording medium storing a bitstream generated by a video encoding method, the video encoding method comprising: Obtaining position information of adjacent reference samples of a reconstructed area that is similar to the current block and far away from the current block; Obtaining type information of adjacent reference samples; Get the intra prediction mode of the current block; By using the position information of the adjacent reference samples and the type information of the adjacent reference samples, the adjacent reference samples of the reconstructed area are copied to the periphery of the current block, thereby generating copied reference samples of the current block; as well as A prediction block of the current block is generated according to an intra prediction mode by using the copied reference sample.

Citation Information

Patent Citations

  • Larazotide derivatives containing D-amino acids

    KR1020230005209A