Method and device for effective motion vector prediction of intra block copy, and recording medium
By employing template matching to generate prediction vectors within the within-screen block copy mode, the method addresses the inefficiencies in motion vector prediction, resulting in improved video encoding and decoding performance.
Patent Information
- Application Number
- PCT/KR2023/022032
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2023-12-29
- Publication Date
- 2025-06-12
AI Technical Summary
The performance of motion vector prediction in the within-screen block copy mode is degraded when there are no available blocks, leading to inefficiencies in video encoding and decoding processes.
The proposed method improves motion vector prediction by generating a prediction vector using template matching within the within-screen block copy mode. This involves configuring a candidate list of predicted motion vectors, obtaining motion vectors from adjacent blocks or pre-encoded blocks, and applying template matching with weights based on distance relationships.
This approach enhances the accuracy and efficiency of motion vector prediction, reducing the need for additional prediction vectors and improving overall video encoding and decoding performance.
Smart Images

Figure KR2023022032_12062025_PF_FP_ABST
Abstract
Description
Method, device and recording medium for effective motion vector prediction of block copy within a screen
[0001] The present invention relates to a method and device for encoding / decoding a video signal.
[0002] The demand for high-resolution, high-quality images is growing across a wide range of applications. As image data becomes higher in resolution and quality, the relative amount of data increases compared to conventional image data. Therefore, transmitting image data using existing media such as wired or wireless broadband lines or storing it using existing storage media incurs increased transmission and storage costs. To address these issues arising from the increasing resolution and quality of image data, high-efficiency image compression technologies can be utilized.
[0003] In a video encoding / decoding method, the intra block copy mode can perform motion prediction by referencing a previously decoded area of the same frame to generate a prediction block for the current decoded block, and generate a prediction block based on the result. At this time, if the upper block or the left block among the surrounding previously decoded blocks is predicted by intra block copy, the motion vector of the corresponding block is used as the prediction vector for the current decoded block. However, if there is no available block, the motion vector of the last decoded block or a specific vector can be used as the prediction vector. In this case, the performance of the motion vector prediction may be degraded.
[0004] The video encoding / decoding method, device, and recording medium of the present invention may include a step of determining whether to encode a current encoding block in a block-in-picture copy mode, a step of constructing a candidate list of a predicted motion vector of the current encoding block in response to a decision to encode the current encoding block in the block-in-picture copy mode, a step of obtaining a motion vector of the current encoding block based on a candidate of the predicted motion vector candidate list, and a step of encoding the current encoding block based on the motion vector.
[0005] In the video encoding / decoding method, device and recording medium of the present invention, information indicating the number of candidates in the predicted motion vector candidate list can be encoded in a bitstream.
[0006] In the video encoding / decoding method, device, and recording medium of the present invention, the predicted motion vector candidate list may include at least one of motion vectors of adjacent blocks encoded in a block copy mode within a screen, motion vectors of blocks pre-encoded in a block copy mode within a screen, or motion vectors generated through template matching.
[0007] In the video encoding / decoding method, device and recording medium of the present invention, the adjacent block may be at least one of a left adjacent block adjacent to the left of the current encoding block or an upper adjacent block adjacent to the top of the current encoding block.
[0008] In the video encoding / decoding method, device, and recording medium of the present invention, the left adjacent block may be the lowermost block among the plurality of blocks when there are a plurality of blocks adjacent to the left of the current encoding block.
[0009] In the video encoding / decoding method, device and recording medium of the present invention, the upper adjacent block may be the rightmost block among the plurality of blocks when there are a plurality of blocks adjacent to the upper side of the current encoding block.
[0010] In the video encoding / decoding method, device and recording medium of the present invention, the motion vector of a block pre-encoded in the block copy mode within the screen may be any one of the blocks determined based on the distance relationship with the current encoding block or the encoding order among the plurality of blocks when there are a plurality of blocks pre-encoded in the block copy mode within the screen.
[0011] In the video encoding / decoding method, device and recording medium of the present invention, the template matching can be performed with a template determined based on surrounding pixels of the current encoding block.
[0012] In the video encoding / decoding method, device and recording medium of the present invention, the template matching can be performed by a weight determined according to the distance relationship between the surrounding pixels and the current encoding block.
[0013] In the video encoding / decoding method, device and recording medium of the present invention, the motion vector generated through the template matching can be obtained based on motion vectors obtained by applying a weight to each template when there are multiple templates of the current encoding block.
[0014] In the image encoding / decoding method, device and recording medium of the present invention, the template matching target area of the template matching can be determined using the decimal point position of the pixel.
[0015] In the video encoding / decoding method, device, and recording medium of the present invention, the predicted motion vector candidate list may further include a motion vector to which a correction offset is applied to a candidate of the predicted motion vector candidate list.
[0016] The proposed invention aims to improve the performance of motion vector prediction in the within-screen block copy mode by generating a prediction vector in the within-screen block copy mode based on template matching.
[0017] FIG. 1 is a schematic block diagram of an encoding device as an embodiment of the present invention.
[0018] FIG. 2 is a schematic block diagram of a decryption device as an embodiment of the present invention.
[0019] FIG. 3 illustrates a flowchart of an encoder according to one embodiment of the present invention.
[0020] Figure 4 illustrates a process of obtaining a template from surrounding pixels of a current encoding block for template matching in a block copy within a screen.
[0021] Figure 5 illustrates an embodiment in which motion prediction is performed using a surrounding pre-decoding block of the current encoding block as a template.
[0022] Figure 6 illustrates a decoder flowchart according to one embodiment of the present invention.
[0023] Fig. 7(a) illustrates a method for generating a prediction block in a block copy mode within a screen, and Fig. 7(b) illustrates the location of a pre-decoded block used for motion vector prediction.
[0024] Figure 8 illustrates an example of performing template matching based on a template.
[0025] Figure 9 illustrates an example of obtaining a predicted motion vector candidate using a base-based / decoded surrounding block.
[0026] The present invention is susceptible to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present invention. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.
[0027] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, a first component may be referred to as a "second component," and similarly, a second component may also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0028] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0029] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not exclude in advance the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0030] Hereinafter, with reference to the attached drawings, preferred embodiments of the present invention will be described in more detail. Hereinafter, identical components in the drawings will be designated by the same reference numerals, and redundant descriptions of identical components will be omitted.
[0031] In the present invention, the terms current encoding block, current decoding block, and current block may be used interchangeably.
[0032]
[0033] FIG. 1 is a schematic block diagram of an encoding device as an embodiment of the present invention.
[0034] Referring to FIG. 1, the encoding device (100) may include a picture segmentation unit (110), a prediction unit (120, 125), a transformation unit (130), a quantization unit (135), a reordering unit (160), an entropy encoding unit (165), an inverse quantization unit (140), an inverse transformation unit (145), a filter unit (150), and a memory (155).
[0035] Each component shown in Fig. 1 is independently depicted to represent different characteristic functions in the video encoding device, which may mean that each component is composed of separate hardware. However, each component is listed and included as a separate component for convenience of explanation, and at least two of each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0036] Additionally, some components may not be essential components that perform essential functions of the present invention, but may be optional components merely used to enhance performance. The present invention may be implemented by including only components essential to implementing the essence of the present invention, excluding components used solely for performance enhancement. A structure that includes only essential components, excluding optional components used solely for performance enhancement, is also within the scope of the present invention.
[0037] The picture segmentation unit (110) can segment the input picture into at least one block. At this time, the block may mean a coding unit (CU), a prediction unit (PU), or a transformation unit (TU). The segmentation may be performed based on at least one of a quadtree, a binary tree, and a triple tree. A quadtree is a method of dividing an upper block into four lower blocks whose width and height are half of those of the upper block. A binary tree is a method of dividing an upper block into lower blocks whose width or height is half of that of the upper block. In a binary tree, the upper block has a height of half. Through the segmentation based on the binary tree described above, a block may have a shape that is not only square but also non-square.
[0038] Hereinafter, in the embodiments of the present invention, the encoding unit may be used to mean a unit that performs encoding or may be used to mean a unit that performs decoding.
[0039] The prediction unit (120, 125) may include an inter prediction unit (120) that performs inter prediction and an intra prediction unit (125) that performs intra prediction. It may be determined whether to use inter prediction or intra prediction for a prediction unit, and specific information (e.g., intra prediction mode, motion vector, reference picture, etc.) according to each prediction method may be determined. At this time, the processing unit where prediction is performed and the processing unit where the prediction method and specific contents are determined may be different. For example, the prediction method and prediction mode may be determined in the prediction unit, and the prediction may be performed in the transformation unit. The residual value (residual block) between the generated prediction block and the original block may be input to the transformation unit (130). In addition, the prediction mode information, motion vector information, etc. used for prediction may be encoded together with the residual value in the entropy encoding unit (165) and transmitted to the decoding device. When using a specific encoding mode, it is also possible to encode the original block as is and transmit it to the decoding unit without generating a prediction block through the prediction unit (120, 125).
[0040] The inter prediction unit (120) may predict a prediction unit based on information of at least one picture among the previous or subsequent pictures of the current picture, and in some cases, may predict a prediction unit based on information of a portion of an encoded region within the current picture. The inter prediction unit (120) may include a reference picture interpolation unit, a motion prediction unit, and a motion compensation unit.
[0041] The reference picture interpolation unit can receive reference picture information from the memory (155) and generate pixel information less than an integer pixel from the reference picture. In the case of luminance pixels, a DCT-based 8-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 4 pixels. In the case of a chrominance signal, a DCT-based 4-tap interpolation filter with different filter coefficients can be used to generate pixel information less than an integer pixel in units of 1 / 8 pixels.
[0042] The motion prediction unit can perform motion prediction based on a reference picture interpolated by the reference picture interpolation unit. Various methods can be used to derive a motion vector, such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm). The motion vector can have a motion vector value in units of 1 / 2 or 1 / 4 pixels based on the interpolated pixel. The motion prediction unit can predict the current prediction unit by using different motion prediction methods. Various methods can be used as motion prediction methods, such as the Skip method, the Merge method, and the AMVP (Advanced Motion Vector Prediction) method.
[0043] The intra prediction unit (125) can generate a prediction unit based on reference pixel information surrounding the current block, which is pixel information within the current picture. If the surrounding block of the current prediction unit is a block on which inter prediction has been performed and the reference pixel is a pixel on which inter prediction has been performed, the reference pixel included in the block on which inter prediction has been performed can be replaced and used with reference pixel information of the surrounding block on which intra prediction has been performed. That is, if the reference pixel is not available, the unavailable reference pixel information can be replaced and used with at least one reference pixel among the available reference pixels.
[0044] In intra prediction, the prediction mode can have a directional prediction mode that uses reference pixel information according to the prediction direction, and a non-directional mode that does not use directional information when performing prediction. The mode for predicting luminance information and the mode for predicting chrominance information can be different, and the intra prediction mode information used to predict luminance information or the predicted luminance signal information can be utilized to predict chrominance information.
[0045] The intra prediction method can generate a prediction block after applying an Adaptive Intra Smoothing (AIS) filter to a reference pixel according to a prediction mode. The type of AIS filter applied to the reference pixel may be different. In order to perform the intra prediction method, the intra prediction mode of the current prediction unit can be predicted from the intra prediction modes of prediction units existing around the current prediction unit. When the prediction mode of the current prediction unit is predicted using the mode information predicted from the surrounding prediction units, if the intra prediction modes of the current prediction unit and the surrounding prediction units are the same, information indicating that the prediction modes of the current prediction unit and the surrounding prediction units are the same can be transmitted using predetermined flag information, and if the prediction modes of the current prediction unit and the surrounding prediction units are different, entropy encoding can be performed to encode the prediction mode information of the current block.
[0046] Additionally, a residual block containing residual value information, which is the difference between the prediction unit that performed the prediction based on the prediction unit generated in the prediction unit (120, 125) and the original block of the prediction unit, can be generated. The generated residual block can be input to the transformation unit (130).
[0047] In the transformation unit (130), a residual block including residual data can be transformed using a transformation type such as DCT, DST, etc. At this time, the transformation method can be determined based on the intra prediction mode of the prediction unit used to generate the residual block.
[0048] The quantization unit (135) can quantize values converted to the frequency domain by the transformation unit (130). The quantization coefficients can vary depending on the block or the importance of the image. The values produced by the quantization unit (135) can be provided to the dequantization unit (140) and the reordering unit (160).
[0049] The rearrangement unit (160) can perform rearrangement of coefficient values for quantized residual values.
[0050] The reordering unit (160) can change a two-dimensional block-shaped coefficient into a one-dimensional vector form through a coefficient scanning method. For example, the reordering unit (160) can scan from the DC coefficient to the high-frequency region coefficient using a predetermined scan type and change it into a one-dimensional vector form.
[0051] The entropy encoding unit (165) can perform entropy encoding based on the values produced by the rearrangement unit (160). Entropy encoding can use various encoding methods such as, for example, Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC).
[0052] The entropy encoding unit (165) can encode various information such as residual value coefficient information of an encoding unit, block type information, prediction mode information, division unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information from the rearrangement unit (160) and the prediction unit (120, 125).
[0053] The entropy encoding unit (165) can entropy encode the coefficient values of the encoding unit input from the rearrangement unit (160).
[0054] The inverse quantization unit (140) and the inverse transformation unit (145) inversely quantize the values quantized in the quantization unit (135) and inversely transform the values transformed in the transformation unit (130). The residual values generated in the inverse quantization unit (140) and the inverse transformation unit (145) can be combined with the predicted prediction units predicted through the motion estimation unit, motion compensation unit, and intra prediction unit included in the prediction unit (120, 125) to generate a reconstructed block.
[0055] The filter unit (150) may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter).
[0056] A deblocking filter can remove block distortion caused by boundaries between blocks in a reconstructed picture. To determine whether to perform deblocking, a deblocking filter can be applied to the current block based on the pixels contained in several columns or rows within the block. When applying a deblocking filter to a block, a strong filter or a weak filter can be applied depending on the required deblocking filtering strength. Furthermore, when applying a deblocking filter, horizontal and vertical filtering can be processed in parallel when performing vertical and horizontal filtering.
[0057] The offset correction unit can correct the offset from the original image on a pixel-by-pixel basis for an image that has undergone deblocking. To perform offset correction for a specific picture, the pixels contained in the image can be divided into a certain number of regions, the regions to be offset can be determined, and the offset can be applied to those regions. Alternatively, the offset can be applied by considering the edge information of each pixel.
[0058] Adaptive Loop Filtering (ALF) can be performed based on the comparison of the filtered restored image with the original image. After dividing the pixels included in the image into predetermined groups, a filter to be applied to each group can be determined, and filtering can be performed differentially for each group. Information regarding whether to apply ALF can be transmitted by luminance signal for each coding unit (CU), and the shape and filter coefficients of the ALF filter to be applied can vary depending on each block. Furthermore, an ALF filter of the same form (fixed form) can be applied regardless of the characteristics of the target block.
[0059] The memory (155) can store a restored block or picture produced through the filter unit (150), and the stored restored block or picture can be provided to the prediction unit (120, 125) when performing inter prediction.
[0060]
[0061] FIG. 2 is a schematic block diagram of a decryption device as an embodiment of the present invention.
[0062] Referring to FIG. 2, the decryption device (200) may include an entropy decryption unit (210), a rearrangement unit (215), an inverse quantization unit (220), an inverse transformation unit (225), a prediction unit (230, 235), a filter unit (240), and a memory (245).
[0063] Each component shown in FIG. 2 is independently depicted to represent different characteristic functions in the decryption device, which may mean that each component is composed of separate hardware. However, each component is listed and included as a separate component for convenience of explanation, and at least two of each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0064] The entropy decoding unit (210) can perform entropy decoding on the input bitstream. For example, various methods such as Exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), and Context-Adaptive Binary Arithmetic Coding (CABAC) can be applied for entropy decoding.
[0065] The entropy decoding unit (210) can decode information related to intra prediction and inter prediction performed in the encoding device.
[0066] The reordering unit (215) can reorder the bitstream that has been entropy-decoded by the entropy decoding unit (210). The coefficients expressed in the form of a one-dimensional vector can be reordered by restoring them back to coefficients in the form of a two-dimensional block. The reordering unit (215) can receive information related to coefficient scanning performed by the encoding device and perform reordering by performing reverse scanning based on the scanning order performed by the corresponding encoding device.
[0067] The inverse quantization unit (220) can perform inverse quantization based on the quantization parameters and the coefficient values of the rearranged block.
[0068] The inverse transform unit (225) can perform inverse transform on the inverse quantized transform coefficients using a predetermined transform method. At this time, the transform method can be determined based on information regarding the prediction method (inter / intra prediction), block size / shape, intra prediction mode, etc.
[0069] The prediction unit (230, 235) can generate a prediction block based on the prediction block generation related information provided by the entropy decoding unit (210) and the previously decoded block or picture information provided by the memory (245).
[0070] The prediction unit (230, 235) may include a prediction unit determination unit, an inter prediction unit, and an intra prediction unit. The prediction unit determination unit may receive various information such as prediction unit information input from the entropy decoding unit (210), prediction mode information of an intra prediction method, and motion prediction-related information of an inter prediction method, and may distinguish a prediction unit in a current coding unit (CU) and determine whether the prediction unit performs inter prediction or intra prediction. The inter prediction unit (230) may perform inter prediction on the current prediction unit based on information included in at least one of a previous picture or a subsequent picture of the current picture including the current prediction unit, by using information necessary for inter prediction of the current prediction unit provided from the encoding device. Alternatively, inter prediction may be performed based on information on a pre-restored portion of the current picture including the current prediction unit.
[0071] In order to perform inter prediction, it is possible to determine whether the motion prediction method of the prediction unit included in the encoding unit is Skip Mode, Merge Mode, or AMVP Mode based on the encoding unit.
[0072] The intra prediction unit (235) can generate a prediction block based on pixel information in the current picture. If the prediction unit is a prediction unit that has performed intra prediction, intra prediction can be performed based on intra prediction mode information of the prediction unit provided by the encoding device. The intra prediction unit (235) can include an AIS (Adaptive Intra Smoothing) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is a part that performs filtering on the reference pixels of the current block, and can determine and apply whether to apply the filter according to the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixels of the current block using the prediction mode and AIS filter information of the prediction unit provided by the encoding device. If the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.
[0073] The reference pixel interpolation unit can generate a reference pixel of a pixel unit less than an integer value by interpolating the reference pixel when the prediction mode of the prediction unit is a prediction unit that performs intra prediction based on the pixel value interpolated from the reference pixel. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, the reference pixel may not be interpolated. The DC filter can generate a prediction block through filtering when the prediction mode of the current block is the DC mode.
[0074] The restored block or picture may be provided to a filter unit (240). The filter unit (240) may include a deblocking filter, an offset correction unit, and an ALF.
[0075] Information regarding whether a deblocking filter has been applied to a block or picture from an encoding device and, if so, whether a strong or weak filter has been applied can be provided. The deblocking filter of a decoding device can receive information regarding the deblocking filter provided by the encoding device, and the decoding device can perform deblocking filtering on the block.
[0076] The offset correction unit can perform offset correction on the restored image based on the type of offset correction applied to the image during encoding and offset value information.
[0077] ALF can be applied to a coding unit based on information such as whether ALF is applied and ALF coefficient information provided from the encoder. This ALF information can be provided by being included in a specific parameter set.
[0078] The memory (245) can store a restored picture or block so that it can be used as a reference picture or reference block, and can also provide the restored picture to an output unit.
[0079]
[0080] FIG. 3 illustrates a flowchart of an encoder according to one embodiment of the present invention.
[0081] The order of FIG. 3 refers to the order in which all of the various embodiments specified above are applied, and when only some embodiments are reflected in the method of adding a candidate list when determining a predicted motion vector candidate, steps related to unused methods may be omitted. For example, when offset correction is not performed, the offset may be set to 0 and the step of calculating the corresponding value may be omitted. The details of the predicted motion vector candidates are described below. In addition, the order of FIG. 3 illustrates adding candidates to the predicted motion vector candidate list in the order of a motion vector obtained by checking the pre-coding mode of an adjacent block, a motion vector obtained through template matching, and a motion vector obtained by applying an offset to the candidates.
[0082] When it is determined that the current encoding block is encoded in a block copy mode within a screen, the number of candidates and candidate targets of a candidate list of a prediction motion vector of the current encoding block can be determined in order to determine candidates for a prediction motion vector of the current encoding block. The number of candidates in the list can be a natural number equal to or greater than 1, 2, 3, 4, 5, 6, 7, or 8. The number of candidates in the list can be determined as a fixed value pre-agreed in the encoder / decoder, or can be adaptively determined based on the availability of adjacent blocks, position information of the current encoding block, size information of the current encoding block, etc. When the number of candidates in the list is adaptively determined, information indicating the number of candidates in the list can be encoded in the encoder and transmitted to the decoder through a bitstream.
[0083] The above predicted motion vector candidate list may include at least one of motion vectors of adjacent blocks encoded in the intra-screen block copy mode, motion vectors of blocks pre-encoded in the intra-screen block copy mode, or motion vectors generated through template matching as candidates. Here, the motion vectors of adjacent blocks encoded in the intra-screen block copy mode may include at least one of motion vectors of left blocks encoded in the intra-screen block copy mode or motion vectors of upper blocks encoded in the intra-screen block copy mode.
[0084] Additionally, the above predicted motion vector candidate list may further include as a candidate a motion vector that has been corrected by applying an offset to a motion vector already added to the above predicted motion vector candidate list.
[0085] <Motion vectors of adjacent blocks encoded in block copy mode within the screen>
[0086] If adjacent blocks at a specific position of the current encoding block are encoded in a block copy mode within the screen and the motion vectors of the corresponding blocks are available, the motion vectors of the corresponding blocks can be added to a predicted motion vector candidate list. For example, at least one of the motion vector of the left adjacent block or the motion vector of the upper adjacent block can be added to the predicted motion vector candidate list in order of priority. Here, the left adjacent block may be the uppermost block or the lowermost block among the left adjacent blocks when there are multiple blocks adjacent to the left of the current block. Here, the upper adjacent block may be the rightmost block or the leftmost block among the upper adjacent blocks when there are multiple blocks adjacent to the top of the current block. In addition, the number of motion vectors or the adjacent direction of the motion vectors of the adjacent blocks at a specific position of the current encoding block to be added to the predicted motion vector candidate list may be determined according to the number of predicted motion vector candidate lists. For example, if the number of predicted motion vector candidate lists is 4 or more, both the motion vector of the left adjacent block and the motion vector of the right adjacent block can be added to the predicted motion vector candidate list. For example, if the number of predicted motion vector candidate lists is less than 4, only one of the motion vector of the left adjacent block and the motion vector of the right adjacent block can be added to the predicted motion vector candidate list.
[0087] <Motion vector of a block encoded in block copy mode within the screen>
[0088] If there is a block encoded by intra-picture block copy before the current encoding block and the motion vector of the block is available, the motion vector of the block can be added as a candidate to the predicted motion vector candidate list. At this time, blocks adjacent to the current encoding block can be excluded from the encoded blocks. For example, blocks adjacent to the left and top of the current encoding block can be excluded. If there are multiple blocks encoded by intra-picture block copy mode in a pre-decoded picture, the multiple blocks are sorted based on the distance relationship between the corresponding block and the current encoding block, the encoding order, the number of occurrences of the same value of the predicted vector, etc., and a candidate to be added to the predicted motion vector candidate list can be determined based on rules such as the block closest to the encoding target block, the block with the most recent encoding order, and the value with the largest number of occurrences of the predicted vector.
[0089] Additionally, candidates added based on the above rules can be restricted to those belonging to the same CTU, block group, or slice as the current coded block. That is, in addition to the above rules, the above restrictions can be further considered to determine candidates to be added to the predicted motion vector candidate list.
[0090] <Motion vector generated through template matching>
[0091] If the base / decoding region adjacent to the current encoding block can be used as a template, a region to be used as a template among the base / decoding regions is determined to generate a candidate to be added to the list of predicted motion vector candidates of the current encoding block, and template matching can be performed by determining template matching weights and template matching target regions according to the positions of pixels within the region.
[0092] Figure 4 illustrates a process of obtaining a template from surrounding pixels of a current encoding block for template matching in a block copy within a screen.
[0093] The region to be used as a template includes at least one pixel and may be part of the peripheral pre-decoding region of the current encoding block, as shown in FIG. 4. The shape, size, and number of regions to be used as templates may vary depending on the embodiment.
[0094] For example, in block-based encoding, when the encoding order progresses from the upper left to the lower right, the area to be used as a template may be a part of the N available pixel lines of the upper, left, and upper left sub-encoding areas based on the current encoding block. Here, N is a natural number such as 1, 2, 3, 4, or 5, and may have the same value for the upper, left, and upper left of the current encoding block. Alternatively, N may be defined for each of the upper, left, and upper left, and may have different values. For example, the left side may be an area where one pixel line is used as a template, and the upper side may be an area where three pixel lines are used as templates. Alternatively, the upper side may be an area where one pixel line is used as a template, and the left side may be an area where three pixel lines are used as templates.
[0095] Depending on the embodiment, template matching may be performed using all or part of the pixels in the area to be used as a template. For example, only some of the pixels forming a pixel line within the area to be used as a template may be used as pixels for template matching.
[0096] The template matching of the present disclosure can be performed by assigning different weights to different pixel locations. The weights can be determined differently depending on the position and distance relationship between the current encoding block and the corresponding pixel. This can be performed identically in the encoder / decoder. Alternatively, information indicating the optimal weights calculated by the encoder can be transmitted to the decoder, and the decoder can determine the optimal weights based on this information without the need for the same calculations as the encoder.
[0097] Figure 5 illustrates an embodiment in which motion prediction is performed using a surrounding pre-decoding block of the current encoding block as a template.
[0098] Weights can be determined differently for each partition unit of the pre-decoded block to which a pixel belongs. If weights are determined differently for each partition unit of the pre-decoded block to which a pixel belongs, the calculated motion vectors may differ depending on the number of cases in which the weights are configured.
[0099] FIG. 5 can be an example of motion vector candidates for four cases where the weight of one of the regions (a), (b), (c), and (d) constituting the template is set to 100 and the weight of the other region is set to 0. Conversely, if there are multiple regions constituting the template with non-zero weights, the final motion vector candidate can be generated based on the value obtained by applying the weight to each motion vector. For example, if the weights for each of the regions (a), (b), (c), and (d) are the same at 25, the final motion vector candidate can be determined as 1 / 4 * mv_(a) + 1 / 4 * mv_(b) + 1 / 4 * mv_(c) + 1 / 4 * mv_(d). Here, mv_(x) can be a motion vector for the x region.
[0100] The template matching target region may be a part or all of the region excluding the region used as a template among the base / decoded regions in the same frame to which the current encoding block belongs. Here, the template matching target region may be a region matched based on a motion vector obtained based on a value obtained by applying a weight to each motion vector of a pixel.
[0101] In some embodiments, the template matching target region may be determined using all or part of the integer and decimal positions of pixels. When using decimal positions, interpolation filtering may be performed to generate the decimal positions. If interpolation filtering is used in a fixed manner by a pre-arrangement between the encoder and decoder, the information may not be transmitted from the encoder to the decoder.
[0102] In embodiments where interpolation filtering is used adaptively, information about the interpolation filtering can be transmitted from the encoder to the decoder via the bitstream.
[0103] According to the above method, when a motion vector candidate is obtained through template matching, it can be added to a list of predicted motion vector candidates for block copying within the screen of the current block. Here, the motion vector obtained through template matching can be any one of a motion vector pointing to a block to be matched with a template from a template, a motion vector of the block to be matched with a template, and a motion vector of a block adjacent to the block to be matched with a template (corresponding to a position of the current encoding block adjacent to the template).
[0104] In order to improve the accuracy of the predicted motion vector of the current encoding block, a compensation offset can be applied to the predicted motion vector obtained through the predicted motion vector candidate list. When the compensation offset is applied, the encoder can calculate the difference between both the predicted motion vector and the compensation offset from the motion vector of the current encoding block as a motion vector difference (MVD, or differential motion vector) and transmit it to the decoder. In other words, the motion vector of the current encoding block can be expressed as 'predicted motion vector + compensation offset + motion vector difference'.
[0105] In some embodiments, if the motion vector is determined only by the predicted motion vector and the correction offset without transmitting the motion vector difference value, the encoder may not transmit the differential motion vector to the decoder.
[0106] If the current coded block is an intra-picture prediction copy block, an adjacent block of the current coded block is coded through inter-picture prediction, there is a block coded through an available intra-picture block copy before the current coded block, and an adjacent block of the block coded through intra-picture prediction is coded through inter-picture prediction, an offset can be calculated through a relationship between motion vectors of inter-picture prediction blocks of the current coded block and the adjacent blocks of the current coded block and the motion vectors of adjacent blocks of the decoded intra-picture block copy block and the decoded intra-picture block copy block. If the condition is not satisfied, the offset can be a specific value including 0.
[0107]
[0108] Figure 6 illustrates a decoder flowchart according to one embodiment of the present invention.
[0109] FIG. 6 also illustrates the order of decoding by constructing a list of predicted motion vector candidates, similar to the flowchart of the encoder. The list of predicted motion vector candidates may be constructed in the following order: a motion vector determined by checking the motion vectors of surrounding blocks, a motion vector generated by template matching, and a motion vector obtained by adding an offset to the candidate. In addition, each process of FIG. 6 may be omitted.
[0110] In one proposed embodiment, if the current decoding block is encoded in a block-in-picture copy mode, motion vector information can be determined for motion prediction of the block. For example, a list of predicted motion vector candidates can be determined to determine the predicted motion vector. The list of predicted motion vector candidates can be processed in the same manner as the encoder. Once the list of predicted motion vector candidates is determined, the related information of the predicted motion vector can be decoded to determine the predicted motion vector (PMV).
[0111] At this time, if the predicted motion vector is determined as a candidate obtained through template matching, the decoder can also generate a motion vector through template matching in the same process as the encoder. In addition, in the embodiment, if a correction offset is applied, the offset can be calculated in the same way as in the encoder. Once the predicted motion vector and the correction offset are determined through the above process, the motion vector can be calculated by adding the differential motion vector transmitted from the encoder, and a reference block can be generated.
[0112] In some embodiments, when the differential motion vector is not transmitted, the motion vector may be determined using only the predicted motion vector and the correction offset.
[0113]
[0114] Figure 7(a) illustrates a method for generating a prediction block in a block copy mode within a screen. Figure 7(b) illustrates the location of a pre-decoded block used for motion vector prediction.
[0115] In prediction using the block-in-picture copy mode, the encoder can perform prediction according to a set rule in the base / decoded area of the same frame as the current coding block, and transmit the motion vector for the corresponding position to the decoder. At this time, if the block above or to the left of the current coding block (C) is predicted by the block-in-picture copy mode and the motion vector is available, the motion vector can be used as a predicted motion vector and a differential motion vector can be transmitted to the decoder. When decoding the current coding block (C), the decoder can calculate the motion vector of C by adding the transmitted differential motion vector using the vector of A or B as the predicted motion vector, and then generate the reference block (R) and perform decoding. If neither A nor B of the current coding block (C) are available, the encoder can use the motion vector of the block most recently encoded by the block-in-picture copy mode as the predicted motion vector, or use a specific motion vector determined by an agreement between the base / decoder as the predicted motion vector, depending on the embodiment.
[0116]
[0117] Referring to FIG. 4, the encoder can use the motion vector of the block including the blocks at positions A, B, 1, 2, 3, 4, and 5 as the predicted motion vector of the current encoding block (C). At this time, if the blocks to which A, B, 1, 2, 3, 4, and 5 belong are not predicted in the block copy mode within the screen and thus the motion vector for the block copy mode within the screen is not available, the block to which the corresponding position belongs can be set as a template as shown in FIG. 4, and the template can be searched in the pre-decoded area to generate a motion vector.
[0118] Fig. 5 illustrates an embodiment in which motion prediction is performed using the surrounding pre-decoded blocks of the current encoding block (C) as templates. If the predicted motion vector candidates belong to the same decoded block, template matching can be performed only once. In the embodiments of Figs. 4 and 5, there are seven predicted motion vector candidates, but if positions belonging to the same block are removed, template matching is performed only four times in total, and the motion vectors generated through template matching can be used as candidates for the predicted motion vector.
[0119]
[0120] Figure 8 illustrates an example of performing template matching based on a template.
[0121] In some embodiments, when performing template matching, template matching may be performed using a weighted SAD determined based on the number of blocks constituting the template. For example, if the template of the current encoding block C of FIG. 8(a) is configured as in FIG. 8(b), template matching may be performed with differential weights for each of blocks (A), (B), (C), and (D). The number of motion vectors obtained through template matching may be determined as many times as the number of cases in which the weights are differential, and the motion vectors obtained in this way may be used as candidates for predicted motion vectors.
[0122]
[0123] Figure 9 illustrates an example of obtaining a predicted motion vector candidate using a base-based / decoded surrounding block.
[0124] In some embodiments, if a block currently being encoded / decoded is encoded in an intra-block copy mode and a surrounding pre-pred / decoded block is encoded / decoded by inter-screen prediction, the motion vector of the surrounding block may be used as a predicted motion vector. In another embodiment, if there is an intra-block copy block (I) that is pre-pred / decoded in the same frame before the current encoded block is encoded / decoded by intra-block copy, and a surrounding block of the block is encoded / decoded by inter-screen prediction, the relationship between the motion vector of I and the motion vector of the surrounding block may be used in block C to predict the motion vector.
[0125] In the method, when calculating the difference between the motion vector for the adjacent block of the I block and the motion vector for the adjacent block of the C block according to an embodiment, a correction offset can be calculated through direct difference only when the two blocks are the same reference frame. If they are not the same reference frame, the correction offset can be determined as 0. Alternatively, if the motion vector for the adjacent block of the I block and the motion vector for the adjacent block of the C block refer to different reference frames, the correction offset can be calculated by projecting them to a specific frame and using the difference between the projected motion vectors.
[0126] While the exemplary methods of this disclosure are presented as a series of operations for clarity of description, this is not intended to limit the order in which the steps are performed, and individual steps may be performed simultaneously or in different orders, if desired. To implement a method according to this disclosure, additional steps may be included in addition to the steps illustrated, some steps may be excluded and the remaining steps included, or some steps may be excluded and additional steps included.
[0127] The various embodiments of the present disclosure are not intended to list all possible combinations but rather to illustrate representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combinations of two or more.
[0128] Additionally, various embodiments of the present disclosure may be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, the embodiments may be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc.
[0129] The scope of the present disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of various embodiments to be executed on a device or a computer, and a non-transitory computer-readable medium having such software or instructions stored thereon and executable on the device or computer.
[0130] The present invention can be utilized in the field of encoding / decoding video signals.
Claims
1. A step for determining whether to encode the current encoding block in a screen block copy mode; In response to determining that the current encoding block is encoded in the within-screen block copy mode, a step of constructing a candidate list of predicted motion vectors of the current encoding block; A step of obtaining a motion vector of the current encoding block based on a candidate of the above predicted motion vector candidate list; and An image encoding method, comprising a step of encoding the current encoding block based on the motion vector.
2. In paragraph 1, A video encoding method, wherein information indicating the number of candidates in the above predicted motion vector candidate list is encoded in a bitstream.
3. In paragraph 1, A video encoding method, wherein the above predicted motion vector candidate list includes at least one of motion vectors of adjacent blocks encoded in a block copy mode within a screen, motion vectors of blocks pre-encoded in a block copy mode within a screen, or motion vectors generated through template matching.
4. In paragraph 3, A method for encoding an image, wherein the adjacent block is at least one of a left adjacent block adjacent to the left of the current encoding block or an upper adjacent block adjacent to the top of the current encoding block.
5. In paragraph 4, A method for encoding an image, wherein the left adjacent block is, when there are multiple blocks adjacent to the left of the current encoded block, the lowermost block among the multiple blocks.
6. In paragraph 4, A method for encoding an image, wherein the upper adjacent block is, when there are multiple blocks adjacent to the upper side of the current encoding block, the rightmost block among the multiple blocks.
7. In paragraph 3, A video encoding method, wherein the motion vector of a block encoded in the above-described block copy mode within a screen is determined based on a distance relationship with the current encoded block or an encoding order among a plurality of blocks when there are a plurality of blocks encoded in the above-described block copy mode within a screen.
8. In paragraph 3, An image encoding method, wherein the above template matching is performed using a template determined based on surrounding pixels of the current encoding block.
9. In paragraph 8, An image encoding method, wherein the above template matching is performed by a weight determined according to the distance relationship between the surrounding pixels and the current encoding block.
10. In paragraph 3, A video encoding method, wherein the motion vector generated through the above template matching is obtained based on motion vectors obtained by applying weights to each template when the current encoding block has multiple templates.
11. In paragraph 3, A method for encoding an image, wherein the template matching target area of the above template matching is determined using the decimal point position of a pixel.
12. In paragraph 3, A video encoding method, wherein the above predicted motion vector candidate list further includes a motion vector to which a correction offset is applied to a candidate of the above predicted motion vector candidate list.
13. A step for determining whether to decrypt the current decryption block in the on-screen block copy mode; In response to determining that the current decoding block is to be decoded in the within-screen block copy mode, a step of constructing a candidate list of predicted motion vectors of the current decoding block; A step of obtaining a motion vector of the current decoding block based on a candidate of the above predicted motion vector candidate list; and An image decoding method, comprising a step of decoding the current decoding block based on the motion vector.
14. In a computer-readable recording medium storing a bitstream generated by a video encoding method, The above image encoding method comprises the steps of: determining whether to encode a current encoding block in a screen block copy mode; In response to determining that the current encoding block is encoded in the within-screen block copy mode, a step of constructing a candidate list of predicted motion vectors of the current encoding block; A step of obtaining a motion vector of the current encoding block based on a candidate of the above predicted motion vector candidate list; and A computer-readable recording medium comprising a step of encoding the current encoding block based on the motion vector.
Citation Information
Patent Citations
High pressure hose coupling connector
KR1020240030394A
Sweat absorption headband and its manufacturing method
KR1020240159995A
Multi-microplate tower assembly and dispensing system comprising the same
KR1020250061090A
Making equipment of on side no sewing bedclothes and manufacturing method of double side no sewing bedclothes
KR102404254B1
Unify intra block copy and inter prediction
KR102447297B1