Encoding method, decoding method, computer-readable storage medium, and transmission method
By constructing candidate lists with diverse block vectors, the method enhances coding efficiency for high-resolution images in intra-TMP and IBC modes, addressing inefficiencies in existing image compression technologies.
Patent Information
- Application Number
- PCT/KR2025/001643
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-05
- Filing Date
- 2025-02-04
- Publication Date
- 2025-08-14
AI Technical Summary
Existing image compression technologies struggle to efficiently encode and decode high-resolution, high-quality images, particularly in modes like intra-TMP and IBC, limiting coding efficiency.
A method and device that construct a candidate list by utilizing various block vectors, allowing access to more diverse reference block locations, enhancing coding efficiency through techniques such as intra-TMP mode and IBC merge mode.
Improves coding efficiency by effectively utilizing block vectors, enabling better encoding and decoding of high-resolution images in intra-TMP and IBC modes.
Smart Images

Figure KR2025001643_14082025_PF_FP_ABST
Abstract
Description
Encoding method, decoding method, computer-readable storage medium and transmission method
[0001] The present disclosure relates to a method for encoding / decoding video information, a computer-readable storage medium for storing a bitstream, and a method for transmitting the bitstream.
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various application fields, and accordingly, high-efficiency image compression technologies are being discussed.
[0003] There are various technologies such as inter prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture, intra prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, and entropy coding technology that assigns short codes to values with high frequency of appearance and long codes to values with low frequency of appearance, and these video compression technologies can be used to effectively compress and transmit or store video data.
[0004] Accordingly, a highly efficient image compression technology is required to effectively transmit, store, and play high-resolution, high-quality image information.
[0005] The present disclosure provides a method and device for constructing a candidate list by utilizing various block vectors when encoding still images or moving images within a screen.
[0006] The present disclosure provides a method and device that can utilize such a method in intra-TMP mode, IBC merge mode, and IBC mode.
[0007] The present disclosure provides a method and device capable of accessing and utilizing reference block locations at more diverse locations by effectively utilizing block vectors according to such a method, thereby improving coding efficiency.
[0008] A decoding method according to one aspect comprises the steps of: obtaining information about a prediction mode from a bitstream; determining a prediction mode to be applied to a current block as a prediction mode using a block vector based on the information about the prediction mode; constructing a candidate list including at least one block vector candidate for the current block; and generating a prediction block based on at least one block vector candidate included in the candidate list; wherein the step of constructing the candidate list comprises adding block vector information of a previously reconstructed neighboring block of the current block or block vector information derived from the block vector information of the previously reconstructed neighboring block to the candidate list, and wherein the previously reconstructed neighboring block of the current block is characterized in that it is reconstructed by a combination of at least one prediction tool or at least one prediction mode and at least one of an Intra Block Copy (IBC) mode or an Intra Template Matching Prediction (TMP) mode.
[0009] An encoding method according to one aspect includes the steps of: determining a prediction mode to be applied to a current block as a prediction mode using a block vector; constructing a candidate list including block vector candidates for the current block; and generating a prediction block based on at least one block vector candidate included in the candidate list; wherein the step of constructing the candidate list includes adding block vector information of a previously reconstructed neighboring block of the current block or block vector information derived from the block vector information of the previously reconstructed neighboring block to the candidate list, and wherein the previously reconstructed neighboring block of the current block is characterized in that it is reconstructed by a combination of at least one prediction tool or at least one prediction mode and at least one of an Intra Block Copy (IBC) mode or an Intra Template Matching Prediction (TMP) mode.
[0010] A computer-readable storage medium storing a bitstream generated by an encoding method according to one aspect, the encoding method comprising: a step of determining a prediction mode applied to a current block as a prediction mode using a block vector; a step of constructing a candidate list including block vector candidates for the current block; and a step of generating a prediction block based on at least one block vector candidate included in the candidate list; wherein the step of constructing the candidate list includes adding block vector information of a previously reconstructed neighboring block of the current block or block vector information derived from the block vector information of the previously reconstructed neighboring block to the candidate list, wherein the previously reconstructed neighboring block of the current block is characterized in that it is reconstructed by a combination of at least one prediction tool or at least one prediction mode and at least one of an Intra Block Copy (IBC) mode or an Intra Template Matching Prediction (TMP) mode.
[0011] A method for transmitting data for an image according to one aspect, comprising: obtaining a bitstream for the image, wherein the bitstream is generated based on a step of determining a prediction mode applied to a current block as a prediction mode using a block vector; a step of constructing a candidate list including block vector candidates for the current block; and a step of generating a prediction block based on at least one block vector candidate included in the candidate list; and a step of transmitting the data including the bitstream; wherein the step of constructing the candidate list includes adding block vector information of a previously reconstructed neighboring block of the current block or block vector information derived from the block vector information of the previously reconstructed neighboring block to the candidate list, and wherein the previously reconstructed neighboring block of the current block is characterized in that it is reconstructed by a combination of at least one prediction tool or at least one prediction mode and at least one of an Intra Block Copy (IBC) mode or an Intra Template Matching Prediction (TMP) mode.
[0012] According to the present disclosure, when encoding still images or moving images within a screen, a candidate list can be constructed by utilizing various block vectors.
[0013] According to the present disclosure, this method can be utilized in intra-TMP mode, IBC merge mode, IBC mode, etc.
[0014] According to the present disclosure, by effectively utilizing block vectors according to this method, reference block locations at more diverse locations can be accessed and utilized, thereby improving coding efficiency.
[0015] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.
[0016] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0017] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0018] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0019] FIG. 4 illustrates an example of a video / image decoding method to which an embodiment of the present disclosure can be applied.
[0020] FIG. 5 illustrates an example of a video / image encoding method to which an embodiment of the present disclosure can be applied.
[0021] Figure 6 is a diagram showing an example of a search area used in intra template matching.
[0022] Figure 7 is a diagram showing the use of intra TMP block vectors for IBC blocks.
[0023] Figure 8 is a diagram expressing information indicated by a block vector in IBC mode.
[0024] Figure 9 is a drawing showing a reference area applied to the IBC mode.
[0025] Figure 10 is a diagram showing an IBC reference area according to the current CU location.
[0026] Figure 11 shows examples of BV, BVD and BVP for the current PU.
[0027] Figure 12 is a schematic diagram showing the configuration method of HoG.
[0028] Figure 13 is a diagram showing the configuration of a prediction block in DIMD mode.
[0029] FIG. 14 is a diagram illustrating an example of a decoding method according to one embodiment of the present disclosure.
[0030] Figure 15 is a flowchart of an encoding method according to one embodiment.
[0031] FIG. 16 is a diagram illustrating an example of a decoding method according to one embodiment of the present disclosure.
[0032] Figure 17 is a drawing showing an example of a template shape used for template matching.
[0033] Fig. 18 is a drawing showing the location of a sub-pel applicable to the disclosed embodiment.
[0034] Figure 19 is a diagram showing a method for deriving filter coefficients of a filter model applied to intra TMP.
[0035] FIG. 20 is a diagram illustrating a process of serially deriving reference blocks of surrounding blocks in the disclosed embodiment.
[0036] Figure 21 is a diagram showing an example of the locations of surrounding blocks used to construct an intra TMP candidate list of the current block.
[0037] FIG. 22 is a drawing showing an example of an encoding method according to one embodiment of the present disclosure.
[0038] FIG. 23 is a diagram illustrating an example of a decoding method according to one embodiment of the present disclosure.
[0039] FIG. 24 is a drawing showing an example of an encoding method according to one embodiment of the present disclosure.
[0040] FIG. 25 is a diagram illustrating an example of a decoding method according to one embodiment of the present disclosure.
[0041] FIG. 26 is a drawing showing an example of an encoding method according to one embodiment of the present disclosure.
[0042] FIG. 27 is a diagram illustrating an example of a content streaming system to which an embodiment according to the present disclosure can be applied.
[0043] The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Throughout the description of each drawing, similar reference numerals have been used to designate similar components.
[0044] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0045] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0046] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0047] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the versatile video coding (VVC) standard. In addition, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVS2), or the next generation of video / image coding standards (e.g., H.267 or H.268).
[0048] This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0049] In this specification, video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs that has a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs that has a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged consecutively according to the CTU raster scan, while tiles within a picture may be arranged consecutively according to the tile raster scan. A slice may contain an integer number of complete tiles or an integer number of contiguous complete CTU rows within a picture, which may be exclusively contained within a single NAL unit. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture.
[0050] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.
[0051] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0052] In this specification, “A or B” can mean “only A,” “only B,” or “both A and B.” In other words, “A or B” in this specification can be interpreted as “A and / or B.” For example, “A, B or C” in this specification can mean “only A,” “only B,” “only C,” or “any combination of A, B, and C.”
[0053] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0054] In this specification, “at least one of A and B” may mean “only A,” “only B,” or “both A and B.” Additionally, in this specification, the expressions “at least one of A or B” or “at least one of A and / or B” may be interpreted identically to “at least one of A and B.”
[0055] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0056] Additionally, parentheses used herein may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction."
[0057] Technical features individually described in a single drawing in this specification may be implemented individually or simultaneously.
[0058] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0059] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device).
[0060] A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting device. The receiving device may include a receiving device, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.
[0061] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0062] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0063] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The storage medium can be a computer-readable storage medium. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0064] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0065] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0066] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0067] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to an embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0068] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units (PUs). For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-Tree Binary-Tree Ternary-Tree) structure.
[0069] For example, a single coding unit may be split into multiple coding units with deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quad-tree structure. The coding procedure according to the present specification may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0070] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be split or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0071] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component. A sample can be used as a term corresponding to a pixel or pel of a picture (or image).
[0072] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).
[0073] The prediction unit (220) can perform a prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit (220) can generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information related to prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0074] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples, i.e., the reference samples, may be located in the neighborhood of the current block or may be located a certain distance away from the current block depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of the DC mode or the planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail in the prediction direction. However, this is merely an example, and a greater or lesser number of directional modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0075] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring block can include a spatial neighboring block existing in the current picture and a temporal neighboring block existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. Temporal neighboring blocks may be called collocated reference blocks, collocated CUs (colCUs), etc., and reference pictures including temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit (221) may construct a motion information candidate list based on neighboring blocks, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of neighboring blocks as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0076] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction to predict a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called a combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can perform an intra block copy (IBC) prediction mode to predict a block. The IBC prediction mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described herein. The prediction signal generated through the prediction unit (220) can be used to generate a restored signal or a residual signal.
[0077] The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously restored pixels. In addition, the transform process can be applied to a pixel block having a square size and the same size, or can be applied to a block of a non-square variable size.
[0078] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0079] The entropy encoding unit (240) can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) can also encode information necessary for video / image restoration (e.g., values of syntax elements, etc.) together or separately from quantized transform coefficients.
[0080] Encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, or SSD. For example, the storage medium may be a medium that stores the bitstream non-statutory.
[0081] A transmission unit (not shown) for transmitting a signal output from an entropy encoding unit (240) and / or a storage unit (not shown) for storing the signal may be configured as an internal / external element of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).
[0082] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0083] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit the information to the entropy encoding unit (240). The information regarding filtering may be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0084] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency.
[0085] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in a picture that has already been restored. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (270) can store restored samples of restored blocks in the current picture and transfer them to the intra prediction unit (222).
[0086] Image information output in the form of a bitstream from the encoding device (200) can be transmitted to the decoding device (300) through the transmission unit.
[0087] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0088] Image information transmitted in the form of a bitstream from the encoding device (200) can be received by the decoding device (300).
[0089] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (332) and an intra-prediction unit (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).
[0090] The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0091] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit of decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output by the decoding device (300) can be reproduced through a reproduction device.
[0092] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded and obtained from the bitstream through the decoding procedure. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of a decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310).
[0093] Meanwhile, a decoding device according to the present specification may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331).
[0094] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0095] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0096] The prediction unit (320) can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit (320) can determine whether intra-prediction or inter-prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter-prediction mode.
[0097] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be signaled and included in the video / image information.
[0098] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit (331) may also determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0099] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the neighboring blocks, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block.
[0100] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the restoration block.
[0101] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0102] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0103] The (modified) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of blocks in an already reconstructed picture. The stored motion information can be transmitted to the inter prediction unit (332) to be used as motion information of a spatial neighboring block or motion information of a temporal neighboring block. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra prediction unit (331).
[0104] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.
[0105] FIG. 4 illustrates an example of a video / image decoding method to which an embodiment of the present disclosure can be applied.
[0106] In image / video coding, the pictures that make up an image / video can be encoded / decoded according to a series of decoding orders. The picture order corresponding to the output order of the decoded pictures can be set differently from the decoding order, and based on this, not only forward prediction but also backward prediction can be performed during inter prediction.
[0107] In FIG. 4, S400 may be performed in the entropy decoding unit (310) of the aforementioned decoding device (300), S410 may be performed in the prediction unit (330), S420 may be performed in the residual processing unit (320), S430 may be performed in the addition unit (340), and S440 may be performed in the filtering unit (350). S400 may include a decoding procedure according to the present disclosure, S410 may include an inter / intra prediction procedure according to the present disclosure, S420 may include a residual processing procedure according to the present disclosure, S430 may include a block / picture restoration procedure according to the present disclosure, and S440 may include an in-loop filtering procedure according to the present disclosure.
[0108] Referring to FIG. 4, the decoding device obtains image / video information from a bitstream (S400), performs prediction based on the obtained image / video information (S410), and restores a picture through residual processing (S420, inverse quantization for quantized transform coefficients, inverse transformation) (S430).
[0109] A modified restored picture can be generated by applying an in-loop filtering procedure (S440) to a restored picture generated through the above restoration procedure, and the modified restored picture can be output as a decoded picture and can be stored in a buffer or memory of a decoding device to be used as a reference picture in an inter prediction procedure when decoding a next picture. In some cases, the in-loop filtering procedure can be omitted, in which case the restored picture can be output as a decoded picture and can be stored in a buffer or memory of a decoding device to be used as a reference picture in an inter prediction procedure when decoding a next picture.
[0110] The in-loop filtering procedure (S440) may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, an adaptive loop filter (ALF) procedure, and / or a bi-lateral filter procedure, and some or all of them may be omitted. In addition, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bi-lateral filter procedure may be sequentially applied, or all of them may be sequentially applied. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture. Or, for example, the ALF procedure may be performed after the deblocking filtering procedure is applied to the restored picture. This may also be performed in an encoding device.
[0111] FIG. 5 illustrates an example of a video / image encoding method to which an embodiment of the present disclosure can be applied.
[0112] In FIG. 5, the prediction step (S500) may be performed in the prediction unit (220) of the encoding device (200) described above, residual processing (S510) based on the prediction result may be performed in the residual processing unit (230), and the step (S520) of encoding image information including prediction information and residual information may be performed in the entropy encoding unit (240). S500 may include an inter / intra prediction procedure according to the present disclosure, S510 may include a residual processing procedure according to the present disclosure, and S520 may include an encoding procedure according to the present disclosure.
[0113] The encoding procedure may optionally include a procedure for encoding information for picture restoration (e.g., prediction information, residual information, partitioning information, etc.) and outputting it in the form of a bitstream, as well as a procedure for generating a restored picture for the current picture and a procedure for applying in-loop filtering to the restored picture.
[0114] The encoding device (200) can derive (corrected) residual samples from the quantized transform coefficients through the inverse quantization unit (234) and the inverse transformation unit (235), and can generate a restored picture based on the prediction samples and (corrected) residual samples, which are outputs of S500. The restored picture generated in this way can be the same as the restored picture generated by the decoding device (300) described above. A modified restored picture can be generated through an in-loop filtering procedure for the restored picture, which can be stored in a buffer or memory, and, as in the case of the decoding device, can be used as a reference picture in the inter prediction procedure when encoding a subsequent picture.
[0115] As described above, some or all of the in-loop filtering procedure may be omitted in some cases. When the in-loop filtering procedure is performed, (in-loop) filtering-related information (parameters) may be encoded by the entropy encoding unit (240) and output in the form of a bitstream, and the decoding device (300) may perform the in-loop filtering procedure in the same manner as the encoding device based on the filtering-related information.
[0116] Through this in-loop filtering procedure, noise occurring during image / video coding, such as blocking artifacts and ringing artifacts, can be reduced, and subjective / objective image quality can be improved. In addition, by performing the in-loop filtering procedure in both the encoding device (200) and the decoding device (300), the same prediction results can be derived from the encoding device (200) and the decoding device (300), thereby increasing the reliability of picture coding and reducing the amount of data that must be transmitted for picture coding.
[0117] As described above, the picture restoration procedure can be performed not only in the decoding device (300) but also in the encoding device (200). A restoration block can be generated based on intra-prediction / inter-prediction for each block, and a restoration picture including the restoration blocks can be generated. If the current picture / slice / tile group is an I picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based only on intra-prediction. On the other hand, if the current picture / slice / tile group is a P or B picture / slice / tile group, the blocks included in the current picture / slice / tile group can be restored based on intra-prediction or inter-prediction. In this case, inter-prediction may be applied to some blocks in the current picture / slice / tile group, and intra-prediction may be applied to some remaining blocks.
[0118] The color component of a picture may include a luma component and a chroma component, and embodiments according to the present disclosure may be applied to the luma component and the chroma component unless explicitly limited in the present disclosure.
[0119] Meanwhile, the prediction unit (220, 330) of the encoding device (200) / decoding device (300) can derive a reference sample according to the intra prediction mode of the current block among the surrounding samples of the current block, and can generate a prediction sample of the current block based on the reference sample.
[0120] For example, (i) the prediction sample can be derived based on the average or interpolation of neighboring reference samples of the current block, and (ii) the prediction sample can be derived based on a reference sample existing in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. The case of (i) can be called a non-directional mode or a non-angular mode, and the case of (ii) can be called a directional mode or an angular mode.
[0121] Additionally, linear interpolation intra prediction (LIP) may be applied to perform intra prediction on the current block by linearly interpolating prediction sample values generated based on the intra prediction mode of the current block.
[0122] Additionally, a temporary prediction sample of the current block may be derived based on filtered peripheral reference samples, and a prediction sample of the current block may be derived by weighting at least one reference sample derived according to an intra prediction mode among existing peripheral reference samples, i.e., unfiltered peripheral reference samples, and the temporary prediction sample. Such prediction may be referred to as Position Dependent Intra Prediction Combination (PDPC).
[0123] In addition, intra prediction encoding can be performed by selecting a reference sample line with the highest prediction accuracy among the surrounding multiple reference sample lines of the current block, deriving a prediction sample using the reference sample located in the prediction direction of the selected line, and then instructing (signaling) the used reference sample line to the decoding device. This case can be referred to as multi-reference line intra prediction (MRL) or MRL-based intra prediction.
[0124] Additionally, the current block can be divided into vertical or horizontal subpartitions, and intra prediction can be performed based on the same intra prediction mode, while peripheral reference samples can be derived and utilized for each subpartition. In other words, in this case, the intra prediction mode for the current block is applied equally to the subpartitions, but peripheral reference samples can be derived and utilized for each subpartition, thereby improving intra prediction performance in some cases. This prediction method can be called intra subpartitions (ISP) or ISP-based intra prediction.
[0125] In addition, if the prediction direction based on the prediction sample points between surrounding reference samples, that is, if the prediction direction points to a fractional sample location, the value of the prediction sample can also be derived through interpolation of multiple reference samples located around the prediction direction (around the fractional sample location).
[0126] Information regarding intra prediction modes may be included in prediction information encoded by an encoding device and transmitted to a decoding device in a bitstream. Information regarding intra prediction modes may be implemented and transmitted in various forms, such as flag information indicating whether each intra prediction mode is applied or index information indicating one of several intra prediction modes.
[0127] The MPM list for deriving the intra prediction mode described above may be configured differently depending on the intra prediction mode. Alternatively, the MPM list may be configured in common regardless of the intra prediction mode.
[0128] Intra TMP (Intra Template Matching Prediction)
[0129] Intra-template matching prediction (IntraTMP) is a special intra-prediction mode that copies the optimal prediction block where the L-shaped template matches the current template within the reconstructed portion of the current frame. Within a predefined search range, the encoder searches the reconstructed region of the current frame for the template most similar to the current template and uses that block as the prediction block. The encoder then signals the use of this mode, and the decoder performs the same prediction operation.
[0130] Figure 6 is a diagram showing an example of a search area used in intra template matching.
[0131] The prediction signal is generated by matching the L-shaped, top-only, or left-only causal neighbors of the current block with other blocks within the predefined search regions of Fig. 6. As illustrated in Fig. 6, there can be a total of six predefined search regions (i.e., R1 to R6), which include not only some of the reconstructed samples within the current CTU located above, left, bottom-left, and top-right of the current block, but also samples reconstructed from the upper CTU and the left CTU:
[0132] The sum of absolute differences (SAD) is used as a cost function.
[0133] A given search order is utilized across six search regions (i.e., R4, R5, R6, R1, R2, R3). Within each region, the decoder generates a list of up to 19 template-matching block vector candidates, sorted in ascending order by template cost (SAD). The supported modes are:
[0134] 1. Single Predictor: A single predictor is selected from the candidate list.
[0135] 2. Fusion of Multiple Predictors: Multiple predictors are fused to derive the final prediction block. The fusion weights can be calculated based on the template matching cost of each predictor, or a weight derivation method based on a Wiener filter can be used.
[0136] 3. Sub-pixel Precision: When using a single predictor, it supports 1 / 2 pixel, 1 / 4 pixel, and 3 / 4 pixel precision, and provides 8 directions each.
[0137] 4. Linear Filter Model: Applies a linear filter learned between the reference template and the current template to the reference block. This mode can be applied to a single predictor that does not use subpixel precision.
[0138] To ensure a fixed number of SAD comparisons per pixel, the sizes of all search regions (SearchRange_w, SearchRange_h) are set to be proportional to the block sizes (BlkW, BlkH). That is:
[0139] SearchRange_w = min(64,a * BlkW)
[0140] SearchRange_h = min(64,a * BlkH)
[0141] Here, 'a' is a constant that controls the trade-off between gain and complexity, and can be set to a = 5.
[0142] To accelerate the template matching process, the search range of each search area can be subsampled by a factor of three. After finding the optimal match, a refinement process is performed. This refinement is achieved through a second template matching search around the optimal match in the reduced range.
[0143] Intra-template matching can be enabled in CUs with a width and height of 64 or less. The maximum CU size for intra-template matching is configurable.
[0144] Intra template matching prediction mode can be signaled at the CU level via a dedicated flag when DIMD is not used for the current CU.
[0145] Intra TMP derived block vector candidates for IBC
[0146] In this method, block vectors (BVs) derived from intra-template matching prediction (Intra-TMP) are used for intra-block copying (IBC). The intra-TMP block vectors of stored neighboring blocks are used as spatial block vector candidates in constructing the IBC candidate list, along with the IBC block vectors.
[0147] Figure 7 is a diagram showing the use of intra TMP block vectors for IBC blocks.
[0148] The intra TMP block vector is stored in the IBC block vector buffer, and as illustrated in FIG. 7, the current IBC block can use both the IBC block vector and the intra TMP block vector of the surrounding blocks as block vector candidates in the IBC block vector candidate list.
[0149] Intra TMP block vectors can be added to the IBC block vector candidate list as spatial candidates.
[0150] Intra Block Copy (IBC)
[0151] In IBC mode, block matching is performed between the current block and the reference block to find the optimal block vector. Furthermore, IBC mode can manage multiple block vector candidates in a candidate list. One of the block vector candidates included in the candidate list can be signaled, and block vector information can be signaled.
[0152] Figure 8 is a diagram expressing information indicated by a block vector in IBC mode.
[0153] Referring to FIG. 8, a block vector is used to indicate a displacement from the current block to a reference block already restored within the current picture.
[0154] At the CU level, IBC mode can be signaled with a flag, and can be signaled as IBC AMVP mode or IBC skip / merge mode as follows:
[0155] - IBC Skip / Merge Mode: Indicates the block vector used to predict the current block among the block vectors in the adjacent candidate IBC coding block list using the merge candidate index.
[0156] - IBC AMVP mode: Block vector differences are coded. A flag indicating the block vector prediction index is signaled.
[0157] An IBC Merge / AMVP candidate can only be inserted into the IBC Merge / AMVP candidate list if it is valid.
[0158] The top-right, bottom-left, and top-left spatial candidates (belonging to the adjacent spatial candidate category) and a pair of average candidates can be added to the IBC merge / AMVP candidate list.
[0159] Template-based adaptive reordering (ARMC-TM) can be applied to IBC merge lists.
[0160] Candidates for non-adjacent spatial neighboring blocks (also known as non-adjacent candidates) can be added to the candidate lists for IBC merge mode and IBC AMVP. These non-adjacent candidates are inserted between adjacent spatial candidates and HBVP candidates for both IBC merge and IBC AMVP. In general inter mode, the same reference region for non-adjacent merge is reused for IBC.
[0161] IBC reference area
[0162] Figure 9 is a drawing showing an example of a reference area applied to the IBC mode.
[0163] The reference region for IBC extends to the upper two CTU rows. Referring to the example in Fig. 9, for a CTU (m,n) being coded, the reference region includes CTUs with indices (m-2, n-2)… (W, n-2),(0, n-1)… (W, n-1),(0, n)… (m, n), where W represents the maximum horizontal index within the current tile, slice, or picture.
[0164] When the CTU size is 256, the reference region can be limited to one upper CTU row. This setting can prevent IBC from requiring additional memory on current ETM platforms for CTU sizes of 128 or 256. The block vector search (also called local search) region per sample is limited to [-(C << 1), C >> 2] horizontally and [-C, C >> 2] vertically to accommodate the reference region expansion, where C represents the size of the CTU.
[0165] Filtered IBC prediction
[0166] A filtered IBC mode is additionally proposed, which applies a filter to the IBC predictor derived by minimizing the MSE between the current template and the reference template.
[0167] The output of the filter can be produced as follows:
[0168] predLumaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B
[0169] The nonlinear term P is expressed as a power of two central samples C, scaled to the range of sample values in the content:
[0170] P = ( C*C + midVal ) >> bitDepth
[0171] The bias term B represents a scalar offset between input and output, and is set to the middle luma value (512 for 10-bit content).
[0172] This filtering mode is used as an additional mode for non-merge IBC blocks and is not used with IBC-LIC, IBC-CIIP, or RR-IBC. For IBC merge mode, this filtering mode is inherited when the merge mode list is constructed. The mode flag is signaled before the IBC-LIC flag.
[0173] IBC with Template Matching
[0174] Template matching can be used for both IBC merge mode and IBC AMVP mode.
[0175] The IBC-TM merge list is modified compared to the one used in the regular IBC merge mode so that candidates are selected based on a pruning method with motion distances between candidates, as in the regular TM merge mode. The final zero-motion transition is replaced with motion vectors for the left (-W, 0), up (0, -H), and up-left (-W, -H), where W is the width of the current CU and H is its height.
[0176] In IBC-TM merge mode, selected candidates are refined using a template matching method before the RDO or decoding process. IBC-TM merge mode competes with the standard IBC merge mode, and the TM merge flag is signaled.
[0177] In IBC-TM AMVP mode, up to three candidates are selected from the IBC-TM merge list. Each of the three selected candidates is refined using a template matching method and sorted according to the resulting template matching cost. Then, as usual, only the first two are considered in the motion estimation process.
[0178] Figure 10 is a diagram showing an IBC reference area according to the current CU location.
[0179] Template matching improvement in IBC-TM merge and AMVP modes is very simple, as illustrated in Figure 10, because the IBC motion vectors are (i) integers and (ii) must be within the reference region. Therefore, in IBC-TM merge mode, all improvements are performed in integer precision, while in IBC-TM AMVP mode, they are performed in integer or 4-pel precision, depending on the AMVR value. This refinement accesses only samples without interpolation. In both cases, the refined motion vectors and templates used at each refinement step must adhere to the constraints of the reference region.
[0180] Block Vector Difference (BVD) prediction
[0181] The possible BVD code combinations in IBC mode are sorted by template matching cost. Additionally, the first four most frequent signification suffix bins of the exponential Golomb code used to indicate the BVD size are also sorted by TM cost. Figure 11 illustrates examples of BV, BVD, and BVP for the current PU.
[0182] A template matching operation is used to determine the BVD candidate with the best cost, and an indication is given in the bitstream whether the optimal candidate was predicted correctly.
[0183] DIMD(Decoder side Intra Mode Derivation)
[0184] The DIMD mode can be derived and used by the encoder and decoder without directly transmitting intra prediction mode information. First, horizontal and vertical gradients are obtained from the second neighboring sample column and row, and a Histogram of Gradients (HoG) can be constructed from these.
[0185] Fig. 12 is a drawing schematically showing a method of configuring HoG, and Fig. 13 is a drawing showing the configuration of a prediction block in DIMD mode.
[0186] Referring to Figure 12, the HoG can be obtained by applying a Sobel filter using L-shaped rows and columns of 3 pixels around the current block. If the block boundaries exist in different CTUs, they are not used for texture analysis.
[0187] Afterwards, as shown in Figure 13, the two intra modes with the largest histogram amplitudes are selected, and the predicted blocks predicted using these modes are blended with the planar mode to form the final predicted block. The weights can be derived from the histogram amplitude. Additionally, the DIMD flag can be transmitted on a block-by-block basis to determine whether DIMD is being used.
[0188] FUSION FOR TEMPLATE-BASED INTRA MODE DERIVATION (TIMD)
[0189] For each intra prediction in the MPM, the Sum of Absolute Transformed Differences (SATD) between the predicted sample and the reconstructed sample in the template is calculated. The first two intra prediction modes with the smallest SATD can be selected as the TIMD mode. These two TIMD modes are weighted and fused, and the weighted intra prediction is used to encode the current CU. PDPC can be included in the derivation of the TIMD mode.
[0190] The costs of the two selected modes are compared against a threshold and a cost factor of 2 is applied in the test as follows:
[0191] costMode2 < 2*costMode1.
[0192] If this condition is true, fusion is applied, otherwise only mode 1 is used.
[0193] The weights of the modes are calculated from the SATD cost as follows:
[0194] weight1 = costMode2 / (costMode1+ costMode2)
[0195] weight2 = 1 - weight1
[0196] Spatial Geometric Partitioning Mode (SGPM)
[0197] SGPM is an intra mode similar to GPM's inter-coding tool, where two prediction parts are generated through the intra prediction process. In this mode, a candidate list is created for each entry, containing one partition and two intra prediction modes. The partition mode and three intra prediction modes are used to form a combination. The candidate list length can be set to 16, and the selected candidate index can be signaled.
[0198] The candidate list is reordered using templates, where the SAD between the template's prediction and reconstruction is used for alignment. The template size can be fixed to 1.
[0199] For each partition mode, an Intra Prediction Mode (IPM) list is derived for each part using the same intra-inter GPM list derivation. The size of the IPM list can be set to 3. In the list, the TIMD derived mode can be replaced by two derived modes in the horizontal and vertical directions.
[0200] SGPM mode can be applied with limited block sizes as follows:
[0201] 4<=width<=64, 4<=height<=64, width <height*8, height<width*8, width*height> =32
[0202] Adaptive blending is also used in spatial GPM, and the blending depth τ can be derived as follows:
[0203] If min(width, height)==4, 1 / 2 τ is selected
[0204] else if min(width, height)==8, τ is selected
[0205] else if min(width, height)==16, 2 τ is selected
[0206] else if min(width, height)==32, 4 τ is selected
[0207] else, 8 τ is selected
[0208] Geometric partitioning mode (GPM)
[0209] GPM can be supported for inter prediction. GPM is a type of merge mode that can be signaled using a CU-level flag. Other merge modes can include regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. Each possible CU size is wxh = 2. m x 2 n A total of 64 partitions can be supported, where {3, ..., 6} ∋ m,n, excluding 8x64 and 64x8.
[0210] When this mode is used, a CU can be divided into two parts by a geometrically positioned straight line. The location of the division line can be mathematically derived from the angle and offset parameters of a specific partition. Each part of the geometric partition of the CU can be inter-predicted using its own motion. Only a single prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. A uni-prediction motion constraint can be applied to ensure that only two motion-compensated predictions are required for each CU, similar to conventional bi-prediction.
[0211] Combined inter and intra prediction (CIIP)
[0212] Combined inter and intra prediction (CIIP) may be applied to the current block. An additional flag (e.g., ciip_flag) may be signaled to indicate whether the CIIP mode applies to the current CU. For example, when the CU is coded in merge mode, if the CU contains at least 64 luma samples (i.e., the product of the CU width and the CU height is greater than or equal to 64), and both the CU width and the CU height are less than 128 luma samples, the additional flag may be signaled to indicate whether the CIIP mode applies to the current CU.
[0213] CIIP prediction combines inter-prediction signals and intra-prediction signals. For example, the inter-prediction signal P_inter in CIIP mode can be derived using the same inter-prediction process applied in regular merge mode, and the intra-prediction signal P_intra can be derived according to the regular intra-prediction process in planar mode. The intra- and inter-prediction signals can then be combined using a weighted average.
[0214] GPM including inter and intra prediction
[0215] In a GPM that includes inter- and intra-prediction, final predicted samples can be generated by weighting the inter-predicted samples and intra-predicted samples for each GPM-segmented region. The inter-predicted samples are derived from the inter-GPM, while the intra-predicted samples can be derived from an intra-prediction mode (IPM) candidate list and an index signaled from the encoder. The IPM candidate list size can be predefined as 3.
[0216] Local Illumination Compensation (LIC)
[0217] LIC is an inter-prediction technique for modeling the local illumination change between the current block and the predicted block as a function of the change between the current block template and the reference block template. The parameters of this function can be expressed as a scaling factor α and an offset β that form a linear equation: α*p[x]+β, which compensates for the illumination change, where p[x] is the reference sample pointed to by the MV at position x in the reference picture. If wrap-around motion compensation is enabled, the MV can be clipped considering the wrap-around offset. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them other than signaling the LIC flag to indicate the use of LIC in AMVP mode.
[0218] LIC can be used in inter CUs with the following aspects:
[0219] - Intra-neighbor samples can be used in LIC parameter derivation.
[0220] - LIC may be disabled in blocks with fewer than 32 luma samples.
[0221] - For both non-sub-block and affine modes, LIC parameter derivation can be performed based on template block samples corresponding to the current CU instead of partial template block samples corresponding to the first upper left 16x16 unit.
[0222] - Samples of a reference block template can be generated using block MV and MC without rounding them to integer-pel resolution (or precision).
[0223] For bi-prediction inter-CUs, two sets of LIC parameters can be derived separately for L0 and L1 prediction samples. An iterative method can be applied to derive the L0 and L1 LIC parameters. Specifically, the L0 LIC parameters can first be derived by minimizing the difference between the L0 template prediction T0 and the template T, and the samples of T can be updated by subtracting the corresponding samples of T0. Then, the L1 parameters can be computed to minimize the difference between the L1 template prediction T1 and the updated template. Finally, the L0 parameters can be re-tuned in the same manner.
[0224] Example
[0225] The disclosed embodiment provides a method for more efficiently constructing a candidate list composed of block vector candidates when applying a prediction mode that uses a block vector for prediction of a current block.
[0226] Fig. 14 is a flowchart of a decoding method according to one embodiment, and Fig. 15 is a flowchart of an encoding method according to one embodiment. The decoding method of Fig. 14 can be performed by the decoding device (300) described above, and the encoding method of Fig. 15 can be performed by the encoding device (200) described above. The decoding method of Fig. 14 and the encoding method of Fig. 15 correspond to each other.
[0227] Referring to FIG. 14, a decoding method according to one embodiment may include a step of obtaining information about a prediction mode from a bitstream (S600), and, based on the obtained information about the prediction mode, determining a prediction mode to be applied to a current block as a prediction mode using a block vector (S610). Here, the prediction mode using a block vector may include an intra TMP mode, an IBC merge mode, an IBC mode, etc.
[0228] If the prediction mode for the current block is determined to be a prediction mode using a block vector, a candidate list including block vector candidates for the current block is constructed (S620), and a prediction block can be generated based on at least one block vector candidate included in the candidate list (S630).
[0229] The step (S620) of constructing the above candidate list may include adding block vector information of a previously restored surrounding block of the current block or block vector information derived from the block vector information to the candidate list.
[0230] Additionally, the restored surrounding blocks of the current block may be blocks restored by a combination of one or more prediction tools or one or more prediction modes and at least one of the Intra Block Copy (IBC) mode or the Intra Template Matching Prediction (TMP) mode.
[0231] The encoding method according to the embodiment of FIG. 15 may include a step of determining a prediction mode applied to a current block as a prediction mode using a block vector (S700), a step of constructing a candidate list including block vector candidates for the current block (S710), a step of generating a prediction block based on at least one block vector candidate included in the candidate list (S720), and a step of encoding image information based on the generated prediction block (S730).
[0232] The step (S710) of constructing the above candidate list may include adding block vector information of a previously restored surrounding block of the current block or block vector information derived from the block vector information to the candidate list.
[0233] Additionally, the restored surrounding blocks of the current block may be blocks restored by a combination of one or more prediction tools or one or more prediction modes and at least one of the Intra Block Copy (IBC) mode or the Intra Template Matching Prediction (TMP) mode.
[0234] Below, the details will be explained through specific examples.
[0235] Example 1
[0236] According to one embodiment, a method is provided for more effectively applying intra-TMP using block vectors. Specifically, the present embodiment describes a case where the prediction mode applied to the current block is intra-TMP mode. As described above, intra-TMP is a technique for finding a reconstructed block similar to the current block within a template search region and using it for a prediction block. However, due to the complex search process, the search region is very limited in its use, and block vector information outside the search region is difficult to utilize.
[0237] Meanwhile, IBC mode supports a search area that is broader than that of intra-TMP. Therefore, utilizing the block vector information supported by IBC mode in intra-TMP can generate more accurate prediction blocks.
[0238] FIG. 16 is a diagram illustrating an example of a decoding method according to one embodiment of the present disclosure. The decoding method of FIG. 16 can be performed by the aforementioned decoding device (300). While each step is illustrated in a flowchart for convenience of explanation, the order in which each step is performed is not necessarily limited to the order in the flowchart. Depending on variations in the embodiment, the order may be performed differently, some steps may be omitted, or other steps may be added.
[0239] As previously described with reference to FIG. 4, the decoding device (300) can obtain image information from a bitstream (S800). According to the present embodiment, the image information may include information related to intra TMP. For example, the information related to intra TMP may be information indicating whether intra TMP is applied to the current block, and this may be expressed as an intra TMP flag. Hereinafter, a specific embodiment will be described using an example in which the information related to intra TMP is an intra TMP flag.
[0240] The decoding device (300) can decode the acquired intra TMP flag information and perform the intra TMP mode if the flag indicates that intra TMP is applied to the current block (e.g., if the flag value indicates 1 or true).
[0241] At this time, whether to use intra TMP mode may be determined based on the size or shape of the block to reduce signaling overhead and improve compression performance. For example, for non-square blocks, intra TMP mode may always be applied. As a specific example, when the width of the block is more than twice the height (i.e., width >= 2*height), intra TMP mode may always be applied. Alternatively, if the width*height of the current block is greater than 4096, intra TMP mode may be determined not to be applied (i.e., the value of the intra TMP flag may be 0 or false) or may be induced, in which case intra TMP mode is not performed.
[0242] The above example relates to a case where the application of intra TMP is determined based on the size or shape of the block, but it can also be extended or changed to other values and / or other conditions.
[0243] When intra TMP is applied to the current block, the decoding device (300) can define a search area for template matching (S810). The search area may vary depending on the size of the picture, the size of the CTU, the size of the current block, etc. For example, when the width and height of the current block are W and H, respectively, the search area may be defined as a previously restored area within 5*H above, 5*H below, 5*W to the left, and 5*W to the right based on the upper left of the current block. Alternatively, the search area may be defined as a previously restored area of a CTU row including the current block, or CTU rows adjacent to the upper side of the CTU including the current block. Alternatively, the search area may be defined as a combination of these. This is one example in which the search area is determined in relation to the size of the picture, the size of the block, the size of the CTU, etc., and it is to be understood that the search area may be extended or changed to another search area defined in advance in the encoding device / decoding device.
[0244] The decoding device (300) can perform template matching and candidate list construction (S820). The decoding device (300) can perform template matching in a given search area, find templates similar to the template of the current block, and construct a candidate list using block vectors for blocks having similar templates. For example, the size of the intra-TMP candidate list can be a value predefined in the encoding device / decoding device. For example, the size of the candidate list can be 30.
[0245] For example, candidates in a candidate list can be stored in ascending order based on an error value. Here, the error value can be calculated using an error calculation method between the template of the current block and the template of the reference block indicated by each candidate block vector. Error calculation methods that can be used include SAD (Sum of Absolute Difference), SATD (Sum of Transformed Difference), SSE (Sum of Squared Error), MR-SAD (Mean-Removed Sum of Difference), MR-SSE (Mean-Removed Sum of Squared Error), and MR-SATD (Mean-Removed Sum of Transformed Difference).
[0246] Figure 17 is a drawing showing an example of a template shape used for template matching.
[0247] As illustrated in Fig. 17, template shapes used for template matching may include L-shape, above-shape, left-shape, and above & left-shape. Information regarding these shapes may be signaled, or may be defined in advance in an encoding device / decoding device and applied without signaling.
[0248] Additionally, a variable template size can be used during the template matching process, which can also be signaled, or can be predefined in the encoding / decoding device and applied without signaling. For example, the template size can be determined based on the size of the current block. Specifically, if the width * height of the current block is less than 64, a template consisting of two pixel lines from the top and two pixel lines from the left of the current block can be used. Otherwise, a template consisting of four pixel lines from the top and four pixel lines from the left can be used.
[0249] As another example, the size of the template may be determined based on the shape of the current block. For example, for a non-square block, a template consisting of four pixel lines on the top and four pixel lines on the left may be used. The above-mentioned cases are examples in which the template size is determined based on the size and shape of the current block, and it is of course possible for the template size to be determined based on other values and / or conditions.
[0250] The decoding device (300) can generate a prediction block using a candidate list (S830). For example, N candidates (N is a natural number) within the intra-TMP candidate list can be used. At this time, information regarding which block vector to use in the candidate list and which prediction block to generate can be signaled and transmitted from the encoding device, or can be predefined in the encoding device / decoding device. An example of a prediction block generated according to one embodiment is as follows:
[0251] - Single predictor: A prediction block can be generated by copying the reference block at the location indicated by the block vector.
[0252] - Fusion of multiple predictors: The final predicted block is generated by blending reference blocks at positions indicated by multiple block vectors. At this time, the blending weights can be calculated based on the template matching error value of each reference block, calculated by deriving weights based on a Wiener filter, using values predefined between the encoding device / decoding device, or defining a preset for the weights commonly between the encoding device / decoding device and determining it through explicit signaling.
[0253] - Sub-pel precision predictor: Fig. 18 is a diagram showing the position of a sub-pel applicable to the disclosed embodiment. For example, a prediction block corresponding to the corresponding sub-pel block vector position can be generated by interpolating surrounding samples of a pixel position of a reference block at a position indicated by a block vector (candidate pixel position in Fig. 18). The sub-pel unit can be Half-pel, 1 / 4-pel, 3 / 4-pel as shown in Fig. 18, and can be positioned according to 8 directions. The position shown in Fig. 18 is only one example, and other sub-pels such as 1 / 8-pel, 1 / 16-pel, 1 / 32-pel, or other directions such as 4 or 16 directions can also be used.
[0254] - Filter model: A filter model can be generated using the relationship between the template of the reference block indicated by the block vector and the template of the current block, and a prediction block can be generated by applying the filter model to the reference block. Fig. 19 is a diagram showing a method for deriving filter coefficients of a filter model applied to an intra TMP. For example, as shown in Fig. 19, filter coefficients can be derived so that the error between the current template and the reference template is minimized. As a specific example, Output is a weighted sum of filter coefficients in C, E, W, S, and N in the reference template, as in Equation 1 below. i The filter coefficients can be derived so that P becomes the current template.
[0255] [Formula 1]
[0256] Output i = c0*C+c1*N+c2*S+c3*E+c4*W+c5*B
[0257] Output i : the i-th sample in the current template area
[0258] c0, c1, c2, c3, c4, c5: filter coefficients
[0259] C: Any sample within the reference template area
[0260] N: Upper sample adjacent to C
[0261] S: Lower sample adjacent to C
[0262] W: Left sample adjacent to C
[0263] E: Right sample adjacent to C
[0264] B: bias term, which can have a value of (1 << (bit depth - 1)) for example
[0265] P: Sample at the same location as C in the reference template within the current template area
[0266] The filter model described above is an example of one of the linear filter models applicable to one embodiment, and in addition, the number of samples calculated in the filter model may be different, and non-linear terms such as the square of the sample value may be included in the filter model.
[0267] The method for generating an intra TMP prediction block is not limited to the above-described method. The above-described prediction blocks are examples applicable to the present disclosure, and intra TMP prediction blocks can be generated in various other ways. For example, it is also possible to generate a new prediction block by applying LIC (Local Illumination Compensation) to the above-described prediction block. As described above, LIC is a technique for compensating for illumination changes by modeling local illumination changes between the template of the current block and the template of the reference block, and is a method for compensating for illumination changes with a linear equation α*p[x]+β using scale α and offset β. Here, p[x] represents a reference sample within the template of the reference block.
[0268] As mentioned above, the typical intra-TMP mode has limitations because it can only be applied within a limited search area. However, the IBC mode can be applied to a wider search area than the intra-TMP mode. Furthermore, the complex template matching in the intra-TMP mode can sometimes result in sparse searches, which can result in failure to find the optimal block vector. Therefore, embodiments according to the present disclosure can enhance compression performance by utilizing block vectors, such as IBC block vectors, in the intra-TMP mode, which are not supported in the typical intra-TMP mode.
[0269] According to one embodiment, in the template matching and candidate list construction step (S820), previously restored intra TMP block vector information may be included in the intra TMP candidate list. Here, the previously restored intra TMP block vector information may include information about a block vector used for prediction (restoration) of a previously restored neighboring block of the current block when the neighboring block is restored by applying the intra TMP mode. In this case, the neighboring block may include a block located within the search area of the current block.
[0270] In addition, in the template matching and candidate list construction step (S820), previously restored IBC block vector information may be included in the intra TMP candidate list. Here, the previously restored IBC block vector information may include information about a block vector used for prediction (restoration) of a previously restored neighboring block of the current block when the neighboring block is restored by applying the IBC mode. At this time, the neighboring block may include a block located within the search area of the current block. In the present embodiment, the IBC mode may include an IBC mode (=IBC AMVP mode) that transmits block vector information and block vector difference (BVD) information, and an IBC merge mode that transmits block vector information but does not transmit BVD information. However, even if the BVD value is not directly transmitted in the IBC merge mode, BVD-related information (distance and direction information) for supplementing the block vector may be transmitted.
[0271] Meanwhile, in the disclosed embodiments, the IBC mode can be applied in combination with one or more prediction tools or prediction modes. Examples of possible combinations are listed below:
[0272] - GPM-IBC: GPM (Geometric partitioning mode) is one of the inter technologies that performs different predictions on each divided area through partitioning. Each divided area can perform predictions in various combinations such as inter / inter (e.g. prediction using different motion vectors), inter / intra, etc., and GPM-IBC is a mode that performs inter prediction and IBC prediction, or IBC prediction and IBC prediction for each divided area.
[0273] - SGPM-IBC: SGPM (Spatial GPM) is one of the intra techniques that performs different predictions on each region divided by partitioning. Each divided region can perform predictions of an intra / intra combination (e.g., prediction using different intra modes), and SGPM-IBC is a mode that performs intra prediction and IBC prediction, or IBC prediction and IBC prediction, for each divided region.
[0274] When applying the IBC mode together to GPM and SGPM, which perform different predictions in each region divided by partitioning, the IBC / IBC combination is possible in both GPM and SGPM, as in the example above. In this case, the encoding device / decoding device can define the combination in advance so that the combination of prediction modes does not overlap. For example, in an I-slice (or I-frame) that can only perform intra prediction, the SGPM mode can allow the IBC / IBC combination. In addition, in an I-slice (or I-frame), p-slice (or p-frame), and b-slice (or b-frame) that can perform intra prediction and inter prediction, the IBC / IBC combination can be allowed only for the GPM mode. Alternatively, the IBC / IBC combination can be allowed only for the SGPM mode. Here, the SGPM mode that allows combination with the IBC mode may mean the SGPM-IBC mode, and the GPM mode that allows combination with the IBC mode may mean the GPM-IBC mode.
[0275] In addition, if either the GPM-IBC mode or the SGPM-IBC mode, which allows combination with the IBC mode at the sequence parameter set (SPS) or picture parameter set (PPS) level, is turned off, the IBC / IBC combination can be allowed for the mode that is not turned off. For example, if GPM-IBC is off and SGPM-IBC is on, SGPM-IBC can allow the IBC / IBC combination. The above embodiments are examples that allow the IBC combination depending on the on / off status of the two modes, GPM-IBC and SGPM-IBC, and in addition, the combination of GPM-IBC and SGPM-IBC can be allowed so that they do not overlap between the encoding device / decoding device in another way.
[0276] - DIMD-IBC: This is a mode in which the DIMD mode and the IBC mode are blended together to generate the final prediction block, or a non-directional mode (Planar or DC mode) is replaced with the IBC mode and the final prediction block is generated after blending.
[0277] - TIMD-IBC: This is a mode in which the TIMD mode and the IBC mode are blended together to generate the final prediction block, or a non-directional mode (Planar or DC mode) is replaced with the IBC mode and the final prediction block is generated after blending.
[0278] - CIIP-IBC: CIIP (Combined Intra-Inter prediction) is a technology that generates prediction blocks by blending intra-prediction and inter-prediction. In CIIP-IBC mode, the final prediction block can be generated by blending a combination of intra- and inter-prediction.
[0279] - LIC-IBC: This is a mode that models local illumination variation between the template of the current block and the template of the IBC prediction block using LIC (Local Illumination Compensation) technology, and then compensates for the illumination variation to generate the final prediction block.
[0280] That is, block vector information of a neighboring block previously restored by the IBC mode used in combination with other tools, such as the example described above, for example, block vector information used to restore the neighboring block, may be included in the intra TMP candidate list. For example, block vector information of a neighboring block previously restored by GPM-IBC may be included in the intra TMP candidate list. In addition, block vector information of a neighboring block previously restored by SGPM-IBC may be included in the intra TMP candidate list. In addition, block vector information of a neighboring block previously restored by DIMD-IBC may be included in the intra TMP candidate list. In addition, block vector information of a neighboring block previously restored by TIMD-IBC may be included in the intra TMP candidate list. In addition, block vector information of a neighboring block previously restored by CIIP-IBC may be included in the intra TMP candidate list. In addition, block vector information of a neighboring block previously restored by LIC-IBC may be included in the intra TMP candidate list.
[0281] The above combinations are just examples, and it is also possible for IBC mode to be combined with other tools not mentioned above.
[0282] Additionally, the above example describes an IBC mode using block vectors as an example of a prediction mode applicable to the disclosed embodiment. Other prediction modes using block vectors may also be applicable to the disclosed embodiment. For example, an intra-TMP used in combination with other tools, such as GPM-IntaTMP, SGPM-IntaTMP, DIMD-IntaTMP, TIMD-IntaTMP, CIIP-IntaTMP, or LIC-IntaTMP, may also be applicable to the disclosed embodiment.
[0283] That is, the block vector information of the surrounding blocks restored by the Intra-TMP combined with other tools, for example, the block vector information used to restore the surrounding blocks, may be included in the Intra-TMP candidate list of the current block. For example, the block vector information of the surrounding blocks previously restored by the GPM-IntraTMP may be included in the Intra-TMP candidate list. In addition, the block vector information of the surrounding blocks previously restored by the SGPM-IntraTMP may be included in the Intra-TMP candidate list. In addition, the block vector information of the surrounding blocks previously restored by the DIMD-IntraTMP may be included in the Intra-TMP candidate list. In addition, the block vector information of the surrounding blocks previously restored by the TIMD-IntraTMP may be included in the Intra-TMP candidate list. In addition, the block vector information of the surrounding blocks previously restored by the CIIP-IntraTMP may be included in the Intra-TMP candidate list. In addition, the block vector information of the surrounding blocks previously restored by the LIC-IntraTMP may be included in the Intra-TMP candidate list.
[0284] Additionally, two or more prediction modes that use block vectors can be combined together. For example, combinations of GPM with IBC and IntaTMP, SGPM with IBC and IntaTMP, DIMD with IBC and IntaTMP, or TIMD with IBC and IntaTMP are possible. Here, the combination of Intra TMP and IBC can mean that the block vector of one of the two modes is used, or the block vectors of both modes are used. For example, SGPM with IBC and IntaTMP can be combined as Intra / IBC, Intra / IntraTMP, IntraTMP / IBC, IBC / Intra, intraTMP / Intra, IBC / IntraTMP, IBC / IBC, IntraTMP / IntraTMP, etc.
[0285] That is, block vector information of a neighboring block reconstructed by a combination of two or more prediction modes, for example, block vector information used to reconstruct the neighboring block, may be included in the intra TMP candidate list of the current block. For example, block vector information of a neighboring block previously reconstructed by GPM with IBC and IntaTMP may be included in the intra TMP candidate list. Block vector information of a neighboring block previously reconstructed by SGPM with IBC and IntaTMP may be included in the intra TMP candidate list. Block vector information of a neighboring block previously reconstructed by DIMD with IBC and IntaTMP may be included in the intra TMP candidate list. Block vector information of a neighboring block previously reconstructed by TIMD with IBC and IntaTMP may be included in the intra TMP candidate list. Block vector information of a neighboring block previously reconstructed by CIIP with IBC and IntaTMP may be included in the intra TMP candidate list.
[0286] In addition, in the template matching and candidate list construction step (S820), it is also possible for candidate block vector information of a history-based candidate list to be included in the intra TMP candidate list. For example, as in the above embodiments, if the prediction mode applied to the previously restored surrounding blocks is a mode in which block vectors are derived, such as IntraTMP, IBC, SGPM-IBC, DIMD-IBC, TIMD-IBC, CIIP-IBC, and LIC-IBC, the derived block vector information may be included in the history-based candidate list. Accordingly, if the intra TMP prediction mode is applied to the current block, block vectors stored in the history-based candidate list may be included in the intra TMP candidate list.
[0287] Meanwhile, the size of the history-based candidate list can be predefined in the encoding / decoding device. For example, the size of the history-based candidate list can be defined as 6. In addition, the history-based candidate list can be initialized for each CTU. In addition, the history-based candidate list can be initialized whenever the CTU row position changes.
[0288] In addition, in the template matching and candidate list construction step (S820), a block vector indicating a position indicated by a block vector of a surrounding block (reference block), i.e., a block vector indicating a position of a reference block of a surrounding block, can be derived, and block vector information indicating this derived block vector can be included in the intra TMP candidate list.
[0289] FIG. 20 is a diagram illustrating a process of serially deriving reference blocks of surrounding blocks in the disclosed embodiment.
[0290] For example, as shown in Fig. 20, a block vector BV indicating the position of the restored surrounding block (reference block) B1 of the current block 0,1 When , the block vector stored in B1 (= BV 1,2 ) can be derived so that the position indicated by BV can be represented based on the current block. That is, BV 0,2 is BV 0,1 +BV 1,2 can be derived. In the same way, the block vector (=BV) stored in B2 2,3 ) can be derived so that the position indicated by BV can be represented based on the current block. That is, BV 0,3 Silver BV 0,1 +BV 1,2 +BV 2,3can be derived. In addition, when a surrounding block (reference block) is restored by the IBC mode and / or IntraTMP mode combined with other tools, the stored block vector can be utilized to derive block vector candidates. For example, the block vector information of a surrounding block restored by the IBC mode / IntraTMP combined with other tools, for example, the block vector information used to restore the surrounding block, can be utilized to derive block vector candidates.
[0291] In this way, when block vector information is stored in surrounding blocks, the process of accessing the location indicated by the block vector can be repeated, and the block vector derived through this can be utilized as a block vector candidate. Furthermore, in the process of deriving the block vector, the encoding device / decoding device can limit the number of repetitions to a predefined value. For example, the location of the reference block can be derived using only the block vector of the second reference block.
[0292] In addition, the reference block may store different block vectors in units of subblocks. The size of the subblock may be, for example, 4 by 4. Therefore, in the process of utilizing the block vectors stored in the reference block, the encoding device / decoding device may derive block vector information stored in a predefined order. For example, the block vector candidates may be derived by utilizing the block vectors stored in the order of center - left top - right top - left bottom - right bottom within the reference block. Alternatively, the encoding device / decoding device may derive the block vector candidates by utilizing only the block vectors stored in the predefined positions. For example, the block vector candidates may be derived by utilizing only the block vectors stored in the center position within the reference block.
[0293] Additionally, in the template matching and candidate list construction step (S820), it is also possible to derive or generate new candidates by combining different candidates included in the intra TMP candidate list. For example, a new candidate can be generated by calculating the average of the candidate at any position in the intra TMP candidate list and the sign at another arbitrary position, and this new candidate can be included in the intra TMP candidate list. In other words, a third block vector can be generated by averaging the first and second block vectors included in the intra TMP candidate list, and the generated third block vector can be included in the intra TMP candidate list.
[0294] As described above, the intra TMP candidate list can be sorted based on the error value. For example, the block vector corresponding to the first candidate in the sorted candidate list can be the first block vector, and the block vector corresponding to the second candidate can be the second block vector. A third block vector generated by averaging the first and second block vectors can be included in the intra TMP candidate list. Of course, the derived block vector (the third block vector) can also be sorted based on the error value. The above example corresponds to an example of deriving a new candidate by combining candidate block vectors in the intra TMP candidate list, and in addition, new candidates can be derived by other combinations or other methods in the candidate list according to the predefined configuration of the encoding device / decoding device.
[0295] Meanwhile, in the disclosed embodiment, the surrounding blocks used to construct the intra TMP candidate list of the current block may include blocks located adjacent / non-adjacent to the current block.
[0296] Figure 21 is a diagram showing an example of the locations of surrounding blocks used to construct an intra TMP candidate list of the current block.
[0297] Referring to the example of Fig. 21, when the upper left coordinate of the current block is (x, y), and the width and height of the current block are W, H, the adjacent positions may include, for example, positions (x-1, y-1), (x+W-1, y-1), (x+W, y-1), (x-1, y+H), (x-1, y+H-1). This is one example, and the encoding device / decoding device may define other adjacent positions in advance.
[0298] Additionally, non-adjacent positions can be expressed as NonAdj_pos as shown in [Table 1] below. This is an example, and only some NonAdj_pos positions may be included in the example below, or the encoding device / decoding device may define other non-adjacent positions in advance.
[0299] [Table 1]
[0300]
[0301]
[0302] As another example, a non-adjacent position can be expressed as NonAdj_pos2, as shown in [Table 2] below. This is also an example, and only some NonAdj_pos2 positions may be included in the example below, or the encoding / decoding device may define other non-adjacent positions in advance.
[0303] [Table 2]
[0304]
[0305] For example, the above defined location may be limited to a location within the search area defined in the search area definition step.
[0306] FIG. 22 is a diagram illustrating an example of an encoding method according to one embodiment of the present disclosure. The encoding method of FIG. 22 may be performed by an encoding device (200). While each step is illustrated in a flowchart for convenience of explanation, the order in which each step is performed is not necessarily limited to the order in the flowchart. Depending on variations in the embodiment, the order may be performed differently, some steps may be omitted, or other steps may be added.
[0307] The encoding device (200) can determine a prediction mode for the current block (S900). At this time, if the prediction mode for the current block is determined to be the intra TMP mode, the encoding device (200) can define a search area (S910), configure a template matching and candidate list (S920), and generate a prediction block (S930).
[0308] The encoding method of FIG. 22 corresponds to the decoding method of FIG. 16 described above. The encoding device (200) and the decoding device (300) generate prediction blocks according to the same rules. Therefore, the description of the prediction method among the decoding methods described above with reference to FIG. 16 can be equally applied to the prediction method among the encoding methods according to the present example. For example, in the step (S900) of determining the prediction mode for the current block, as described above with reference to FIG. 16, whether to use the intra-TMP mode may be determined based on the size or shape of the block in order to reduce signaling overhead and improve compression performance. In addition, other descriptions can be equally applied as long as they do not conflict with the operation of the encoding device (200).
[0309] In addition, in the step (S910) of defining a search area based on the application of intra TMP to the current block, the encoding device (200) can define a search area for template matching, and the search area can vary depending on the size of the picture, the size of the CTU, the size of the current block, etc. Other descriptions can be equally applied as long as they do not conflict with the operation of the encoding device (200).
[0310] In addition, in the step (S920) of performing template matching and constructing a candidate list, the encoding device (200) may perform template matching in a given search area, find a template similar to the template of the current block, and construct a candidate list using block vectors for blocks having similar templates. In addition, other descriptions may be equally applicable as long as they do not conflict with the operation of the encoding device (200). In particular, the description of a method of constructing a candidate list using block vectors of surrounding blocks that have been restored by a prediction mode in which block vectors are derived or used, such as IBC mode, intra-TMP mode, IBC mode combined with other tools, intra-TMP mode combined with other tools, or a combination of intra-TMP mode and IBC mode, may be equally applicable to a prediction method performed by the encoding device (200).
[0311] In addition, in the step of generating a prediction block (S930), the encoding device (200) can generate a prediction block for the current block using the generated candidate list. For example, N candidates (N is a natural number) in the intra TMP candidate list can be used. At this time, the encoding device (200) can signal information about which block vector to use in the candidate list, which prediction block to generate, etc., and transmit it to the decoding device (300) in the form of a bitstream. Alternatively, it can be defined in advance in the encoding device / decoding device. In addition, other descriptions can be equally applied as long as they do not conflict with the operation of the encoding device (200).
[0312] In the step of encoding image information (S940), the image information encoded may include information on prediction, residual information, etc. To reduce redundant explanations, the description has been omitted, but it is of course possible to include procedures such as residual processing described above with reference to FIG. 5 in the encoding method according to the example. The residual information generated by the residual processing may be included in the image information and encoded. The information on prediction may refer to information necessary for predicting the current block, and specifically may include information on the prediction mode applied to the current block (or information indicating the prediction mode applied to the current block), index information, etc. The type of information to be encoded may vary depending on the prediction mode applied to the current block, and in the example, intra TMP flag information may be included in the image information.
[0313] When image information encoded according to the above-described procedure is transmitted to a decoding device (300) in the form of a bitstream, the decoding device (300) can obtain image information from the transmitted bitstream and perform the above-described decoding method.
[0314] Example 2
[0315] According to one embodiment, a method for utilizing block vectors to more effectively apply the IBC merge mode is provided. That is, the present embodiment describes a case where the prediction mode applied to the current block is the IBC merge mode. The IBC merge mode generates a prediction block by finding a previously reconstructed block similar to the current block through block matching. However, since the block vector candidates used for block matching are limited, the disclosed embodiment provides a method for improving compression performance by utilizing a wider variety of block vector candidates.
[0316] FIG. 23 is a diagram illustrating an example of a decoding method according to one embodiment of the present disclosure. The decoding method of FIG. 23 can be performed by the aforementioned decoding device (300), and while each step is illustrated in a flowchart for convenience of explanation, the order in which each step is performed is not necessarily limited to the order in the flowchart. Depending on variations in the embodiment, the order may be performed differently, and some steps may be omitted or other steps may be added.
[0317] As previously described with reference to FIG. 4, the decoding device (300) can obtain image information from a bitstream (S1000). According to the present embodiment, the image information may include information related to the IBC merge mode. For example, the information related to the IBC merge mode may be information indicating whether the IBC merge mode is applied to the current block, and this may be expressed as an IBC merge flag. Hereinafter, a specific embodiment will be described using an example in which the information related to the IBC merge mode is an IBC merge flag.
[0318] The decoding device (300) can decode the acquired IBC merge flag information and perform the IBC merge mode when the flag indicates that the IBC merge mode is applied to the current block (e.g., when the flag value indicates 1 or true).
[0319] At this time, whether to use IBC merge mode may be determined based on the size or shape of the block to reduce signaling overhead and improve compression performance. For example, in the case of non-square blocks, IBC merge mode may always be applied. As a specific example, when the width of the block is more than twice the height (i.e., width >= 2*height), IBC merge mode may always be applied. Alternatively, when the width*height of the current block is greater than 4096, IBC merge mode may be determined not to be applied (i.e., the value of the IBC merge flag may be 0 or false) or induced, in which case IBC merge mode is not performed.
[0320] The above example is about a case where whether or not to apply IBC merge mode is determined based on the size or shape of the block, but it can also be extended or changed to other values and / or other conditions.
[0321] When the IBC merge mode is applied to the current block, the decoding device (300) can define a search area for performing the IBC merge mode (S1010). The search area may vary depending on the size of the picture, the size of the CTU, the size of the current block, etc. For example, when the width and height of the current block are W and H, respectively, the search area may be defined as a previously reconstructed area within 5*H above, 5*H below, 5*W to the left, and 5*W to the right based on the upper left of the current block. Alternatively, when the CTU size is C, the search area may be defined as an area within C*2 to the left, C / 4 to the right, C above, and C / 4 below based on the current block. Alternatively, the search area may be defined as a previously reconstructed area of row CTUs including the current block, and adjacent upper row CTUs of row CTUs including the current block. Alternatively, the search area may be defined as a combination of these. This is an example where the search area is defined based on block size, CTU size, etc., and it is also possible to define and use other search areas in advance in the encoding device / decoding device.
[0322] The decoding device (300) may configure a candidate list (S1020). If previously restored surrounding blocks are restored using the IBC mode or the Intra-TMP mode, block vector information stored in the corresponding surrounding blocks may be included in the merge candidate list. In order to utilize the block vectors of previously restored surrounding blocks, block vectors stored in adjacent / non-adjacent locations to the current block may be utilized. The adjacent / non-adjacent locations may be locations predefined by the encoding device / decoding device. For example, the adjacent locations may be the above-right, bottom-left, and above-left locations of the current block. In addition, a pairwise average candidate (a candidate obtained by averaging two different candidates by a predefined combination in the merge candidate list) may be included in the merge candidate list. This is an example, and other candidates may be added to the merge candidate list according to a predefined value in the encoding device / decoding device. Meanwhile, the size of the merge candidate list may be a value predefined by the encoding device / decoding device, for example, 16.
[0323] The decoding device (300) can refine and reorder block vector candidates of the merge candidate list (S1030). Specifically, the decoding device (300) can refine the block vector candidates of the merge candidate list using template matching (TM). An error value can be calculated between the template of the current block and the template at the position indicated by each block vector candidate, and each block vector candidate can be refined based on the error value. In addition, the block vector candidates of the merge candidate list can be reordered based on the error value. Through this process, signaling overhead can be reduced when signaling block vector candidate information. An error value can be calculated between the template at the position indicated by each block vector candidate of the merge candidate list and the template of the current block, and the merge candidate list can be sorted in order of the smallest template error value.
[0324] Error calculation methods that can be used include SAD (Sum of absolute difference), SATD (Sum of transformed difference), SSE (Sum of squared error), MR-SAD (Mean-removed sum of difference), MR-SSE (Mean-removed sum of squared error), and MR-SATD (Mean-removed sum of transformed difference).
[0325] Meanwhile, template shapes for template matching can be L-shape, above-shape, left-shape, above & left-shape, etc. as illustrated in FIG. 17 described above. Information about these shapes can be signaled, or can be defined in advance in an encoding device / decoding device and applied without signaling.
[0326] In addition, a variable template size can be used during the template matching process, which can also be signaled, or can be predefined in the encoding device / decoding device and applied without signaling. For example, the size of the template can be determined according to the size and shape of the block. For example, if the width*height of the current block is less than 64, a template consisting of one pixel line from the top and one pixel line from the left of the current block can be used. Otherwise, a template consisting of two pixel lines from the top and two pixel lines from the left can be used. In the case of non-square blocks, a template consisting of one pixel line from the top and one pixel line from the left can always be used. The above examples are examples in which the size of the template is determined based on the size and shape of the current block, and it is of course possible to determine it according to other values and / or other conditions.
[0327] The decoding device (300) can obtain block vector difference (BVD) information (S1040). Here, the BVD information may refer to BVD-related information. If a prediction block is generated using only block vector candidate information of a merge candidate list without BVD, prediction performance may deteriorate. According to the present embodiment, this drawback can be compensated for by signaling BVD-related information including distance information and / or direction information. For example, the BVD information may include distance information and information on four directions (two horizontal and two vertical directions). This information can be added to the block vector candidate to access a reference block location that is more similar to the current block.
[0328] For example, the distance information can be {1-pel, 2-pel, 4-pel, 8-pel, 12-pel, 16-pel, 24-pel, 32-pel, 40-pel, 48-pel, 56-pel, 64-pel, 72-pel, 80-pel, 88-pel, 96-pel, 104-pel, 112-pel, 120-pel, 128-pel}. However, this is just an example, and various distance information such as 6-pel, 10-pel, etc. can be supported, and various directions such as 8 or 16 directions can be supported.
[0329] In addition, when signaling BVD information, if there is a lot of distance information and direction information, signaling overhead may occur. In this case, by using template matching technology, each BVD candidate (distance and direction information) can calculate an error value between the template of the position added to the block vector candidate and the template of the current block, and signal the BVD candidate with the smallest error value or signal one of the K BVD candidates with the smallest error value (K is a natural number). Error calculation methods that can be used include SAD (Sum of absolute difference), SATD (Sum of transformed difference), SSE (Sum of squared error), MR-SAD (Mean-removed sum of difference), MR-SSE (Mean-removed sum of squared error), and MR-SATD (Mean-removed sum of transformed difference).
[0330] The decoding device (300) can generate a prediction block (S1050). The decoding device (300) can obtain block vector information from the merge candidate list to generate an IBC prediction block. Furthermore, an improved block vector can be derived from BVD information, thereby generating an IBC prediction block. At this time, information regarding which candidate block vector to use from the merge candidate list, which prediction block to generate, and whether to improve the block vector from the BVD information can be obtained through signaling or defined in advance in the encoding device / decoding device. The IBC prediction block can be generated as follows:
[0331] - Single predictor: Generates a predicted block by copying the reference block at the location indicated by the block vector.
[0332] - Single predictor with BVD: Generates a predicted block by copying the reference block at the location indicated by the improved block vector from the BVD information.
[0333] - GPM-IBC predictor: GPM (Geometric partitioning mode) is one of the inter technologies that performs different predictions on each region divided by partitioning. Each divided region can perform predictions in various combinations such as inter / inter (e.g. prediction using different motion vectors), inter / intra, etc., and GPM-IBC performs inter prediction and IBC prediction, or IBC prediction and IBC prediction for each divided region to generate a prediction block.
[0334] - CIIP-IBC predictor: CIIP (Combined Intra-Inter prediction) is one of the inter technologies that generates prediction blocks by blending intra and inter. CIIP-IBC generates the final prediction block after blending the combination of intra and inter-block prediction.
[0335] - Filter Model: A filter model is created based on the relationship between the template of the reference block indicated by the block vector and the template of the current block, and the filter model is applied to the reference block to generate a predicted block. For example, as described with reference to FIG. 17 above, filter coefficients can be derived so that the error between the current template and the reference template is minimized. A detailed description of the method for deriving the filter coefficients is as described above.
[0336] The IBC prediction block in the disclosed embodiment is not limited to the prediction block described above. The above-described prediction blocks are merely examples, and various other IBC prediction blocks may be generated. For example, a new prediction block may be generated by applying a LIC (Local Illumination Compensation) technique to the above-described prediction block. Alternatively, multiple IBC prediction blocks may be blended by a weight value predefined in an encoding device / decoding device. For example, the blending weight may be calculated based on a template matching error value of each reference block, calculated by deriving the weight based on a Wiener filter, or may be a weight value predefined in an encoding device / decoding device, or a preset for the weight may be commonly defined in an encoding device / decoding device and determined through explicit signaling.
[0337] Conventional IBC merge mode has limitations because it uses only a limited number of block vector candidates for prediction. The disclosed embodiment improves compression performance by using a wider variety of block vector candidates.
[0338] Therefore, in the disclosed embodiment, in the candidate list construction step (S1020), block vector information previously restored by the IBC mode may be included in the IBC merge candidate list. Here, the block vector information previously restored by the IBC mode may include information about a block vector used for prediction (restoration) of a previously restored neighboring block of the current block when the neighboring block is restored by applying the IBC mode. At this time, the neighboring block may include a block located within the search area of the current block. In addition, the IBC mode may include an IBC mode (=IBC AMVP mode) that transmits block vector information and block vector difference (BVD) information, and an IBC merge mode that transmits block vector information but does not transmit BVD information. However, even if the BVD value is not directly transmitted in the IBC merge mode, BVD-related information (distance and direction information) for supplementing the block vector may be transmitted.
[0339] Additionally, in the candidate list configuration step (S1020), previously restored intra TMP block vector information may be included in the IBC merge candidate list. Here, the previously restored intra TMP block vector information may include information about the block vector used for prediction (restoration) of previously restored neighboring blocks of the current block when the neighboring blocks are restored by applying the intra TMP mode. In this case, the neighboring blocks may include blocks located within the search area of the current block.
[0340] Meanwhile, in the disclosed embodiments, the IBC mode can be applied in combination with one or more prediction tools or prediction modes. Examples of possible combinations are listed below:
[0341] - GPM-IBC: GPM (Geometric partitioning mode) is one of the inter technologies that performs different predictions on each divided area through partitioning. Each divided area can perform predictions in various combinations such as inter / inter (e.g. prediction using different motion vectors), inter / intra, etc., and GPM-IBC is a mode that performs inter prediction and IBC prediction, or IBC prediction and IBC prediction for each divided area.
[0342] - SGPM-IBC: SGPM (Spatial GPM) is one of the intra techniques that performs different predictions on each region divided by partitioning. Each divided region can perform predictions of an intra / intra combination (e.g., prediction using different intra modes), and SGPM-IBC is a mode that performs intra prediction and IBC prediction, or IBC prediction and IBC prediction, for each divided region.
[0343] When applying the IBC mode together to GPM and SGPM, which perform different predictions in each region divided by partitioning, the IBC / IBC combination is possible in both GPM and SGPM, as in the example above. In this case, the encoding device / decoding device can define the combination in advance so that the combination of prediction modes does not overlap. For example, in an I-slice (or I-frame) that can only perform intra prediction, the SGPM mode can allow the IBC / IBC combination. In addition, in an I-slice (or I-frame), p-slice (or p-frame), and b-slice (or b-frame) that can perform intra prediction and inter prediction, the IBC / IBC combination can be allowed only for the GPM mode. Alternatively, the IBC / IBC combination can be allowed only for the SGPM mode. Here, the SGPM mode that allows combination with the IBC mode may mean the SGPM-IBC mode, and the GPM mode that allows combination with the IBC mode may mean the GPM-IBC mode.
[0344] In addition, if either the GPM-IBC mode or the SGPM-IBC mode, which allows combination with the IBC mode at the sequence parameter set (SPS) or picture parameter set (PPS) level, is turned off, the IBC / IBC combination can be allowed for the mode that is not turned off. For example, if GPM-IBC is off and SGPM-IBC is on, SGPM-IBC can allow the IBC / IBC combination. The above embodiments are examples that allow the IBC combination depending on the on / off status of the two modes, GPM-IBC and SGPM-IBC, and in addition, the combination of GPM-IBC and SGPM-IBC can be allowed so that they do not overlap between the encoding device / decoding device in another way.
[0345] - DIMD-IBC: This is a mode in which the DIMD mode and the IBC mode are blended together to generate the final prediction block, or a non-directional mode (Planar or DC mode) is replaced with the IBC mode and the final prediction block is generated after blending.
[0346] - TIMD-IBC: This is a mode in which the TIMD mode and the IBC mode are blended together to generate the final prediction block, or a non-directional mode (Planar or DC mode) is replaced with the IBC mode and the final prediction block is generated after blending.
[0347] - CIIP-IBC: CIIP (Combined Intra-Inter prediction) is a technology that generates prediction blocks by blending intra-prediction and inter-prediction. In CIIP-IBC mode, the final prediction block can be generated by blending a combination of intra- and inter-prediction.
[0348] - LIC-IBC: This is a mode that models local illumination variation between the template of the current block and the template of the IBC prediction block using LIC (Local Illumination Compensation) technology, and then compensates for the illumination variation to generate the final prediction block.
[0349] That is, block vector information of a neighboring block previously restored by the IBC mode used in combination with other tools, such as the example described above, for example, block vector information used to restore the neighboring block, may be included in the IBC merge candidate list. For example, block vector information of a neighboring block previously restored by GPM-IBC may be included in the IBC merge candidate list. In addition, block vector information of a neighboring block previously restored by SGPM-IBC may be included in the IBC merge candidate list. In addition, block vector information of a neighboring block previously restored by DIMD-IBC may be included in the IBC merge candidate list. In addition, block vector information of a neighboring block previously restored by TIMD-IBC may be included in the IBC merge candidate list. In addition, block vector information of a neighboring block previously restored by CIIP-IBC may be included in the IBC merge candidate list. In addition, block vector information of a neighboring block previously restored by LIC-IBC may be included in the IBC merge candidate list.
[0350] The above combinations are just examples, and it is also possible for IBC mode to be combined with other tools not mentioned above.
[0351] Additionally, the above example describes an IBC mode using block vectors as an example of a prediction mode applicable to the disclosed embodiment. Other prediction modes using block vectors may also be applicable to the disclosed embodiment. For example, an intra-TMP used in combination with other tools, such as GPM-IntaTMP, SGPM-IntaTMP, DIMD-IntaTMP, TIMD-IntaTMP, CIIP-IntaTMP, or LIC-IntaTMP, may also be applicable to the disclosed embodiment.
[0352] That is, the block vector information of the surrounding blocks restored by Intra-TMP combined with other tools, for example, the block vector information used to restore the surrounding blocks, may be included in the IBC merge candidate list of the current block. For example, the block vector information of the surrounding blocks previously restored by GPM-IntraTMP may be included in the IBC merge candidate list. In addition, the block vector information of the surrounding blocks previously restored by SGPM-IntraTMP may be included in the IBC merge candidate list. In addition, the block vector information of the surrounding blocks previously restored by DIMD-IntraTMP may be included in the IBC merge candidate list. In addition, the block vector information of the surrounding blocks previously restored by TIMD-IntraTMP may be included in the IBC merge candidate list. In addition, the block vector information of the surrounding blocks previously restored by CIIP-IntraTMP may be included in the IBC merge candidate list. In addition, the block vector information of the surrounding blocks previously restored by LIC-IntraTMP may be included in the IBC merge candidate list.
[0353] Additionally, two or more prediction modes using block vectors can be combined, and additional combinations with other tools are also possible. For example, combinations of GPM with IBC and IntaTMP, SGPM with IBC and IntaTMP, DIMD with IBC and IntaTMP, or TIMD with IBC and IntaTMP are possible. Here, the combination of Intra TMP and IBC can mean that the block vector of one of the two modes is used, or the block vector of both modes is used. For example, SGPM with IBC and IntaTMP can be combined as Intra / IBC, Intra / IntraTMP, IntraTMP / IBC, IBC / Intra, intraTMP / Intra, IBC / IntraTMP, IBC / IBC, IntraTMP / IntraTMP, etc.
[0354] That is, block vector information of a neighboring block restored by a combination of two or more prediction modes, for example, block vector information used to restore the neighboring block, may be included in the IBC merge candidate list of the current block. For example, block vector information of a neighboring block previously restored by GPM with IBC and IntaTMP may be included in the IBC merge candidate list. In addition, block vector information of a neighboring block previously restored by SGPM with IBC and IntaTMP may be included in the IBC merge candidate list. In addition, block vector information of a neighboring block previously restored by DIMD with IBC and IntaTMP may be included in the IBC merge candidate list. In addition, block vector information of a neighboring block previously restored by TIMD with IBC and IntaTMP may be included in the IBC merge candidate list. In addition, block vector information of a neighboring block previously restored by CIIP with IBC and IntaTMP may be included in the IBC merge candidate list.
[0355] In addition, in the candidate list configuration step (S1020), it is also possible for candidate block vector information of a history-based candidate list to be included in the IBC merge candidate list. For example, as in the above embodiments, if the prediction mode applied to the previously restored surrounding blocks is a mode in which block vectors are derived, such as IntraTMP, IBC, SGPM-IBC, DIMD-IBC, TIMD-IBC, CIIP-IBC, and LIC-IBC, the derived block vector information may be included in the history-based candidate list. Accordingly, if the IBC merge mode is applied to the current block, block vectors stored in the history-based candidate list may be included in the IBC merge candidate list.
[0356] Meanwhile, the size of the history-based candidate list can be predefined in the encoding / decoding device. For example, the size of the history-based candidate list can be defined as 6. In addition, the history-based candidate list can be initialized for each CTU. In addition, the history-based candidate list can be initialized whenever the CTU row position changes.
[0357] Additionally, in the candidate list configuration step (S1020), a block vector indicating a position indicated by a block vector of a surrounding block (reference block), i.e., a block vector indicating a position of a reference block of a surrounding block, can be derived, and block vector information indicating this block vector can be included in the IBC merge candidate list.
[0358] For example, as shown in the aforementioned Figure 20, a block vector BV indicating the location of the restored surrounding block (reference block) B1 of the current block 0,1 When , the block vector stored in B1 (= BV 1,2 ) can be derived so that the position indicated by BV can be represented based on the current block. That is, BV 0,2is BV 0,1 +BV 1,2 can be derived. In the same way, the block vector (=BV) stored in B2 2,3 ) can be derived so that the position indicated by BV can be represented based on the current block. That is, BV 0,3 Silver BV 0,1 +BV 1,2 +BV 2,3 can be derived. In addition, if the surrounding blocks (reference blocks) are restored by the IBC mode and / or IntraTMP mode combined with other tools, the stored block vectors can be utilized to derive block vector candidates. For example, the block vector information of the surrounding blocks restored by the IBC mode / IntraTMP combined with other tools, for example, the block vector information used to restore the surrounding blocks, can be utilized to derive block vector candidates.
[0359] In this way, when block vector information is stored in surrounding blocks, the process of accessing the location indicated by the block vector can be repeated, and the block vector derived through this can be utilized as a block vector candidate. Furthermore, in the process of deriving the block vector, the encoding device / decoding device can limit the number of repetitions to a predefined value. For example, the location of the reference block can be derived using only the block vector of the second reference block.
[0360] In addition, the reference block may store different block vectors in units of subblocks. The size of the subblock may be, for example, 4 by 4. Therefore, in the process of utilizing the block vectors stored in the reference block, the encoding device / decoding device may derive block vector information stored in a predefined order. For example, the block vector candidates may be derived by utilizing the block vectors stored in the order of center - left top - right top - left bottom - right bottom within the reference block. Alternatively, the encoding device / decoding device may derive the block vector candidates by utilizing only the block vectors stored in the predefined positions. For example, the block vector candidates may be derived by utilizing only the block vectors stored in the center position within the reference block.
[0361] Meanwhile, in the disclosed embodiment, the surrounding blocks used to construct the IBC merge candidate list of the current block may include blocks located adjacent / non-adjacent to the current block. The locations of the surrounding blocks have been described previously with reference to FIG. 21, and therefore, a redundant description is omitted here.
[0362] FIG. 24 is a diagram illustrating an example of an encoding method according to one embodiment of the present disclosure. The encoding method of FIG. 24 may be performed by an encoding device (200), and while each step is illustrated in a flowchart for convenience of explanation, the order in which each step is performed is not necessarily limited to the order in the flowchart. Depending on variations in the embodiment, the order may be performed differently, and some steps may be omitted or other steps may be added.
[0363] The encoding device (200) can determine a prediction mode for the current block (S1100). At this time, if the prediction mode for the current block is determined to be IBC merge mode, the encoding device (200) can define a search area (S1110), construct a candidate list (S1120), perform improvement and reordering (S1130), generate BVD information (S1140), and generate a prediction block (S1150). Then, image information can be encoded based on the generated prediction block (S1160).
[0364] The encoding method of FIG. 24 corresponds to the decoding method of FIG. 23 described above. The encoding device (200) and the decoding device (300) generate prediction blocks according to the same rules. Therefore, the description of the prediction method among the decoding methods described above with reference to FIG. 23 can be equally applied to the prediction method among the encoding methods according to the present example. For example, in the step (S1100) of determining the prediction mode for the current block, as described above with reference to FIG. 23, whether to use the IBC merge mode may be determined based on the size or shape of the block in order to reduce signaling overhead and improve compression performance. In addition, other descriptions can be equally applied as long as they do not conflict with the operation of the encoding device (200).
[0365] Additionally, in the step (S1110) of defining a search area based on whether the IBC merge mode is applied to the current block, the search area may vary depending on the size of the picture, the size of the CTU, the size of the current block, etc. Other descriptions may be equally applicable as long as they do not conflict with the operation of the encoding device (200).
[0366] In addition, in the step of constructing a candidate list (S1120), the encoding device (200) may construct a candidate list with block vectors of surrounding blocks that have been restored by a prediction mode in which block vectors are derived or used, such as IBC mode, intra TMP mode, IBC mode combined with other tools, intra TMP mode combined with other tools, or a combination of intra TMP mode and IBC mode. More specific details described with reference to FIG. 23 may also be applied to the present example.
[0367] Additionally, in the improvement and reordering step (S1130), the encoding device (200) can improve the block vector candidates of the merge candidate list using template matching. An error value can be calculated for the template of the current block and the template at the position indicated by each block vector candidate, and each block vector candidate can be improved based on the error value. In addition, the block vector candidates of the merge candidate list can be reordered based on the error value. Through this process, signaling overhead can be reduced when signaling block vector candidate information.
[0368] In addition, in the step of generating BVD information (S1140), BVD-related information described above with reference to FIG. 23 may be generated. That is, in order to solve the problem that prediction performance may be degraded when generating a prediction block using only block vector candidate information of the merge candidate list, BVD-related information may be generated and signaled.
[0369] In addition, in the step of generating a prediction block (S1150), the encoding device (200) can generate a prediction block for the current block using a block vector candidate and a BVD. At this time, the encoding device (200) can signal information about which block vector to use from the candidate list, which prediction block to generate, etc., and transmit it to the decoding device (300) in the form of a bitstream. Alternatively, it can be defined in advance in the encoding device / decoding device. In addition, other descriptions can be equally applied as long as they do not conflict with the operation of the encoding device (200).
[0370] In the step of encoding image information (S1160), the encoded image information may include prediction information, residual information, etc. Although the description has been omitted to reduce redundant explanation, it is of course possible to include procedures such as residual processing described above with reference to FIG. 5 in the encoding method according to the present example. The residual information generated by the residual processing may be included in the image information and encoded. The prediction information may refer to information necessary for predicting the current block, and specifically may include information about the prediction mode applied to the current block (or information indicating the prediction mode applied to the current block), index information, etc. The type of information to be encoded may vary depending on the prediction mode applied to the current block, and in the present example, IBC merge flag information may be included in the image information.
[0371] When image information encoded according to the above-described procedure is transmitted to a decoding device (300) in the form of a bitstream, the decoding device (300) can obtain image information from the transmitted bitstream and perform the above-described decoding method.
[0372] Example 3
[0373] According to one embodiment, a method of using block vectors to more effectively apply the IBC mode (IBC AMVP mode) is provided. That is, the embodiment describes a case where the prediction mode applied to the current block is the IBC mode (IBC AMVP mode). The IBC mode finds a previously reconstructed block similar to the current block through block matching to generate a prediction block. However, since only a very limited number of block vector candidates are used, there may be cases where the BVD value after block matching is large, resulting in increased signaling overhead. Therefore, the disclosed embodiment provides a method of improving compression performance by using a wider variety of block vector candidates.
[0374] FIG. 25 is a diagram illustrating an example of a decoding method according to one embodiment of the present disclosure. The decoding method of FIG. 25 can be performed by the aforementioned decoding device (300), and while each step is illustrated in a flowchart for convenience of explanation, the order in which each step is performed is not necessarily limited to the order in the flowchart. Depending on variations in the embodiment, the order may be performed differently, some steps may be omitted, or other steps may be added.
[0375] As previously described with reference to FIG. 4, the decoding device (300) can obtain image information from a bitstream. According to the present embodiment, the decoding device (300) can obtain information related to the IBC mode from the bitstream (S1200). For example, the information related to the IBC mode may be information indicating whether the IBC mode is applied to the current block, and this may be expressed as an IBC flag. Hereinafter, a specific embodiment will be described using an example in which the information related to the IBC mode is an IBC flag.
[0376] The decoding device (300) can decode the acquired IBC flag information and perform the IBC mode if the flag indicates that the IBC mode is applied to the current block (e.g., if the flag value indicates 1 or true).
[0377] At this time, whether to use IBC mode may be determined based on the size or shape of the block to reduce signaling overhead and improve compression performance. For example, in the case of non-square blocks, IBC mode may always be applied. As a specific example, when the block width is more than twice the height (i.e., width >= 2*height), IBC mode may always be applied. Alternatively, if the width*height of the current block is greater than 4096, IBC mode may be determined not to be applied (i.e., the value of the IBC merge flag may be 0 or false) or induced, in which case IBC mode is not performed.
[0378] The above example relates to a case where the application of IBC mode is determined based on the size or shape of the block, but it can also be extended or changed to other values and / or other conditions.
[0379] When the IBC mode is applied to the current block, the decoding device (300) can define a search area for performing the IBC mode (S1210). The search area may vary depending on the size of the picture, the size of the CTU, the size of the current block, etc. For example, when the width and height of the current block are W and H, respectively, the search area may be defined as a previously reconstructed area within 5*H above, 5*H below, 5*W to the left, and 5*W to the right based on the upper left of the current block. Alternatively, when the CTU size is C, the search area may be defined as an area within C*2 to the left, C / 4 to the right, C above, and C / 4 below based on the current block. Alternatively, the search area may be defined as a previously reconstructed area of row CTUs including the current block, and adjacent upper row CTUs of row CTUs including the current block. Alternatively, the search area may be defined as a combination of these. This is an example where the search area is defined based on block size, CTU size, etc., and it is also possible to define and use other search areas in advance in the encoding device / decoding device.
[0380] The decoding device (300) can perform block matching and candidate list construction (S1220). In this step, block matching can be performed on a given search area to find a reference block similar to the current block. Block vector candidates selected through block matching can be included in the candidate list. If previously restored surrounding blocks are restored by IBC mode or Intra TMP mode, block vector information stored in the corresponding surrounding blocks can be included in the candidate list. In order to utilize the block vector of the previously restored surrounding blocks, block vectors stored in adjacent / non-adjacent locations to the current block can be utilized. The adjacent / non-adjacent locations can be locations predefined by the encoding device / decoding device. For example, the adjacent locations can be Above-right, Bottom-left, and Above-left locations of the current block. In addition, a pairwise average candidate (a candidate obtained by averaging two different candidates by a predefined combination in the candidate list) can also be included in the candidate list. This is an example, and other candidates may be added to the candidate list based on values predefined by the encoding / decoding device. Meanwhile, the size of the candidate list may be a value predefined by the encoding / decoding device, for example, 16.
[0381] The decoding device (300) can refine and reorder block vector candidates in the candidate list (S1230). Specifically, the decoding device (300) can refine the block vector candidates in the candidate list using template matching (TM). An error value can be calculated between the template of the current block and the template at the position indicated by each block vector candidate, and each block vector candidate can be refined based on the error value. In addition, the block vector candidates in the candidate list can be reordered based on the error value. Through this process, signaling overhead can be reduced when signaling block vector candidate information. An error value can be calculated between the template at the position indicated by each block vector candidate in the candidate list and the template of the current block, and the candidate list can be sorted in order of the smallest template error value.
[0382] Error calculation methods that can be used include SAD (Sum of absolute difference), SATD (Sum of transformed difference), SSE (Sum of squared error), MR-SAD (Mean-removed sum of difference), MR-SSE (Mean-removed sum of squared error), and MR-SATD (Mean-removed sum of transformed difference).
[0383] Meanwhile, template shapes for template matching can be L-shape, above-shape, left-shape, above & left-shape, etc. as illustrated in FIG. 17 described above. Information about these shapes can be signaled, or can be defined in advance in an encoding device / decoding device and applied without signaling.
[0384] In addition, a variable template size can be used during the template matching process, which can also be signaled, or can be predefined in the encoding device / decoding device and applied without signaling. For example, the size of the template can be determined according to the size and shape of the block. For example, if the width*height of the current block is less than 64, a template consisting of one pixel line from the top and one pixel line from the left of the current block can be used. Otherwise, a template consisting of two pixel lines from the top and two pixel lines from the left can be used. In the case of non-square blocks, a template consisting of one pixel line from the top and one pixel line from the left can always be used. The above examples are examples in which the size of the template is determined based on the size and shape of the current block, and it is of course possible to determine it according to other values and / or other conditions.
[0385] The decoding device (300) can obtain block vector difference (BVD) information (S1240). If the block vector obtained through block matching is directly signaled, signaling overhead may occur. Therefore, in the disclosed embodiment, signaling overhead can be reduced by transmitting the difference (i.e., BVD) between the block vector obtained by performing block matching and the block vector candidates of the candidate list. For example, the BVD may be signaled information by coding direction information and the BVD size. In addition, template matching may be used to effectively signal the BVD. For example, the combination of direction information (x direction, y direction) of the BVD may be (+, +), (+, -),(-, +),(-, -), and the template error value may be obtained between the template of the current block and the template for each direction combination and sorted in order of the lowest error value, or one of the sorted direction information may be signaled. Alternatively, template matching may be used. Error calculation methods that can be used include SAD (Sum of absolute difference), SATD (Sum of transformed difference), SSE (Sum of squared error), MR-SAD (Mean-removed sum of difference), MR-SSE (Mean-removed sum of squared error), and MR-SATD (Mean-removed sum of transformed difference).
[0386] The decoding device (300) can generate a prediction block (S1250). The decoding device (300) can generate an IBC prediction block using a block vector candidate and a BVD value. At this time, information such as which candidate block vector to use from the candidate list, which prediction block to generate, and the BVD value can be obtained through signaling or defined in advance in the encoding device / decoding device. The IBC prediction block can be generated as follows:
[0387] - Single predictor: Generates a predicted block by copying the reference block at the location indicated by the block vector.
[0388] - CIIP-IBC predictor: CIIP (Combined Intra-Inter prediction) is one of the inter technologies that generates prediction blocks by blending intra and inter. CIIP-IBC generates the final prediction block after blending the combination of intra and inter-block prediction.
[0389] - Filter Model: A filter model is created based on the relationship between the template of the reference block indicated by the block vector and the template of the current block, and the filter model is applied to the reference block to generate a predicted block. For example, as described with reference to FIG. 17 above, filter coefficients can be derived so that the error between the current template and the reference template is minimized. A detailed description of the method for deriving the filter coefficients is as described above.
[0390] The IBC prediction block in the disclosed embodiment is not limited to the prediction block described above. The above-described prediction blocks are merely examples, and various other IBC prediction blocks may be generated. For example, a new prediction block may be generated by applying a LIC (Local Illumination Compensation) technique to the above-described prediction block. Alternatively, multiple IBC prediction blocks may be blended by a weight value predefined in an encoding device / decoding device. For example, the blending weight may be calculated based on a template matching error value of each reference block, calculated by deriving the weight based on a Wiener filter, or may be a weight value predefined in an encoding device / decoding device, or a preset for the weight may be commonly defined in an encoding device / decoding device and determined through explicit signaling.
[0391] Conventional IBC mode has limitations because it uses only a limited number of block vector candidates for prediction. The disclosed embodiment improves compression performance by using a wider variety of block vector candidates.
[0392] Therefore, in the disclosed embodiment, in the candidate list construction step (S1220), block vector information previously restored by the IBC mode may be included in the IBC candidate list. Here, the block vector information previously restored by the IBC mode may include information about a block vector used for prediction (restoration) of a previously restored neighboring block of the current block when the neighboring block is restored by applying the IBC mode. At this time, the neighboring block may include a block located within the search area of the current block. In addition, the IBC mode may include an IBC mode (=IBC AMVP mode) that transmits block vector information and block vector difference (BVD) information, and an IBC merge mode that transmits block vector information but not BVD information. However, even if the BVD value is not directly transmitted in the IBC merge mode, BVD-related information (distance and direction information) for supplementing the block vector may be transmitted.
[0393] Additionally, in the candidate list configuration step (S1220), previously restored intra TMP block vector information may be included in the IBC candidate list. Here, the previously restored intra TMP block vector information may include information about the block vector used for prediction (restoration) of previously restored neighboring blocks of the current block when the neighboring blocks are restored by applying the intra TMP mode. In this case, the neighboring blocks may include blocks located within the search area of the current block.
[0394] Meanwhile, the disclosed embodiments may be applied in combination with one or more prediction tools or prediction modes. Examples of possible combinations are listed below:
[0395] - GPM-IBC: GPM (Geometric partitioning mode) is one of the inter technologies that performs different predictions on each divided area through partitioning. Each divided area can perform predictions in various combinations such as inter / inter (e.g. prediction using different motion vectors), inter / intra, etc., and GPM-IBC is a mode that performs inter prediction and IBC prediction, or IBC prediction and IBC prediction for each divided area.
[0396] - SGPM-IBC: SGPM (Spatial GPM) is one of the intra techniques that performs different predictions on each region divided by partitioning. Each divided region can perform predictions of an intra / intra combination (e.g., prediction using different intra modes), and SGPM-IBC is a mode that performs intra prediction and IBC prediction, or IBC prediction and IBC prediction, for each divided region.
[0397] When applying the IBC mode together to GPM and SGPM, which perform different predictions in each region divided by partitioning, the IBC / IBC combination is possible in both GPM and SGPM, as in the example above. In this case, the encoding device / decoding device can define the combination in advance so that the combination of prediction modes does not overlap. For example, in an I-slice (or I-frame) that can only perform intra prediction, the SGPM mode can allow the IBC / IBC combination. In addition, in an I-slice (or I-frame), p-slice (or p-frame), and b-slice (or b-frame) that can perform intra prediction and inter prediction, the IBC / IBC combination can be allowed only for the GPM mode. Alternatively, the IBC / IBC combination can be allowed only for the SGPM mode. Here, the SGPM mode that allows combination with the IBC mode may mean the SGPM-IBC mode, and the GPM mode that allows combination with the IBC mode may mean the GPM-IBC mode.
[0398] In addition, if either the GPM-IBC mode or the SGPM-IBC mode, which allows combination with the IBC mode at the sequence parameter set (SPS) or picture parameter set (PPS) level, is turned off, the IBC / IBC combination can be allowed for the mode that is not turned off. For example, if GPM-IBC is off and SGPM-IBC is on, SGPM-IBC can allow the IBC / IBC combination. The above embodiments are examples that allow the IBC combination depending on the on / off status of the two modes, GPM-IBC and SGPM-IBC, and in addition, the combination of GPM-IBC and SGPM-IBC can be allowed so that they do not overlap between the encoding device / decoding device in another way.
[0399] - DIMD-IBC: This is a mode in which the DIMD mode and the IBC mode are blended together to generate the final prediction block, or a non-directional mode (Planar or DC mode) is replaced with the IBC mode and the final prediction block is generated after blending.
[0400] - TIMD-IBC: This is a mode in which the TIMD mode and the IBC mode are blended together to generate the final prediction block, or a non-directional mode (Planar or DC mode) is replaced with the IBC mode and the final prediction block is generated after blending.
[0401] - CIIP-IBC: CIIP (Combined Intra-Inter prediction) is a technology that generates prediction blocks by blending intra-prediction and inter-prediction. In CIIP-IBC mode, the final prediction block can be generated by blending a combination of intra- and inter-prediction.
[0402] - LIC-IBC: This is a mode that models local illumination variation between the template of the current block and the template of the IBC prediction block using LIC (Local Illumination Compensation) technology, and then compensates for the illumination variation to generate the final prediction block.
[0403] That is, block vector information of a neighboring block previously restored by the IBC mode used in combination with other tools, such as the example described above, for example, block vector information used to restore the neighboring block, may be included in the IBC candidate list. For example, block vector information of a neighboring block previously restored by GPM-IBC may be included in the IBC candidate list. In addition, block vector information of a neighboring block previously restored by SGPM-IBC may be included in the IBC candidate list. In addition, block vector information of a neighboring block previously restored by DIMD-IBC may be included in the IBC candidate list. In addition, block vector information of a neighboring block previously restored by TIMD-IBC may be included in the IBC candidate list. In addition, block vector information of a neighboring block previously restored by CIIP-IBC may be included in the IBC candidate list. In addition, block vector information of a neighboring block previously restored by LIC-IBC may be included in the IBC candidate list.
[0404] The above combinations are just examples, and it is also possible for IBC mode to be combined with other tools not mentioned above.
[0405] Additionally, the above example describes an IBC mode using block vectors as an example of a prediction mode applicable to the disclosed embodiment. Other prediction modes using block vectors may also be applicable to the disclosed embodiment. For example, an intra-TMP used in combination with other tools, such as GPM-IntaTMP, SGPM-IntaTMP, DIMD-IntaTMP, TIMD-IntaTMP, CIIP-IntaTMP, or LIC-IntaTMP, may also be applicable to the disclosed embodiment.
[0406] That is, the block vector information of the surrounding blocks restored by the Intra-TMP combined with other tools, for example, the block vector information used to restore the surrounding blocks, may be included in the IBC candidate list of the current block. For example, the block vector information of the surrounding blocks previously restored by the GPM-IntraTMP may be included in the IBC candidate list. In addition, the block vector information of the surrounding blocks previously restored by the SGPM-IntraTMP may be included in the IBC candidate list. In addition, the block vector information of the surrounding blocks previously restored by the DIMD-IntraTMP may be included in the IBC candidate list. In addition, the block vector information of the surrounding blocks previously restored by the TIMD-IntraTMP may be included in the IBC candidate list. In addition, the block vector information of the surrounding blocks previously restored by the CIIP-IntraTMP may be included in the IBC candidate list. In addition, the block vector information of the surrounding blocks previously restored by the LIC-IntraTMP may be included in the IBC candidate list.
[0407] Additionally, two or more prediction modes that use block vectors can be combined together. For example, combinations of GPM with IBC and IntaTMP, SGPM with IBC and IntaTMP, DIMD with IBC and IntaTMP, or TIMD with IBC and IntaTMP are possible. Here, the combination of Intra TMP and IBC can mean that the block vector of one of the two modes is used, or the block vectors of both modes are used. For example, SGPM with IBC and IntaTMP can be combined as Intra / IBC, Intra / IntraTMP, IntraTMP / IBC, IBC / Intra, intraTMP / Intra, IBC / IntraTMP, IBC / IBC, IntraTMP / IntraTMP, etc.
[0408] That is, block vector information of a neighboring block restored by a combination of two or more prediction modes, for example, block vector information used to restore the neighboring block, may be included in the IBC candidate list of the current block. For example, block vector information of a neighboring block previously restored by GPM with IBC and IntaTMP may be included in the IBC candidate list. In addition, block vector information of a neighboring block previously restored by SGPM with IBC and IntaTMP may be included in the IBC candidate list. In addition, block vector information of a neighboring block previously restored by DIMD with IBC and IntaTMP may be included in the IBC candidate list. In addition, block vector information of a neighboring block previously restored by TIMD with IBC and IntaTMP may be included in the IBC candidate list. In addition, block vector information of a neighboring block previously restored by CIIP with IBC and IntaTMP may be included in the IBC candidate list.
[0409] In addition, in the candidate list configuration step (S1220), it is also possible for candidate block vector information of a history-based candidate list to be included in the IBC candidate list. For example, as in the above embodiments, if the prediction mode applied to the previously restored surrounding blocks is a mode in which block vectors are derived, such as IntraTMP, IBC, SGPM-IBC, DIMD-IBC, TIMD-IBC, CIIP-IBC, and LIC-IBC, the derived block vector information may be included in the history-based candidate list. Accordingly, if the IBC mode is applied to the current block, block vectors stored in the history-based candidate list may be included in the IBC candidate list.
[0410] Meanwhile, the size of the history-based candidate list can be predefined in the encoding / decoding device. For example, the size of the history-based candidate list can be defined as 6. In addition, the history-based candidate list can be initialized for each CTU. In addition, the history-based candidate list can be initialized whenever the CTU row position changes.
[0411] Additionally, in the candidate list configuration step (S1220), a block vector indicating a position indicated by a block vector of a surrounding block (reference block), i.e., a block vector indicating a position of a reference block of a surrounding block, can be derived, and block vector information indicating this block vector can be included in the IBC candidate list.
[0412] For example, as shown in the aforementioned Figure 20, a block vector BV indicating the location of the restored surrounding block (reference block) B1 of the current block 0,1 When , the block vector stored in B1 (= BV 1,2 ) can be derived so that the position indicated by BV can be represented based on the current block. That is, BV 0,2 is BV 0,1 +BV 1,2 can be derived. In the same way, the block vector (=BV) stored in B2 2,3 ) can be derived so that the position indicated by BV can be represented based on the current block. That is, BV 0,3 Silver BV 0,1 +BV 1,2 +BV 2,3can be derived. In addition, if the surrounding blocks (reference blocks) are restored by the IBC mode and / or IntraTMP mode combined with other tools, the stored block vectors can be utilized to derive block vector candidates. For example, the block vector information of the surrounding blocks restored by the IBC mode / IntraTMP combined with other tools, for example, the block vector information used to restore the surrounding blocks, can be utilized to derive block vector candidates.
[0413] In this way, when block vector information is stored in surrounding blocks, the process of accessing the location indicated by the block vector can be repeated, and the block vector derived through this can be utilized as a block vector candidate. Furthermore, in the process of deriving the block vector, the encoding device / decoding device can limit the number of repetitions to a predefined value. For example, the location of the reference block can be derived using only the block vector of the second reference block.
[0414] In addition, the reference block may store different block vectors in units of subblocks. The size of the subblock may be, for example, 4 by 4. Therefore, in the process of utilizing the block vectors stored in the reference block, the encoding device / decoding device may derive block vector information stored in a predefined order. For example, the block vector candidates may be derived by utilizing the block vectors stored in the order of center - left top - right top - left bottom - right bottom within the reference block. Alternatively, the encoding device / decoding device may derive the block vector candidates by utilizing only the block vectors stored in the predefined positions. For example, the block vector candidates may be derived by utilizing only the block vectors stored in the center position within the reference block.
[0415] Meanwhile, in the disclosed embodiment, the surrounding blocks used to construct the IBC candidate list of the current block may include blocks located adjacent / non-adjacent to the current block. Since the locations of the surrounding blocks have been described previously with reference to FIG. 21, a redundant description is omitted here.
[0416] Fig. 26 is a diagram illustrating an example of an encoding method according to one embodiment of the present disclosure. The encoding method of Fig. 26 may be performed by an encoding device (200), and while each step is illustrated in a flowchart for convenience of explanation, the order in which each step is performed is not necessarily limited to the order in the flowchart. Depending on variations in the embodiment, the order may be performed differently, and some steps may be omitted or other steps may be added.
[0417] The encoding device (200) can determine a prediction mode for the current block (S1300). At this time, if the prediction mode for the current block is determined to be the IBC mode, the encoding device (200) can define a search area (S1310), configure block matching and a candidate list (S1320), perform improvement and reordering (S1330), generate BVD information (S1340), and generate a prediction block (S1350). Then, image information can be encoded based on the generated prediction block (S1360).
[0418] The encoding method of FIG. 26 corresponds to the decoding method of FIG. 25 described above. The encoding device (200) and the decoding device (300) generate prediction blocks according to the same rules. Therefore, the description of the prediction method among the descriptions of the decoding method described above with reference to FIG. 25 can be equally applied to the prediction method among the encoding methods according to the present example. For example, in the step (S1300) of determining the prediction mode for the current block, as described above with reference to FIG. 25, whether to use the IBC mode may be determined based on the size or shape of the block in order to reduce signaling overhead and improve compression performance. In addition, other descriptions can be equally applied as long as they do not conflict with the operation of the encoding device (200).
[0419] Additionally, in the step (S1310) of defining a search area based on whether the IBC mode is applied to the current block, the search area may vary depending on the size of the picture, the size of the CTU, the size of the current block, etc. Other descriptions may be equally applicable as long as they do not conflict with the operation of the encoding device (200).
[0420] In addition, in the step (S1320) of performing block matching and candidate list construction, the encoding device (200) may construct a candidate list with block vectors of surrounding blocks that have been restored by a prediction mode in which block vectors are derived, such as IBC mode, intra-TMP mode, IBC mode combined with other tools, intra-TMP mode combined with other tools, or a combination of intra-TMP mode and IBC mode. More specific details explained with reference to FIG. 25 may also be applied to the present example.
[0421] Additionally, in the improvement and reordering step (S1330), the encoding device (200) can improve the block vector candidates of the merge candidate list using template matching. An error value can be calculated for the template of the current block and the template at the position indicated by each block vector candidate, and each block vector candidate can be improved based on the error value. In addition, the block vector candidates of the merge candidate list can be reordered based on the error value. Through this process, signaling overhead can be reduced when signaling block vector candidate information.
[0422] Additionally, in the step of generating BVD information (S1340), the BVD information described above with reference to FIG. 25 may be generated.
[0423] In addition, in the step of generating a prediction block (S1350), the encoding device (200) can generate a prediction block for the current block using a block vector candidate and a BVD. At this time, the encoding device (200) can signal information about which block vector to use from the candidate list, which prediction block to generate, etc., and transmit it to the decoding device (300) in the form of a bitstream. Alternatively, it can be defined in advance in the encoding device / decoding device. In addition, other descriptions can be equally applied as long as they do not conflict with the operation of the encoding device (200).
[0424] In the step of encoding image information (S1360), the image information encoded may include information on prediction, residual information, etc. Although the description has been omitted to reduce redundant explanation, it is of course possible to include procedures such as residual processing described above with reference to FIG. 5 in the encoding method according to the example. The residual information generated by the residual processing may be included in the image information and encoded. The information on prediction may refer to information necessary for predicting the current block, and specifically may include information on the prediction mode applied to the current block (or information indicating the prediction mode applied to the current block), index information, etc. The type of information to be encoded may vary depending on the prediction mode applied to the current block, and in the example, IBC flag information may be included in the image information.
[0425] When image information encoded according to the above-described procedure is transmitted to a decoding device (300) in the form of a bitstream, the decoding device (300) can obtain image information from the transmitted bitstream and perform the above-described decoding method.
[0426] Although the embodiments have been described separately for convenience of explanation so far, a combination of two or more embodiments is possible, and changes required by the combination of embodiments may also be included in the scope of the disclosed invention or disclosed embodiments.
[0427] FIG. 27 is a diagram illustrating an example of a content streaming system to which an embodiment according to the present disclosure can be applied.
[0428] Referring to FIG. 27, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0429] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.
[0430] The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0431] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, in which case the control server controls commands / responses between each device within the content streaming system.
[0432] The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0433] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.
[0434] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0435] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.
[0436] Embodiments according to the present disclosure can be used to encode / decode images.
Claims
1. A step of obtaining information about a prediction mode from a bitstream; A step of determining a prediction mode applied to a current block as a prediction mode using a block vector based on information about the above prediction mode; A step of constructing a candidate list including at least one block vector candidate for the current block; A step of generating a prediction block based on at least one block vector candidate included in the candidate list; The steps for constructing the above candidate list are: Including adding block vector information of a previously restored surrounding block of the current block or block vector information derived from the block vector information of the previously restored surrounding block to the candidate list, The restored surrounding blocks of the current block above are: A decoding method, wherein the decoding method is restored by a combination of one or more prediction tools or one or more prediction modes and at least one of Intra Block Copy (IBC) mode or Intra Template Matching Prediction (TMP) mode.
2. In paragraph 1, Information about the above prediction mode is: A decoding method including information indicating whether the IBC mode is applied to the current block.
3. In paragraph 1, Information about the above prediction mode is: A decoding method comprising information indicating whether the IBC merge mode is applied to the current block.
4. In paragraph 1, Information about the above prediction mode is: A decoding method comprising information indicating whether intra TMP mode is applied to the current block.
5. In paragraph 1, The restored surrounding blocks of the current block above are: A decoding method restored by at least one prediction mode among GPM-IBC mode, SGPM-IBC mode, DIMD-IBC mode, TIMD-IBC mode, CIIP-IBC mode, or LIC-IBC mode.
6. In paragraph 1, The restored surrounding blocks of the current block above are: A decoding method restored by at least one prediction mode among GPM-IntaTMP, SGPM-IntaTMP, DIMD-IntaTMP, TIMD-IntaTMP, CIIP- IntaTMP or LIC-IntaTMP.
7. In paragraph 1, The restored surrounding blocks of the current block above are: A decoding method restored by at least one prediction mode among GPM with IBC and IntaTMP, SGPM with IBC and IntaTMP, DIMD with IBC and IntaTMP, or TIMD with IBC and IntaTMP.
8. A step of determining the prediction mode applied to the current block as a prediction mode using a block vector; A step of constructing a candidate list including block vector candidates for the current block; and A step of generating a prediction block based on at least one block vector candidate included in the candidate list; The steps for forming the above candidate list are: Including adding block vector information of a previously restored surrounding block of the current block or block vector information derived from the block vector information of the previously restored surrounding block to the candidate list, The restored surrounding blocks of the current block above are: An encoding method, wherein the encoding method is restored by a combination of one or more prediction tools or one or more prediction modes and at least one of Intra Block Copy (IBC) mode or Intra Template Matching Prediction (TMP) mode.
9. In paragraph 8, Information about the above prediction mode is: An encoding method including information indicating whether the IBC mode is applied to the current block.
10. In paragraph 8, Information about the above prediction mode is: An encoding method comprising information indicating whether IBC merge mode is applied to the current block.
11. In paragraph 8, Information about the above prediction mode is: An encoding method comprising information indicating whether intra TMP mode is applied to the current block.
12. In paragraph 8, The restored surrounding blocks of the current block above are: An encoding method restored by at least one prediction mode among GPM-IBC mode, SGPM-IBC mode, DIMD-IBC mode, TIMD-IBC mode, CIIP-IBC mode, or LIC-IBC mode.
13. In paragraph 8, The restored surrounding blocks of the current block above are: An encoding method restored by at least one prediction mode among GPM-IntaTMP, SGPM-IntaTMP, DIMD-IntaTMP, TIMD-IntaTMP, CIIP- IntaTMP or LIC-IntaTMP.
14. In paragraph 8, The restored surrounding blocks of the current block above are: An encoding method restored by at least one prediction mode among GPM with IBC and IntaTMP, SGPM with IBC and IntaTMP, DIMD with IBC and IntaTMP, or TIMD with IBC and IntaTMP.
15. In a computer-readable storage medium storing a bitstream generated by an encoding method, The above encoding method is, A step of determining the prediction mode applied to the current block as a prediction mode using a block vector; A step of constructing a candidate list including block vector candidates for the current block; and A step of generating a prediction block based on at least one block vector candidate included in the candidate list; The steps for constructing the above candidate list are: Including adding block vector information of a previously restored surrounding block of the current block or block vector information derived from the block vector information of the previously restored surrounding block to the candidate list, The restored surrounding blocks of the current block above are: A computer-readable storage medium restored by a combination of one or more prediction tools or one or more prediction modes and at least one of an Intra Block Copy (IBC) mode or an Intra Template Matching Prediction (TMP) mode.
16. In the method of transmitting data for video, Obtaining a bitstream for the image, wherein the bitstream is generated based on the steps of: determining a prediction mode applied to a current block as a prediction mode using a block vector; constructing a candidate list including block vector candidates for the current block; and generating a prediction block based on at least one block vector candidate included in the candidate list; and A step of transmitting the data including the bitstream; The steps for forming the above candidate list are: Including adding block vector information of a previously restored surrounding block of the current block or block vector information derived from the block vector information of the previously restored surrounding block to the candidate list, The restored surrounding blocks of the current block above are: A transmission method restored by a combination of one or more prediction tools or one or more prediction modes and at least one of Intra Block Copy (IBC) mode or Intra Template Matching Prediction (TMP) mode.
Citation Information
Patent Citations
Method, apparatus, and medium for video processing
WO2023133443A2
Method, device, and recording medium for image encoding / decoding
WO2024010377A1
Method, apparatus, and medium for video processing
WO2024012460A1