Video encoding / decoding method and apparatus, and recording medium storing bitstreams
The method improves video coding performance by determining reference blocks and applying filters to neighboring samples, enhancing accuracy and optimizing complexity for high-resolution images.
Patent Information
- Application Number
- JP2025541778
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-20
- Filing Date
- 2024-01-22
- Publication Date
- 2026-01-29
AI Technical Summary
Existing video compression techniques face challenges in achieving high coding performance for high-resolution and high-quality images due to inefficiencies in inter-prediction methods, particularly in determining accurate reference samples for block prediction.
A method and apparatus for video decoding that determines a reference block based on motion information, generates a predicted sample using filters applied to reference and neighboring samples, and reconstructs the current block with improved accuracy by considering extended surrounding areas and template types.
Enhances coding performance by deriving a highly accurate compensation model, optimizing computational complexity and performance through selective sampling and filter application, and improving prediction accuracy by adjusting templates based on block size and type.
Smart Images

Figure 2026503498000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a video encoding / decoding method and apparatus, and a recording medium storing a bitstream. [Background technology]
[0002] 2. Description of the Related Art In recent years, the demand for high-resolution, high-quality images such as high-definition (HD) images and ultra-high-definition (UHD) images has increased in various application fields, and as a result, highly efficient image compression techniques have been discussed.
[0003] There are various video compression techniques, such as inter-prediction techniques that predict pixel values contained in a current picture from pictures before or after the current picture, intra-prediction techniques that predict pixel values contained in a current picture using pixel information within the current picture, and entropy coding techniques that assign short codes to values that occur frequently and long codes to values that occur less frequently. Using these video compression techniques, video data can be effectively compressed and transmitted or stored. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure seeks to provide an inter-prediction method and apparatus.
[0005] The present disclosure seeks to provide a method and apparatus for compensating a reference sample. [Means for solving the problem]
[0006] The video decoding method and apparatus according to the present disclosure may determine a reference block for a current block based on motion information of the current block, generate a predicted sample for the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset, and reconstruct the current block based on the predicted sample.
[0007] In the video decoding method and apparatus according to the present disclosure, the peripheral samples may include (may comprise; may constitute; may construct; may be set; may encompass; may contain; may have) at least one of a left peripheral sample, a top peripheral sample, a left peripheral sample, or a bottom peripheral sample.
[0008] In the video decoding method and apparatus according to the present disclosure, the filter coefficients of the filter may be determined based on at least one sample belonging to a first surrounding area of the current block and at least one sample belonging to a second surrounding area of the reference block.
[0009] In the video decoding method and apparatus according to the present disclosure, the second surrounding area may include a template of the reference block and an area extended by N sample lines from the template of the reference block.
[0010] In the video decoding method and apparatus according to the present disclosure, the template may include at least one of a left peripheral region, a top peripheral region, a top left peripheral region, a top right peripheral region, or a bottom left peripheral region.
[0011] In the video decoding method and apparatus according to the present disclosure, at least one of the position or size of the template may be determined based on the size of the reference block.
[0012] In the video decoding method and apparatus according to the present disclosure, the template is determined to be one of a plurality of template candidates based on template type information signaled in a bitstream, and the template type information may indicate at least one of a position or a size of the template.
[0013] In the video decoding method and apparatus according to the present disclosure, the filter coefficients of the filter may be derived using a portion of samples belonging to a second surrounding region of the reference block, and the portion of samples may be identified based on a predetermined sampling rate and a predetermined offset.
[0014] In the video decoding method and apparatus according to the present disclosure, the filter may be determined to be one of a plurality of filters based on the value of the reference sample and a predetermined threshold.
[0015] In the video decoding method and apparatus according to the present disclosure, samples in the template of the reference block may be classified into a plurality of sample intervals based on the predetermined threshold, and the plurality of filters may be derived for the plurality of sample intervals, respectively.
[0016] In the video decoding method and apparatus according to the present disclosure, the prediction samples may be generated based on a plurality of filters including the filter, and the plurality of filters may be derived from a plurality of template candidates, respectively.
[0017] The video encoding method and apparatus according to the present disclosure may determine a reference block for inter-prediction of a current block, generate a predicted sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset, derive a residual sample of the current block based on the predicted sample, and encode the residual sample.
[0018] A computer-readable digital storage medium is provided having encoded video / image information stored thereon that enables a video decoding method to be performed by a decoding device according to the present disclosure.
[0019] A computer-readable digital storage medium is provided having stored thereon video / image information generated by the video encoding method of the present disclosure.
[0020] A method and apparatus for transmitting video / image information generated by a video encoding method according to the present disclosure is provided. [Effects of the Invention]
[0021] According to the present disclosure, coding performance can be improved by deriving a highly accurate compensation model and generating a predicted sample by referring to more surrounding samples in addition to the reference sample corresponding to the current sample.
[0022] According to the present disclosure, by selectively utilizing samples within a template based on a predetermined sampling rate and offset, an optimal tradeoff between computational complexity and performance can be derived.
[0023] According to the present disclosure, by adjusting the template position / size according to the block size, an optimal trade-off between computational complexity and performance can be derived.
[0024] According to the present disclosure, by selectively using one of a plurality of template candidates, a model can be defined that more accurately reflects the characteristics of the current block, thereby improving coding performance.
[0025] According to the present disclosure, prediction performance can be improved by defining multiple filter types and applying a more appropriate filter type to each reference sample.
[0026] According to the present disclosure, a final predicted block can be constructed by reflecting appropriate weights for various relational expressions between a current block and a reference block, thereby improving coding performance. [Brief explanation of the drawings]
[0027] [Figure 1] 1 illustrates a video / image coding system according to the present disclosure. [Figure 2] 1 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied, in which video / image signals are encoded. [Figure 3] 1 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied, in which video / image signals are decoded. [Figure 4] 1 is a diagram illustrating a video decoding method performed in an encoding device according to the present disclosure. [Figure 5] FIG. 10 is a diagram illustrating a schematic configuration of an inter-prediction unit (332) that performs the video decoding method according to the present disclosure. [Figure 6] 1 is a diagram illustrating a video encoding method performed in an encoding device according to the present disclosure. [Figure 7] FIG. 2 is a diagram showing a schematic configuration of an inter-prediction unit (220) that performs the video encoding method according to the present disclosure. [Figure 8] FIG. 1 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0028] While the present disclosure may be modified in various ways and may have various embodiments, specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit the present disclosure to the specific embodiments, and it should be understood that the present disclosure includes all modifications, equivalents, and alternatives within the spirit and technical scope of the present disclosure. In the description of each figure, similar reference numerals are used to refer to similar components.
[0029] Terms such as "first," "second," etc. may be used to describe various components, but these components should not be limited by such terms. These terms are used merely to distinguish one component from another. For example, a first component could be termed a second component, and similarly, a second component could be termed a first component, without departing from the scope of the present disclosure. The term "and / or" includes a combination of multiple associated listed items or any item of multiple associated listed items.
[0030] When a component is referred to as being "coupled" or "connected" to another component, it should be understood that the component may be directly coupled or connected to the other component, and that there may be additional components in between. On the other hand, when a component is referred to as being "directly coupled" or "directly connected" to another component, it should be understood that there are no additional components in between.
[0031] The terms used in this application are merely for the purpose of describing particular embodiments and are not intended to limit the present disclosure. The singular terms also include the plural terms unless the context clearly dictates otherwise. In this application, terms such as "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0032] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the versatile video coding (VVC) standard. The methods / embodiments disclosed herein may also be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation audio video coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0033] This specification presents various embodiments relating to video / image coding, and unless otherwise stated, the above embodiments may be performed in combination with each other.
[0034] In this specification, video may refer to a collection of a series of images over time. A picture generally refers to a unit representing an image at a specific time period, and a slice / tile is a unit constituting part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). One picture may be composed of one or more slices / tiles. A tile is a rectangular area composed of multiple CTUs in a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs having the same height as the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs having the same height as the picture and a width specified by the picture parameter set. CTUs within a tile may be arranged consecutively by CTU raster scanning, while tiles within a picture may be arranged consecutively by tile raster scanning. A slice may contain an integer number of complete tiles or an integer number of consecutive complete CTU rows within the tiles of a picture that may be contained exclusively in a single NAL unit, while a picture may be partitioned into two or more sub-pictures, which may be rectangular regions of one or more slices in a picture.
[0035] A picture element, pixel, or pel can refer to the smallest unit that makes up a picture (or an image). A "sample" can also be used as a term corresponding to a pixel. A sample can generally indicate a pixel or a pixel value, and may indicate only a pixel / pixel value of a luminance (luma) component, or may indicate only a pixel / pixel value of a chrominance (chroma) component.
[0036] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to that region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may sometimes be used interchangeably with terms such as block or area. In general, an MxN block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0037] As used herein, "A or B" can mean "A only," "B only," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B, or C" can mean "A only," "B only," "C only," or "any combination of A, B, and C."
[0038] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Thus, "A / B" can mean "A only," "B only," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0039] As used herein, "at least one of A and B" can mean "A only," "B only," or "both A and B." Furthermore, as used herein, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as being the same as "at least one of A and B."
[0040] Furthermore, in this specification, "at least one of A, B, and C" can mean "A only," "B only," "C only," or "any combination of A, B, and C." Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C."
[0041] Furthermore, parentheses used in this specification may mean "for example." Specifically, when "prediction (intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra prediction," and "intra prediction" may be suggested as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is displayed, "intra prediction" may be suggested as an example of "prediction."
[0042] In this specification, technical features that are individually described in the same drawing may be embodied individually or simultaneously.
[0043] FIG. 1 is a diagram illustrating a video / image coding system according to this disclosure.
[0044] Referring to FIG. 1, a video / image coding system may include a first device (a source device) and a second device (a receiving device).
[0045] A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.
[0046] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include a computer, tablet, smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced by a process in which the associated data is generated.
[0047] An encoding device may encode input video / images. The encoding device may perform a series of steps such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) may be output in the form of a bitstream.
[0048] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to a receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray (registered trademark: the same applies hereinafter), HDD, SSD, etc. The transmitting unit can include elements for generating a media file according to a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0049] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, which correspond to the operations of the encoding device.
[0050] The renderer can render the decoded video / image, and the rendered video / image can be displayed on a display unit.
[0051] FIG. 2 is a schematic block diagram of an encoding device to which the embodiments of the present disclosure can be applied, in which video / image signals are encoded.
[0052] Referring to FIG. 2, the encoding apparatus 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The above-described image divider 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filterer 260 may be configured by one or more hardware components (e.g., an encoding device chipset or processor) depending on the embodiment. In addition, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0053] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided into coding tree units (CTUs) or largest coding units (LCUs) according to a QTBTTT (Quad-tree, Binary-tree, Ternary-tree) structure.
[0054] For example, one coding unit may be divided into multiple coding units having deeper depths based on a quadtree structure, a binary tree structure, and / or a tertiary structure. In this case, for example, the quadtree structure may be applied first, and then the binary tree structure and / or the tertiary structure may be applied later. Alternatively, the binary tree structure may be applied before the quadtree structure. The coding procedure according to the present specification may be performed based on a final coding unit that is not further divided. In this case, based on coding efficiency according to video characteristics, the largest coding unit may be immediately used as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0055] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the final coding unit. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0056] The term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample may be used in terms corresponding to one picture (or image), pixel, or pel.
[0057] The encoding apparatus 200 may subtract a prediction signal (prediction block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input video signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, a unit in the encoding apparatus 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) may be referred to as a subtraction unit 231.
[0058] The prediction unit 220 may perform prediction on a current block (hereinafter referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit 220 may determine whether intra prediction or inter prediction is applied to the current block or CU. The prediction unit 220 may generate various information related to prediction, such as prediction mode information, as will be described later in the description of each prediction mode, and transmit the information related to prediction to the entropy encoding unit 240. The entropy encoding unit 240 may encode the information related to prediction and output it in the form of a bitstream.
[0059] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional modes may include at least one of DC mode and planar mode. The directional modes may include 33 directional modes or 65 directional modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0060] The inter prediction unit 221 may derive a prediction block for a current block based on a reference block (reference sample array) identified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring blocks may be the same or different. The temporal neighboring blocks may be called collocated reference blocks, collocated control units (colCUs), etc., and the reference picture including the temporal neighboring blocks may be called collocated pictures (colPic). For example, the inter prediction unit 221 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidates are used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of a skip mode or a merge mode, the inter prediction unit 221 may use motion information of neighboring blocks as motion information for the current block. In the case of the skip mode, unlike in the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of the surrounding block as a motion vector predictor and signaling the motion vector difference.
[0061] The prediction unit 220 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as a combined inter and intra prediction (CIIP) mode. The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for coding content images / videos, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, a sample value within the picture may be signaled based on information about a palette table and a palette index. The predicted signal generated by the prediction unit 220 may be used to generate a reconstructed signal or a residual signal.
[0062] The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), and a Conditionally Non-Linear Transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. The transform process may be applied to square pixel blocks of the same size, or to non-square blocks of variable sizes.
[0063] The quantization unit 233 quantizes the transform coefficients and transmits the quantized signal to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0064] The entropy encoding unit 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit 240 can encode information necessary for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients.
[0065] Encoded information (e.g., encoded video / video information) may be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) units. The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. In this specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / video information. The video / video information may be encoded using the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 may be transmitted to a transmitting unit (not shown) and / or stored to a storing unit (not shown) configured as an internal / external element of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.
[0066] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, the inverse quantization unit 234 and the inverse transform unit 235 may apply inverse quantization and inverse transform to the quantized transform coefficients to reconstruct a residual signal (residual block or residual sample). The adder 250 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to a prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the current block, such as when a skip mode is applied, the predicted block may be used as the reconstructed block. The adder 250 may be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, or may be used for inter prediction of the next picture after filtering, as described below. Meanwhile, luma mapping with chroma scaling (LMCS) may be applied during picture encoding and / or reconstruction.
[0067] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods may include deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc. The filtering unit 260 may generate various information related to filtering and transmit it to the entropy encoding unit 240. The information related to filtering may be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0068] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter prediction unit 221. This allows the encoding apparatus to avoid prediction mismatch between the encoding apparatus 200 and the decoding apparatus when inter prediction is applied, and also improves coding efficiency.
[0069] The DPB of the memory 270 may store the modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.
[0070] FIG. 3 is a schematic block diagram of a decoding device to which the embodiments of the present disclosure can be applied, in which video / image signals are decoded.
[0071] 3, the decoding device 300 may include an entropy decoding unit (entropy decoder 310), a residual processor (residual processor 320), a predictor (predictor 330), an adder (adder 340), a filter (filter 350), and a memory (memory 360). The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer (dequantizer 321) and an inverse transformer (inverse transformer 322).
[0072] The entropy decoding unit 310, residual processing unit 320, prediction unit 330, addition unit 340, and filtering unit 350 may be configured as a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. Also, the memory 360 may include a decoded picture buffer (DPB) and may be configured as a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0073] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the process by which the video / image information was processed by the encoding apparatus of FIG. 2. For example, the decoding apparatus 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using a processing unit applied by the encoding apparatus. Accordingly, the processing unit for decoding may be a coding unit, which may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a tertiary tree structure. One or more transform units may be derived from the coding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 may be played back by a playback device.
[0074] The decoding apparatus 300 may receive a signal output from the encoding apparatus of FIG. 2 in the form of a bitstream, and the received signal may be decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream and derive information (e.g., video / video information) necessary for video restoration (or picture restoration). The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. The decoding apparatus may decode pictures further based on the information on the parameter sets and / or the general constraint information. Signal / received information and / or syntax elements described later in this specification may be decoded by the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 may decode information in a bitstream based on a coding method such as Exponential-Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded, decoding information on neighboring and current blocks, or information on symbols / bins decoded in previous steps, predicts the occurrence probability of the bins based on the determined context model, and generates symbols corresponding to the values of each syntax element by performing arithmetic decoding of the bins. In this case, after determining the context model, the CABAC entropy decoding method may update the context model using information on the decoded symbols / bins for the context model of the next symbol / bin.Information related to prediction among the information decoded by the entropy decoding unit 310 may be provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to a residual processing unit 320. The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). In addition, information related to filtering among the information decoded by the entropy decoding unit 310 may be provided to a filtering unit 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding apparatus may be further configured as an internal / external element of the decoding apparatus 300, or the receiving unit may be a component of the entropy decoding unit 310.
[0075] Meanwhile, the decoding apparatus according to the present specification may be referred to as a video / image / picture decoding apparatus, and the decoding apparatus may be divided into an information decoding apparatus (video / image / picture information decoding apparatus) and a sample decoding apparatus (video / image / picture sample decoding apparatus). The information decoding apparatus may include the entropy decoding unit 310, and the sample decoding apparatus may include at least one of the inverse quantization unit 321, the inverse transform unit 322, the addition unit 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.
[0076] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding apparatus. The inverse quantization unit 321 may inverse quantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0077] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0078] The prediction unit 320 may perform prediction on a current block and generate a predicted block including prediction samples for the current block. The prediction unit 320 may determine whether intra prediction or inter prediction is applied to the current block based on the prediction information output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.
[0079] The prediction unit 320 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit 320 may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as a combined inter and intra prediction (CIIP) mode. The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode may be used for content video / movie coding, such as games, such as screen content coding (SCC). IBC basically performs prediction within a current picture, but may be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described herein. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information about a palette table and a palette index may be included in the video / picture information and signaled.
[0080] The intra prediction unit 331 may predict a current block by referring to samples in a current picture. The referenced samples may be located in the neighborhood of the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit 331 may also determine a prediction mode to be applied to the current block using prediction modes applied to neighboring blocks.
[0081] The inter prediction unit 332 may derive a prediction block for a current block based on a reference block (reference sample array) identified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.
[0082] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a prediction signal (prediction block, prediction sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when the skip mode is applied, the prediction block may be used as the reconstructed block.
[0083] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in a current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture. Meanwhile, luma mapping with chroma scaling (LMCS) may be applied during picture decoding.
[0084] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods may include deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0085] The (modified) reconstructed picture stored in the DPB of the memory 360 may be used as a reference picture in the inter predictor 332. The memory 360 may store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 260 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.
[0086] In this specification, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 may also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.
[0087] FIG. 4 is a diagram illustrating a video decoding method performed by an encoding device according to the present disclosure.
[0088] This disclosure defines a compensation model and proposes a method for applying the compensation model to inter-predicted blocks. The compensation method according to this disclosure can construct a linear model using a predetermined scaling factor (α) and offset (β) for compensation between a current block and a reference block. That is, a linear model that satisfies the following Equation 1 can be constructed using neighboring samples of the current block and neighboring samples of the reference block.
[0089]
number
[0090] The compensation model in the present disclosure may refer to, but is not limited to, an illumination compensation model between a current block and a reference block. A compensation method for an inter-predicted block will now be described in detail.
[0091] Referring to FIG. 4, a reference block for a current block can be determined based on motion information of the current block (S400).
[0092] A reference block according to the present disclosure may belong to a reference picture. Here, the reference picture may be a picture different from the current picture to which the current block belongs. For example, the reference picture may be a picture having a different output order (picture order count, POC) from the current picture. Or, the reference picture may be a picture having a different decoding order from the current picture. The reference picture may be a picture that has been coded / decoded before the current picture.
[0093] The motion information of the current block may include at least one of a motion vector, a reference picture index, or prediction direction information. The motion vector may identify the position of a reference block in a reference picture. For example, the motion vector may represent the difference between the position of the current block in the current picture and the position of the reference block in the reference picture. The reference picture index may identify one of a plurality of reference pictures belonging to a reference picture list for the current block. The prediction direction information may indicate at least one of whether the current block is predicted based on a reference picture list (List0) in the L0 direction or whether the current block is predicted based on a reference picture list (List1) in the L1 direction.
[0094] Referring to FIG. 4, a predicted sample of the current block can be generated by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset (S410).
[0095] The neighboring samples, which are input to the filter, may be one or more samples adjacent to the reference sample. The neighboring samples may include at least one of the left neighboring sample, the right neighboring sample, the top neighboring sample, the bottom neighboring sample, the top left neighboring sample, the bottom left neighboring sample, the top right neighboring sample, or the bottom right neighboring sample, which are adjacent to the reference sample. The offset, which is input to the filter, may be determined based on the bit depth of the current image to which the current block belongs. However, without being limited thereto, a sample that is not adjacent to the reference sample but is adjacent to at least one of the neighboring samples may also be used. Here, the shape of the area to which the filter is applied within the reference block may be square, non-square, cross-shaped, or diamond-shaped.
[0096] Depending on the position of the reference sample to which the filter is applied, the filter length, or the filter type, at least one of the neighboring samples may not belong to the reference block. For example, the neighboring samples that do not belong to the reference block may belong to one or more sample lines adjacent to at least one of the left, top, right, or bottom edges of the reference block. Even if the neighboring samples that are input to the filter do not belong to the reference block, the predicted sample may be generated using the neighboring samples as they are. Alternatively, compensation may be limited to only using samples that belong to the reference block. When a filter is applied to reference samples located at the boundary of the reference block, some neighboring samples may be unavailable. In this case, the unavailable neighboring samples may be replaced with the nearest available samples before the filter is applied.
[0097] The filter may be a filter for linear compensation. The filter may be a filter for a weighted sum of at least two of a reference sample belonging to the reference block, at least one surrounding sample adjacent to the reference sample, or a predetermined offset. The filter may be an M-tap convolution filter having a filter length (or the number of filter coefficients) of M, where M may be an integer greater than or equal to 2.
[0098] The filter coefficients of the filter may be derived based on one or more samples belonging to a template of the current block and one or more samples belonging to a template of the reference block. Hereinafter, for convenience of explanation, the template of the current block and the template of the reference block will be referred to as the current template and the reference template, respectively.
[0099] As an example, a filter with a filter length of 6 may be used. In this case, the reference samples may be compensated based on a 6-tap convolution filter as shown in Equation 2 below.
[0100]
number
[0101] According to Equation 2, the reference sample (pred) at the (x, y) position in the reference block x,y ) and the surrounding samples, the predicted sample (pred') at the (x,y) position in the current block is calculated. x,y ) can be generated, where the neighboring samples are the left neighboring samples (pred x-1,y ), right peripheral sample (pred x+1,y ), upper edge surrounding sample (pred x,y-1 ), and the bottom edge sample (pred x,y+1) where W and H are the width and height of the current block, respectively. x may be greater than or equal to 0 and less than W, and y may be greater than or equal to 0 and less than H.
[0102] In Equation 2, ci (i = 0...5) represents filter coefficients, which can be calculated using the least squares method using samples belonging to the current template and samples belonging to the reference template. B may be defined as (1 << (BitDepth - 1)), where BitDepth may represent the bit depth of the current image.
[0103] A vector (A) including samples and offsets of the reference template, a vector (w) of filter coefficients, and a sample vector (p) of the current template may be defined as follows:
[0104]
number
[0105] When Equation 3 is expressed as a normal equation, it is an equation consisting of the autocorrelation of A and the crosscorrelation between A and p, and may be expressed as Equation 4 below.
[0106]
number
[0107] Here, various methods may be used to derive the filter coefficients (w). For example, the filter coefficients may be derived by directly calculating the inverse matrix of the autocorrelation matrix. Alternatively, a linear system solution may be used. For example, the Cholesky method may be used to decompose the autocorrelation matrix into a lower triangular matrix and an upper triangular matrix, and then the filter coefficients may be derived by forward substitution and backward substitution. Alternatively, the LDL decomposition method may be used to decompose the autocorrelation matrix into a lower triangular matrix, a diagonal matrix, and an upper triangular matrix, and the filter coefficients may be derived. Alternatively, the filter coefficients may be derived using Gaussian elimination. Alternatively, if a model of Equation 2 is defined for compensation of an inter-predicted block and the filter coefficients of Equation 2 are derived using a specific linear system solution, this may be considered a method consistent with the present disclosure.
[0108] The current template according to the present disclosure may refer to a surrounding area that has already been restored before the current block. For example, the current template may include at least one of the top, left, top-left, bottom-left, or top-right surrounding areas adjacent to the current block. Similarly, the reference template is a region corresponding to the current template and may include at least one of the top, left, top-left, bottom-left, or top-right surrounding areas adjacent to the reference block.
[0109] The height of the top peripheral region, the top left peripheral region, and / or the top right peripheral region may be N. The width of the left peripheral region, the top left peripheral region, and / or the bottom left peripheral region may be N, where N may be an integer greater than or equal to 1. The width of the top peripheral region may be greater than or equal to the width of the current block (or the reference block), and the height of the left peripheral region may be greater than or equal to the height of the current block (or the reference block). However, without being limited thereto, the height of the top peripheral region may be different from the width of the left peripheral region.
[0110] The filter coefficients may be derived by further using a region extended by K sample lines from the reference template. Alternatively, the filter coefficients may be derived by further using a region extended by K sample lines from a region formed by the reference template and the reference block. Here, K may be an integer greater than or equal to 1. The extension direction may include at least one of an upper end direction, a left end direction, a lower end direction, and a right end direction. Whether the extended region is used may be adaptively determined based on at least one of a position of a reference sample to which a filter is applied, a filter length, or a filter shape.
[0111] Depending on at least one of the position of the reference sample to which the filter is applied, the filter length, and the filter type, neighboring samples that are input to the filter may not be available. In such a case, instead of using an extended template, the unavailable neighboring samples may be derived based on one or more samples in the reference template. The unavailable neighboring samples may be replaced with samples belonging to the reference template that are closest to the unavailable neighboring samples.
[0112] Specifically, a reference template may be defined as a set of N sample lines adjacent to the top and left sides of a reference block, where the size of the reference template (N) may be an integer greater than or equal to 1. Assume that the width and height of the reference block are W and H, respectively.
[0113] As an example, the reference template may be configured with a WxN top peripheral region, an NxN top left peripheral region, and an NxH left peripheral region. Alternatively, the reference template may be configured with a 2WxN top peripheral region, an NxN top left peripheral region, and an Nx2H left peripheral region. Alternatively, the reference template may be configured with a WxN top peripheral region, a WxN top right peripheral region, an NxN top left peripheral region, an NxH left peripheral region, and an NxH bottom left peripheral region.
[0114] Alternatively, the reference template may be defined as a set of N sample lines adjacent to the top edge of the reference block, where the size of the reference template (N) may be an integer greater than or equal to 1. Assume that the width and height of the reference block are W and H, respectively.
[0115] As an example, the reference template may be configured with a WxN top peripheral region, or may be configured with 2WxN top peripheral regions, or may be configured with a WxN top peripheral region and a WxN top right peripheral region.
[0116] Alternatively, the reference template may be defined as a set of N sample lines adjacent to the left side of the reference block, where the size of the reference template (N) may be an integer greater than or equal to 1. Assume that the width and height of the reference block are W and H, respectively.
[0117] As an example, the reference template may be configured with an NxH left peripheral region, or an Nx2H left peripheral region, or an NxH left peripheral region and an NxH bottom left peripheral region.
[0118] Alternatively, the reference template may be defined as a set of N sample lines adjacent to the top and left sides of the reference block, where the size of the reference template (N) may be an integer greater than or equal to 1. Assume that the width and height of the reference block are W and H, respectively.
[0119] As an example, the reference template may be configured with a WxN top peripheral region and an NxH left peripheral region, or a 2WxN top peripheral region and an Nx2H left peripheral region, or a WxN top peripheral region, a WxN top right peripheral region, an NxH left peripheral region, and an NxH bottom left peripheral region.
[0120] The above-described method for constructing a reference template may be equally applied to the method for constructing a current template, and a duplicated description will be omitted.
[0121] A plurality of template candidates may be defined for both the encoding device and the decoding device. Here, the plurality of template candidates may include at least one of the exemplary template configurations related to the template positions described above. For example, the plurality of template candidates may include a first template candidate having a WxN top peripheral region, an NxN top-left peripheral region, and an NxH left peripheral region, and a second template candidate having a 2WxN top peripheral region, an NxN top-left peripheral region, and an Nx2H left peripheral region. Alternatively, the plurality of template candidates may include a first template candidate having a WxN top peripheral region and an NxH left peripheral region, a second template candidate having a WxN top peripheral region, and a third template candidate having an NxH left peripheral region.
[0122] Alternatively, the plurality of template candidates may be configured according to any one of the template configuration examples related to the template position described above, but may have different sizes (N). For example, the plurality of template candidates may include a first template candidate including a left peripheral region of N1xH and a second template candidate including a left peripheral region of N2xH, where N1 and N2 may be different integers.
[0123] Alternatively, the plurality of template candidates may be configured using two or more of the exemplary template configurations related to the template positions described above, and one of the template candidates may have a different size (N) from the others. For example, the plurality of template candidates may include a first template candidate configured with a top peripheral region of WxN1 and a left peripheral region of N1xH, a second template candidate configured with a top peripheral region of 2WxN2, and a third template candidate configured with a left peripheral region of N2x2H.
[0124] Any one of the predefined plurality of template candidates may be selectively used. Template type information may be used to select any one of the plurality of template candidates. The template type information may include at least one of information indicating a template position or information indicating a template size. In this case, the template type information may be defined as a single index and may indicate the template position and size. Alternatively, the information indicating the template position and the information indicating the template size may each be defined as separate indexes.
[0125] At least one of the template type information may be signaled by a bitstream. The template type information may be defined in a region unit such as a video sequence, a picture, a slice, etc., and may be signaled at a higher level such as a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), or a slice header (SH). Alternatively, the template type information may be defined in a block unit such as a coding tree unit, a coding unit, or a sub-block, and may be signaled at a lower level such as a coding tree unit syntax or a coding unit syntax. The information indicating the template location and the information indicating the template size may be signaled in the same region unit or block unit. The information indicating the template location and the information indicating the template size may be signaled separately in region unit or block unit. One of the information indicating the template location and the information indicating the template size may be signaled in one of the region units described above, and the other may be signaled in one of the block units described above.
[0126] Alternatively, the template type information (e.g., information indicating the location of the template) may be derived based on the size of the reference block (or the current block), i.e., one of multiple template candidates may be selected based on the size of the reference block (or the current block).
[0127] For example, if the width of the reference block is greater than the height, a template candidate including a top-edge peripheral region may be used. Conversely, if the width of the reference block is less than the height, a template candidate including a left-edge peripheral region may be used. If this is not the case (i.e., if the width and height of the reference block are the same), a template candidate including a left-edge peripheral region, a top-edge peripheral region, and a top-left-top peripheral region may be used, or a template candidate including a top-edge peripheral region and a left-edge peripheral region but not a top-left-top peripheral region may be used.
[0128] Alternatively, template type information (e.g., information indicating the size of the template) may be derived based on the size of the reference block (or the current block). That is, one of multiple template candidates may be selected based on the size of the reference block (or the current block). The size (N) of the template may be variably determined based on the size of the reference block (or the current block).
[0129] For example, if the number of samples belonging to the reference block (or the current block) is less than or equal to 64, a template candidate having a size of N1 may be selected, and if not, a template candidate having a size of N2 may be selected. Alternatively, if the number of samples belonging to the reference block (or the current block) is less than or equal to 64, the template size may be determined to be N1, and if not, the template size may be determined to be N2. Here, N1 may be an integer less than N2. For example, N1 and N2 may be 4 and 6, respectively, but are not limited to this.
[0130] As described above, by selectively using one of multiple template candidates, it is possible to define a model that more accurately reflects the characteristics of the current block, thereby improving coding performance. Also, by adjusting the position / size of the template according to the block size, it is possible to derive an optimal tradeoff between computational complexity and performance.
[0131] All samples belonging to the current and / or reference template may be used to derive the filter coefficients, or only a selected subset of the samples belonging to the current and / or reference template may be used to derive the filter coefficients, i.e., rather than cycling through all samples in the template, a compensation model may be derived based on only a subset of the samples.
[0132] Specifically, assume that a sampling rate of i and an offset of a are applied to the x-axis, and a sampling rate of j and an offset of b are applied to the y-axis. The sample at the (x, y) position in the template is denoted by T x,y Then, the number of samples required to derive the filter coefficients is T i*x+a,j*y+b Here, the offset may be any one of integers including 0. a and b may be different integers or may be the same integer.
[0133] For example, the filter coefficients can be derived using only samples belonging to at least one even row and at least one even column in the template. In this case, the samples for deriving the filter coefficients are T 2x,2y The filter coefficients may be derived using only samples belonging to at least one odd row and at least one odd column in the template. The filter coefficients may be derived using only samples belonging to at least one odd row or at least one odd column in the template. The filter coefficients may be derived using only samples belonging to at least one even row or at least one even column in the template.
[0134] When samples in a template are selectively used based on a predetermined sampling rate and offset, an optimal trade-off between computational complexity and performance can be derived. Also, by referencing more neighboring samples in addition to a reference sample corresponding to a current sample, a highly accurate compensation model can be derived and a predicted sample can be generated, thereby improving coding performance.
[0135] For the current block, two or more compensation models may be derived / defined instead of a single compensation model. One of the multiple compensation models may be selected according to a specific condition. In this case, one of the multiple compensation models may be selected for each reference sample in the reference block. Alternatively, one of the multiple compensation models may be selected for each sample interval of one or more reference samples. Herein, the specific condition may be a magnitude relationship between a specific reference sample in the reference block and a predetermined threshold. Hereinafter, the compensation model will be referred to as a filter.
[0136] For example, a filter may be derived by cycling through all or some samples of a reference template. In this case, the samples in the reference template may be classified into a plurality of sample intervals based on whether the values of the samples in the reference template are smaller than a predetermined threshold. A filter may be derived for each sample interval.
[0137] A sample interval to which a reference sample belongs may be determined based on whether the value of the reference sample in the reference block is smaller than a predetermined threshold. A filter corresponding to the determined sample interval may be applied to the reference sample. For example, if (L-1) thresholds are derived / defined, L filter types may be derived / defined for the current block. If the value of the reference sample in the reference block is greater than or equal to a first threshold, a first filter may be used. If the value of the reference sample in the reference block is less than the first threshold and greater than or equal to a second threshold, a second filter may be used. If the value of the reference sample in the reference block is less than the (L-2)th threshold and greater than or equal to the (L-1)th threshold, an (L-1)th filter may be used. In other cases, an Lth filter may be used. L may be an integer greater than or equal to 1.
[0138] The threshold value according to the present disclosure may be determined based on at least one of the average value of the samples in the reference block, the minimum value, the maximum value, or the median value of the samples in the reference block, the bit depth of the current image, or information signaled to identify the threshold value.
[0139] According to the present disclosure, prediction performance can be improved by defining multiple filter types and applying a more appropriate filter type to each reference sample.
[0140] Alternatively, two or more filters may be derived / defined for the current block, and two or more predicted samples may be generated by applying the two or more filters to the reference samples, respectively, and one final predicted sample may be generated by weighting the two or more predicted samples.
[0141] For example, as in the method described above, two or more filters may be derived / defined based on the values of samples in the reference template and a predetermined threshold. Two or more filtered current templates may be generated by applying the two or more filters to the samples of the current template, respectively. A cost between the filtered current template and the reference template may be calculated. Here, the cost may be a value measuring the difference between the filtered current template and the reference template. The cost may be calculated based on the sum of absolute difference (SAD), sum of absolute transformed difference (SATD), or sum of squared difference (SSE).
[0142] Two or more predicted samples may be generated by applying two or more filters to reference samples in a reference block, respectively. A final predicted sample may be generated by weighting the two or more predicted samples. In this case, weights for the weighted sum may be determined based on the calculated costs. The sum of the costs calculated for the two or more filters may be the denominator of the weight, and a specific cost may be the numerator of the weight. In this case, a relatively large weight may be applied to a predicted sample generated based on a filter having a relatively small cost.
[0143] For example, a final predicted sample based on a weighted sum of two predicted samples may be generated as shown in Equation 5 below.
[0144]
number
[0145] In Equation 5, pred x,y f0(refBlock x,y f1(refBlock) may refer to a first predicted sample generated by applying a first filter to a reference sample at the (x, y) position in the reference block. x,y ) may refer to a second predicted sample generated by applying a second filter to a reference sample at the (x, y) position in the reference block.
[0146] According to Equation 5, if Cost0 is smaller than Cost1, a relatively large weight may be applied to the first predicted sample generated by applying the first filter. That is, the ratio of weights applied to the first predicted sample and the second predicted sample may be Cost1:Cost0. Conversely, if Cost0 is larger than or equal to Cost1, a relatively large weight may be applied to the second predicted sample generated by applying the second filter. That is, the ratio of weights applied to the first predicted sample and the second predicted sample may be Cost1:Cost0.
[0147] According to the above-described method, the final predicted block can be constructed by reflecting appropriate weights for various relationships between the current block and the reference block, thereby improving coding performance.
[0148] Alternatively, two or more filters may be derived / defined for the current coding block instead of a single filter. The two or more filters may be derived based on two or more template candidates (or samples at specific positions within multiple template candidates). The two or more template candidates may be defined according to the above-described template candidate configuration method, and redundant description will be omitted here. Two or more predicted samples may be generated by applying the two or more filters to reference samples of the reference block, respectively. One final predicted sample may be generated by weighting the two or more predicted samples.
[0149] For example, the two or more filters for the current block may include at least one of a first filter derived from the left peripheral region, a second filter derived from the top peripheral region, a third filter derived from the left and top peripheral regions, or a fourth filter derived from the left, top, and top-left peripheral regions. Alternatively, two or more filters may be derived using different sampling rates, offsets, and / or thresholds within the same template candidate.
[0150] For example, a first filter may be derived from a first template candidate including the left peripheral region, and a second filter may be derived from a second template candidate including the top peripheral region. Two filtered current templates may be generated by applying the two filters to samples of the current template, respectively. A cost between the filtered current template and the reference template may be calculated. Here, the cost may be a value measuring the difference between the filtered current template and the reference template. The cost may be calculated based on the sum of absolute differences (SAD), sum of absolute transformed differences (SATD), or sum of squared differences (SSE).
[0151] A first predicted sample may be generated by applying the first filter to a reference sample in a reference block, and a second predicted sample may be generated by applying the second filter to the reference sample. A final predicted sample may be generated by performing a weighted sum of the first predicted sample and the second predicted sample. Weights for the weighted sum may be determined based on the calculated costs. The sum of the costs calculated for the two filters may be the denominator of the weight, and a specific cost may be the numerator of the weight. In this case, a relatively large weight may be applied to a predicted sample generated based on a filter having a relatively small cost.
[0152] For example, a final predicted sample based on a weighted sum of two predicted samples may be generated as shown in Equation 6 below.
[0153]
number
[0154] In Equation 6, pred x,y can mean the final predicted sample at the (x,y) position in the current block. Leftmay represent a first cost between a current template and a reference template to which a first filter among a plurality of filters is applied. Here, the first filter may be derived from a second template candidate including the left peripheral region. Above f may represent a second cost between the current template and the reference template to which a second filter from among the plurality of filters is applied. Here, the second filter may be derived from the first template candidate including the upper edge surrounding region. Left (refBlock x,y ) may refer to a first predicted sample generated by applying a first filter to a reference sample at the (x, y) position in the reference block. Above (refBlock x,y ) may refer to a second predicted sample generated by applying a second filter to a reference sample at the (x, y) position in the reference block.
[0155] According to Equation 6, Cost Left Cost Above If the first predicted sample is smaller than the first predicted sample, a relatively large weight may be applied to the first predicted sample generated by applying the first filter. That is, the ratio of the weights applied to the first predicted sample and the second predicted sample is determined by the Cost Above :Cost Left Conversely, Cost Left Cost Above If the second predicted sample is greater than or equal to the first predicted sample, a relatively large weight may be applied to the second predicted sample generated by applying the second filter. That is, the ratio of the weights applied to the first predicted sample and the second predicted sample is determined by the Cost Above :Cost Left It may be.
[0156] According to the above-described method, the final predicted block can be constructed by reflecting appropriate weights for various relationships between the current block and the reference block, thereby improving coding performance.
[0157] Referring to FIG. 4, the current block can be reconstructed based on the predicted samples of the current block (S420).
[0158] A current sample of the current block may be reconstructed based on a predicted sample of the current block and a residual sample. The residual sample may be derived by deriving transform coefficients based on residual information acquired from a bitstream and performing inverse quantization / inverse transform on the derive transform coefficients.
[0159] Meanwhile, the above-described compensation method may be applied in the same or similar manner to linear model-based intra prediction between a luminance component block and a chrominance component block. In this case, the current block and the reference block in the above-described compensation method may be understood to be replaced with the chrominance component block and the luminance component block of the current block, respectively. Also, the reference block in the above-described compensation method may be understood to belong to the same current picture as the current block.
[0160] FIG. 5 is a diagram showing a schematic configuration of the inter prediction unit 332 that performs the video decoding method according to the present disclosure.
[0161] Referring to FIG. 5, the inter prediction unit 332 may include a reference block determination unit 500 and a reference sample compensation unit 510.
[0162] The reference block determination unit 500 can determine a reference block for the current block based on the motion information of the current block. The method for determining the reference block is the same as that described with reference to FIG.
[0163] The reference sample compensation unit 510 may generate a predicted sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset.
[0164] The filter may have a predetermined filter length and may be a filter for linear compensation, or may be a filter for a weighted sum of at least two of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset.
[0165] The filter coefficients of the filter may be derived based on one or more samples belonging to a template of the current block and one or more samples belonging to a template of the reference block. The template and the method for deriving the filter coefficients according to the present disclosure are as described with reference to FIG.
[0166] Also, multiple filters may be derived / defined for the current block, and the method of generating predicted samples based on the multiple filters is the same as that described with reference to FIG.
[0167] The predicted samples output from the reference sample compensation unit 510 may be input to the adder 340 of the decoding apparatus 300 and used to reconstruct the current block.
[0168] FIG. 6 is a diagram illustrating a video encoding method performed by an encoding device according to the present disclosure.
[0169] Referring to FIG. 6, a reference block for inter-prediction of a current block can be determined (S600).
[0170] A reference block according to the present disclosure may belong to a reference picture. Here, the reference picture may be a picture different from the current picture to which the current block belongs. For example, the reference picture may be a picture having a different output order (picture order count, POC) from the current picture. Or, the reference picture may be a picture having a different decoding order from the current picture. The reference picture may be a picture that has been coded / decoded before the current picture.
[0171] Referring to FIG. 6, a predicted sample of the current block can be generated by applying a filter to at least one of a reference sample belonging to the reference block, at least one surrounding sample adjacent to the reference sample, or a predetermined offset (S610).
[0172] The filter may have a predetermined filter length and may be a filter for linear compensation, or may be a filter for a weighted sum of at least two of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset.
[0173] The filter coefficients of the filter may be derived based on one or more samples belonging to a template of the current block and one or more samples belonging to a template of the reference block. The template and the method for deriving the filter coefficients according to the present disclosure are as described with reference to FIG.
[0174] Also, multiple filters may be derived / defined for the current block, and the method of generating predicted samples based on the multiple filters is the same as that described with reference to FIG.
[0175] Referring to FIG. 6, residual samples of the current block may be derived based on the predicted samples (S620), and the residual samples may be encoded to generate a bitstream (S630).
[0176] FIG. 7 is a diagram showing a schematic configuration of the inter prediction unit 221 that performs the video encoding method according to the present disclosure.
[0177] Referring to FIG. 7, the inter predictor 221 may include a reference block determiner 700 and a reference sample compensator 710.
[0178] The reference block determination unit 700 may determine a reference block for inter-prediction of the current block, as described with reference to FIG.
[0179] The reference sample compensation unit 710 can generate a predicted sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one surrounding sample adjacent to the reference sample, or a predetermined offset.
[0180] The filter may have a predetermined filter length and may be a filter for linear compensation, or may be a filter for a weighted sum of at least two of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset.
[0181] The filter coefficients of the filter may be derived based on one or more samples belonging to a template of the current block and one or more samples belonging to a template of the reference block. The template and the method for deriving the filter coefficients according to the present disclosure are as described with reference to FIG.
[0182] Also, multiple filters may be derived / defined for the current block, and the method of generating predicted samples based on the multiple filters is the same as that described with reference to FIG.
[0183] The prediction samples output from the reference sample compensation unit 710 may be input to the residual processing unit 230 of the encoding apparatus 200. The residual processing unit 230 may derive residual samples based on the prediction samples, perform transformation / quantization on the residual samples to derive transform coefficients, and generate residual information related to the derive transform coefficients. The residual information may be input to the entropy encoding unit 240, which may encode the residual information to generate a bitstream.
[0184] In the above-described embodiments, the method is described based on a flowchart with a series of steps or blocks, but the embodiment is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included, or one or more steps of the flowchart may be deleted without affecting the scope of the embodiments of this document.
[0185] The methods according to the embodiments of the present document described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present document may be included in a device that performs video processing, such as a TV, a computer, a smartphone, a set-top box, or a display device.
[0186] When embodiments in this document are embodied as software, the methods described above may be embodied as modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor by various known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be embodied and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the figures may be embodied and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for the implementation may be stored on a digital storage medium.
[0187] In addition, the decoding device and encoding device to which the embodiments of the present specification are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a custom video (VoD) service providing device, an over-the-top (OTT) video (over-the-top) device, an internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, a video telephone video device, a vehicle terminal (e.g., a vehicle terminal (including an autonomous vehicle), an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process video signals or data signals. For example, over-the-top (OTT) video (over-the-top) video devices may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0188] In addition, a processing method to which the embodiments of the present specification are applied may be produced in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of the present specification may also be stored in a computer-readable recording medium. The computer-readable recording medium may include any type of storage device or distributed storage device in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray Disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium may also include media embodied in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0189] Furthermore, the embodiments of the present specification may be embodied as a computer program product using program code, which may be executed by a computer according to the embodiments of the present specification. The program code may be stored on a computer-readable carrier.
[0190] FIG. 8 illustrates an example of a content streaming system to which the embodiments of the present disclosure can be applied.
[0191] Referring to FIG. 8, a content streaming system to which the embodiments of the present specification are applied may broadly include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0192] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0193] The bitstream may be generated by an encoding method or a bitstream generation method to which the embodiments of this specification are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0194] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0195] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0196] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, and head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signage.
[0197] Each server in the content streaming system may be operated as a distributed server, in which case data received by each server may be processed in a distributed manner.
[0198] The claims described herein may be combined in various ways. For example, technical features of method claims herein may be combined and embodied as an apparatus, and technical features of apparatus claims herein may be combined and embodied as a method. Furthermore, technical features of method claims herein and technical features of apparatus claims herein may be combined and embodied as an apparatus, and technical features of method claims herein and technical features of apparatus claims herein may be combined and embodied as a method.
[0199] [Claims at the time of international application] [Claim 1] A video decoding method, comprising: determining a reference block for the current block based on motion information of the current block; generating a predicted sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset; and reconstructing the current block based on the predicted samples. [Claim 2] The video decoding method of claim 1 , wherein the peripheral samples include at least one of left peripheral samples, top peripheral samples, left peripheral samples, or bottom peripheral samples. [Claim 3] 2. The video decoding method of claim 1, wherein filter coefficients of the filter are determined based on at least one sample belonging to a first peripheral region of the current block and at least one sample belonging to a second peripheral region of the reference block. [Claim 4] 4. The image decoding method of claim 3, wherein the second surrounding area includes the template of the reference block and an area extended by N sample lines from the template of the reference block. [Claim 5] The video decoding method of claim 4, wherein the template includes at least one of a left peripheral region, a top peripheral region, a top left peripheral region, a top right peripheral region, or a bottom left peripheral region. [Claim 6] The video decoding method of claim 4 , wherein at least one of a position or a size of the template is determined based on a size of the reference block. [Claim 7] The template is determined to be one of a plurality of template candidates based on template type information signaled in a bitstream; The video decoding method of claim 4, wherein the template type information indicates at least one of a position or a size of the template. [Claim 8] filter coefficients of the filter are derived using some samples belonging to a second neighboring region of the reference block; 4. The video decoding method of claim 3, wherein the subset of samples is identified based on a predetermined sampling rate and a predetermined offset. [Claim 9] The video decoding method of claim 1, wherein the filter is determined to be one of a plurality of filters based on the value of the reference sample and a predetermined threshold. [Claim 10] The samples in the template of the reference block are classified into a plurality of sample sections based on the predetermined threshold; 10. The video decoding method of claim 9, wherein the plurality of filters are derived for the plurality of sample intervals, respectively. [Claim 11] the predicted samples are generated based on a plurality of filters including the filter; The video decoding method of claim 1 , wherein the plurality of filters are derived from a plurality of candidate templates, respectively. [Claim 12] 1. A video encoding method, comprising: determining a reference block for inter prediction of a current block; generating a predicted sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset; deriving residual samples of the current block based on the predicted samples; encoding the residual samples. [Claim 13] A computer-readable recording medium, comprising: A computer-readable recording medium storing a bitstream generated by the video encoding method of claim 12. [Claim 14] 1. A data transmission method, comprising: obtaining a bitstream for video information; The bitstream comprises: determining a reference block of the current block for inter prediction of the current block; generating a predicted sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset; deriving residual samples of the current block based on the predicted samples; generated by encoding the residual samples, transmitting data including the bitstream.
Claims
1. A video decoding method, comprising: determining a reference block for the current block based on motion information of the current block; generating a predicted sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset; and reconstructing the current block based on the predicted samples.
2. The video decoding method of claim 1 , wherein the peripheral samples include at least one of left peripheral samples, top peripheral samples, left peripheral samples, and bottom peripheral samples.
3. 2. The image decoding method of claim 1, wherein filter coefficients of the filter are determined based on at least one sample belonging to a first peripheral region of the current block and at least one sample belonging to a second peripheral region of the reference block.
4. The image decoding method of claim 3 , wherein the second surrounding area includes a template of the reference block and an area extended by N sample lines from the template of the reference block.
5. The image decoding method of claim 4 , wherein the template includes at least one of a left peripheral region, a top peripheral region, a top left peripheral region, a top right peripheral region, or a bottom left peripheral region.
6. The video decoding method of claim 4 , wherein at least one of the position or size of the template is determined based on the size of the reference block.
7. The template is determined to be one of a plurality of template candidates based on template type information signaled in a bitstream; The video decoding method of claim 4 , wherein the template type information indicates at least one of a position or a size of the template.
8. filter coefficients of the filter are derived using some samples belonging to a second surrounding region of the reference block; The video decoding method of claim 3 , wherein the subset of samples is identified based on a predetermined sampling rate and a predetermined offset.
9. The image decoding method of claim 1 , wherein the filter is determined to be one of a plurality of filters based on the value of the reference sample and a predetermined threshold.
10. The samples in the template of the reference block are classified into a plurality of sample sections based on the predetermined threshold; The video decoding method of claim 9, wherein the plurality of filters are derived for the plurality of sample intervals, respectively.
11. the predicted samples are generated based on a plurality of filters including the filter; The video decoding method of claim 1 , wherein the plurality of filters are derived from a plurality of candidate templates, respectively.
12. 1. A video encoding method, comprising: determining a reference block for inter prediction of a current block; generating a predicted sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset; deriving residual samples of the current block based on the predicted samples; encoding the residual samples.
13. A computer-readable recording medium, comprising: A computer-readable recording medium storing a bitstream generated by the video encoding method of claim 12.
14. 1. A data transmission method, comprising: obtaining a bitstream for video information; The bitstream comprises: determining a reference block of the current block for inter prediction of the current block; generating a predicted sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset; deriving residual samples of the current block based on the predicted samples; generated by encoding the residual samples, transmitting data including the bitstream.