Image encoding / decoding method and apparatus, and recording medium for storing bit stream
Through the filter generation based on motion information and the filter coefficient derivation method of reference blocks, the efficiency problems of inter prediction and reference sample compensation in high-resolution image coding are solved, and higher encoding accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202480008355.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-20
- Filing Date
- 2024-01-22
- Publication Date
- 2025-08-19
AI Technical Summary
Existing image compression techniques are insufficiently efficient in high resolution and high-quality image coding, especially in inter-frame prediction and reference sample compensation.
By determining the reference block based on the motion information of the current block, using a filter to generate a predicted sample, and reconstructing the current block based on the predicted sample, derive filter coefficients using adjacent samples and template regions, and selectively using in-template samples to improve coding performance.
Improve the accuracy and efficiency of image encoding, and improve the encoding performance by more accurately reflecting the current block characteristics, optimizing the balance between operational complexity and performance.
Smart Images

Figure CN120513633A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image encoding / decoding method and apparatus, and a recording medium storing a bitstream. Background Art
[0002] Recently, demands for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images have been increasing in various application fields, and therefore, efficient image compression technology is being discussed.
[0003] There are various technologies, such as inter-frame prediction technology that uses video compression technology to predict pixel values included in the current picture from pictures before or after the current picture, intra-frame prediction technology that predicts pixel values included in the current picture by using pixel information in the current picture, entropy coding technology that assigns short symbols to values that occur frequently and long symbols to values that occur infrequently, etc., and these image compression technologies can be used to effectively compress image data and transmit or store it. Summary of the Invention
[0004] Technical issues
[0005] The present invention aims to provide an inter-frame prediction method and device.
[0006] The present invention aims to provide a reference sample compensation method and device.
[0007] Technical Solution
[0008] According to the image decoding method and apparatus of the present disclosure, a reference block of the current block can be determined based on the motion information of the current block, prediction samples of the current block are generated by applying a filter to at least one of a reference sample belonging to the reference block, at least one adjacent sample adjacent to the reference sample, or a predetermined offset, and the current block is reconstructed based on the prediction samples.
[0009] In the image decoding method and apparatus according to the present disclosure, the adjacent samples may include at least one of left adjacent samples, above adjacent samples, left adjacent samples, or below adjacent samples.
[0010] In the image decoding method and apparatus according to the present disclosure, filter coefficients of the filter may be determined based on at least one sample of a first neighboring region belonging to a current block and at least one sample of a second neighboring region belonging to a reference block.
[0011] In the image decoding method and apparatus according to the present disclosure, the second adjacent area may include a template of the reference block and an area extending from the template of the reference block through N sample lines.
[0012] In the image decoding method and apparatus according to the present disclosure, the template may include at least one of a left adjacent region, an upper adjacent region, an upper left adjacent region, an upper right adjacent region, or a lower left adjacent region.
[0013] In the image decoding method and apparatus according to the present disclosure, at least one of the position or the size of the template may be determined based on the size of the reference block.
[0014] In the image decoding method and apparatus according to the present disclosure, a template may be determined as any one of a plurality of template candidates based on template type information transmitted through a bitstream signal, and the template type information may indicate at least one of a position or a size of the template.
[0015] In the image decoding method and apparatus according to the present disclosure, filter coefficients of a filter may be derived by using some samples belonging to the second neighboring region of a reference block, and these samples may be specified based on a predetermined sampling rate and a predetermined offset.
[0016] In the image decoding method and apparatus according to the present disclosure, the filter may be determined to be any one of a plurality of filters based on a value of a reference sample and a predetermined threshold.
[0017] In the image decoding method and apparatus according to the present disclosure, samples within a template of a reference block may be classified into a plurality of sample intervals based on a predetermined threshold, and a plurality of filters may be derived for the plurality of sampling intervals, respectively.
[0018] In the image decoding method and apparatus according to the present disclosure, prediction samples may be generated based on a plurality of filters including a filter, and the plurality of filters may be derived from a plurality of template candidates, respectively.
[0019] According to the image encoding method and apparatus of the present disclosure, a reference block for inter-frame prediction of a current block can be determined, prediction samples of the current block can be generated by applying a filter to at least one of a reference sample belonging to the reference block, at least one adjacent sample adjacent to the reference sample, or a predetermined offset, residual samples of the current block can be derived based on the prediction samples, and the residual samples can be encoded.
[0020] A computer-readable digital storage medium storing encoded video / image information is provided, which enables a decoding device according to the present disclosure to perform an image decoding method.
[0021] A computer-readable digital storage medium storing video / image information generated according to an image encoding method is provided according to the present disclosure.
[0022] Provided are a method and apparatus for transmitting video / image information generated according to an image encoding method according to the present disclosure.
[0023] Beneficial effects
[0024] According to the present invention, encoding performance can be improved by deriving a compensation model with high accuracy and generating a prediction sample by referring to more neighboring samples in addition to a reference sample corresponding to a current sample.
[0025] According to the present disclosure, when samples within a template are selectively used based on a predetermined sampling rate and offset, an optimal trade-off between operation complexity and performance can be derived.
[0026] According to the present disclosure, the best trade-off between operation complexity and performance can be derived by adjusting the position / size of the template according to the block size.
[0027] According to the present disclosure, by selectively using any one of a plurality of template candidates, a model can be defined by more accurately reflecting the characteristics of the current block, thereby improving encoding performance.
[0028] According to the present disclosure, prediction performance can be improved by defining multiple filter types and applying a more appropriate filter type to each reference sample.
[0029] According to the present disclosure, a final prediction block may be configured by reflecting appropriate weights of various equations between a current block and a reference block, thereby improving encoding performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A video / image encoding system according to the present disclosure is shown.
[0031] Figure 2 A schematic block diagram illustrating an encoding device to which embodiments of the present disclosure are applicable and which performs encoding of a video / image signal.
[0032] Figure 3 A schematic block diagram illustrating a decoding device to which embodiments of the present disclosure are applicable and which performs decoding of a video / image signal.
[0033] Figure 4 An image decoding method performed by an encoding device according to the present disclosure is shown.
[0034] Figure 5 A schematic configuration of the inter-frame predictor 332 that performs the image decoding method according to the present disclosure is shown.
[0035] Figure 6 An image encoding method performed by an encoding device according to the present disclosure is shown.
[0036] Figure 7 A schematic configuration of the inter-frame predictor 220 performing the image encoding method according to the present disclosure is shown.
[0037] Figure 8 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown. DETAILED DESCRIPTION
[0038] Because the present disclosure can be modified in various ways and has several embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. However, it is not intended to limit the present disclosure to specific embodiments, and it should be understood that all variations, equivalents, and alternatives are included in the spirit and technical scope of the present disclosure. When describing each of the drawings, similar reference numerals are used for similar components.
[0039] Terms such as first, second, etc. may be used to describe various components, but components should not be limited by these terms. These terms are only used to distinguish one component from other components. For example, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component without departing from the scope of the present disclosure. Terms and / or combinations of any one or more related statement items include multiple related statement items.
[0040] When a component is referred to as being "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to the other component, but another component may also exist in between. On the other hand, when a component is referred to as being "directly connected" or "directly linked" to another component, it should be understood that another component does not exist in between.
[0041] The terms used in this application are only used to describe specific embodiments and are not intended to limit the present disclosure. Unless the context clearly indicates otherwise, singular expressions include plural expressions. In this application, it should be understood that terms such as "including" or "having" are intended to designate the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0042] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the Versatile Video Coding (VVC) standard. Furthermore, the methods / embodiments disclosed herein may be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second-generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0043] This specification proposes various embodiments of video / image encoding, and unless otherwise specified, these embodiments may be performed in combination with each other.
[0044] Here, video can refer to a collection of images over time. A picture generally refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms part of a picture during coding. A slice / tile may include at least one Coding Tree Unit (CTU). A picture may consist of at least one slice / tile. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and tile row of a picture. A tile column is a rectangular area of CTUs with the same height as the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs with a height specified by the picture parameter set and the same width as the picture. CTUs within a tile may be arranged consecutively according to a CTU raster scan, while tiles within a picture may be arranged consecutively according to a tile raster scan. A slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that may be exclusively included in a single NAL unit. A picture may also be divided into at least two sub-pictures. A sub-picture may be a rectangular area of at least one slice within a picture.
[0045] Pixel, pixel, or picture element can refer to the smallest unit that constitutes a picture (or image). In addition, "sample" can be used as a term corresponding to pixel. Sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.
[0046] A unit can represent the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the corresponding region. A unit can include a luma block and two chroma (e.g., CB, CR) blocks. In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." In general, an MxN block can include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.
[0047] Here, "A or B" can mean "only A," "only B," or "both A and B." In other words, herein, "A or B" can be interpreted as "A and / or B." For example, herein, "A, B, or C" can mean "only A," "only B," "only C," or "any combination of A, B, and C."
[0048] As used herein, a slash mark ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B, or C."
[0049] Here, “at least one of A and B” may refer to “only A,” “only B,” or “both A and B.” In addition, herein, expressions such as “at least one of A or B” or “at least one of A and / or B” can be interpreted in the same manner as “at least one of A and B.”
[0050] In addition, herein, “at least one of A, B, and C” may refer to “only A,” “only B,” “only C,” or “any combination of A, B, and C.” In addition, “at least one of A, B, or C” or “at least one of A, B, and / or C” may refer to “at least one of A, B, and C.”
[0051] Additionally, parentheses used herein may refer to "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction." In other words, "prediction" here is not limited to "intra-frame prediction," and "intra-frame prediction" may be provided as an example of "prediction." Furthermore, even when "prediction (i.e., intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction."
[0052] Here, technical features described individually in one drawing may be implemented individually or simultaneously.
[0053] Figure 1 A video / image encoding system according to the present disclosure is shown.
[0054] refer to Figure 1 , a video / image encoding system may include a first device (source device) and a second device (receiving device).
[0055] The source device can transmit the encoded video / image information or data to the receiving device in the form of a file or stream transmission via a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0056] A video source can obtain video / images through a process of capturing, synthesizing, or generating video / images. A video source can include both devices that capture video / images and devices that generate video / images. Devices that capture video / images may include at least one camera, a video / image archive containing previously captured video / images, and the like. Devices that generate video / images may include computers, tablets, smartphones, and the like, and can (electronically) generate video / images. For example, a virtual video / image can be generated by a computer, etc., and in this case, the process of capturing video / images can be replaced by a process of generating related data.
[0057] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0058] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device in the form of a file or stream transmission via a digital storage medium or network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include components for generating a media file in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0059] The decoding device may decode the video / image by performing a series of processes such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0060] The renderer may render the decoded video / image, and the rendered video / image may be displayed through a display unit.
[0061] Figure 2 A rough block diagram showing an encoding device to which an embodiment of the present disclosure can be applied and which performs encoding of a video / image signal is shown.
[0062] refer to Figure 2The encoding device 200 may include an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, an inverse quantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Furthermore, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0063] The image splitter 210 may partition an input image (or picture, or frame) input to the encoding apparatus 200 into at least one processing unit. For example, a processing unit may be referred to as a coding unit (CU). In this case, the CU may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quad-tree binary-tree ternary-tree (QTBTTT) structure.
[0064] For example, one coding unit can be split into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure can be applied first, and the binary tree structure and / or the ternary structure can be applied later. Alternatively, the binary tree structure can be applied before the quadtree structure. The encoding process according to this specification can be performed based on the final coding unit that is no longer split. In this case, based on the encoding efficiency according to the image characteristics, etc., the maximum coding unit can be directly used as the final coding unit, or if necessary, the coding unit can be recursively split into coding units of a deeper depth, and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which are described later.
[0065] As another example, a processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or partitioned from the final coding unit. A prediction unit may be a unit for sample prediction, and a transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0066] In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." In general, an MxN block can represent a set of transform coefficients or samples consisting of M columns and N rows. A sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample can be used as a term to refer to a picture (or image) corresponding to a pixel or picture element.
[0067] The encoding device 200 may subtract the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within the encoding device 200 may be referred to as a subtractor 231.
[0068] The predictor 220 may perform prediction on a block to be processed (hereinafter referred to as a current block) and generate a predicted block including prediction samples for the current block. The predictor 220 may determine whether intra prediction or inter prediction is applied in units of the current block or CU. The predictor 220 may generate various information about the prediction, such as prediction mode information, and transmit it to the entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0069] The intra-frame predictor 222 can predict the current block by referencing samples within the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or can be located a certain distance away from the current block. In intra-frame prediction, the prediction mode may include at least one non-directional mode and multiple directional modes. The non-directional mode may include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional mode may include 33 directional modes or 65 directional modes. However, this is just an example, and more or fewer directional modes may be used depending on the configuration. The intra-frame predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0070] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter-frame prediction direction information (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in a reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture containing temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as the motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0071] The predictor 220 can generate prediction signals based on various prediction methods described later. For example, the predictor can apply not only intra prediction or inter prediction to predict a block, but also both intra and inter prediction simultaneously. This is referred to as a combined inter and intra prediction (CIIP) mode. Alternatively, the predictor can be based on an intra block copy (IBC) prediction mode or a palette mode for block-specific prediction. The IBC prediction mode or palette mode can be used for content image / video coding, such as screen content coding (SCC), for gaming and the like. IBC essentially performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives reference blocks within the current picture. In other words, IBC can utilize at least one of the inter prediction techniques described herein. The palette mode can be considered an example of intra coding or intra prediction. When palette mode is applied, sample values within the picture can be signaled based on information about a palette table and palette index. The prediction signal generated by the predictor 220 can be used to generate a reconstructed signal or a residual signal.
[0072] Transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when relationship information between pixels is expressed as a graph. CNT refers to a transform obtained by generating a prediction signal using all previously reconstructed pixels. Furthermore, the transform process can be applied to square pixel blocks of the same size or to non-square blocks of variable size.
[0073] The quantizer 233 may quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scanning order, and may generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0074] The entropy encoder 240 may perform various encoding methods such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. The entropy encoder 240 may encode information necessary for video / video image reconstruction (eg, values of syntax elements, etc.) in addition to transform coefficients quantized together or individually.
[0075] Encoded information (e.g., encoded video / image information) can be transmitted or stored in a bitstream in units of Network Abstraction Layer (NAL) units. The video / image information may further include information regarding various parameter sets, such as the Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may further include general constraint information. Information and / or syntax elements transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the aforementioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media, such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (not shown) for transmission and / or the storage unit (not shown) for storing the signal output from the entropy encoder 240 may be configured as internal or external components of the encoding device 200, or the transmission unit may also be included in the entropy encoder 240.
[0076] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual samples) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantizer 234 and the inverse transformer 235. The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, or reconstructed sample array). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current picture and can also be used for inter-frame prediction of the next picture through filtering, which will be described later. Luma mapping with chroma scaling (LMCS) can also be applied during the picture encoding and / or reconstruction process.
[0077] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and the modified reconstructed picture can be stored in the memory 270, specifically in the DPB of the memory 270. Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information about filtering and send it to the entropy encoder 240. The information about filtering can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0078] The modified reconstructed picture sent to the memory 270 may be used as a reference picture in the inter-frame predictor 221. When inter-frame prediction is applied therethrough, the encoding apparatus can avoid prediction mismatch in the encoding apparatus 200 and the decoding apparatus, and can also improve encoding efficiency.
[0079] The DPB of the memory 270 can store the modified reconstructed picture for use as a reference picture in the inter-frame predictor 221. The memory 270 can store motion information of the block from which the motion information in the current picture was derived (or encoded) and / or motion information of blocks in pre-reconstructed pictures. The stored motion information can be sent to the inter-frame predictor 221 to be used as motion information for spatially neighboring blocks or motion information for temporally neighboring blocks. The memory 270 can store reconstructed samples of the reconstructed blocks in the current picture and send them to the intra-frame predictor 222.
[0080] Figure 3 A rough block diagram showing a decoding device to which an embodiment of the present disclosure can be applied and which performs decoding of a video / image signal is shown.
[0081] refer to Figure 3 , the decoding apparatus 300 may be configured by including an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 321.
[0082] According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 described above may be configured by a single hardware component (e.g., a decoder chipset or processor). Furthermore, the memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.
[0083] When a bit stream including video / image information is input, the decoding apparatus 300 may generate a decoded image in response to the bit stream received in the decoded image. Figure 2 The image is reconstructed by processing video / image information in the encoding device of the decoding apparatus. For example, the decoding apparatus 300 can derive a unit / block based on relevant information of block segmentation obtained from the bit stream. The decoding apparatus 300 can perform decoding by using a processing unit applied in the encoding apparatus. Therefore, the processing unit of decoding can be a coding unit, and the coding unit can be divided from a coding tree unit or a maximum coding unit according to a quadtree structure, a binary tree structure and / or a ternary tree structure. At least one transform unit can be derived from the coding unit. And, the reconstructed image signal decoded and output by the decoding apparatus 300 can be played by a playback device.
[0084] The decoding device 300 may receive the data in the form of a bit stream from Figure 2 The received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the Adaptation Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Furthermore, the video / image information may further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or general constraint information. Signaled / received information and / or syntax elements described later herein can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 can decode information in the bitstream based on coding methods such as Exponential Golomb coding, CAVLC, or CABAC, and output values for syntax elements necessary for image reconstruction and quantized values for residual transform coefficients. In more detail, the CABAC entropy decoding method receives bins corresponding to each syntax element from the bitstream, determines a context model using information about the syntax element to be decoded, decoded information about surrounding blocks and the block to be decoded, or information about symbols / bins decoded in a previous step, performs arithmetic decoding on the bins by predicting the probability of occurrence of the bins based on the determined context model, and generates symbols corresponding to the values of each syntax element. After determining the context model, the CABAC entropy decoding method updates the context model using information about the decoded symbols / bins for the context model for the next symbol / bin. Information decoded in the entropy decoder 310 regarding prediction is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values entropy-decoded in the entropy decoder 310, namely, quantized transform coefficients and related parameter information, are input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). Furthermore, information regarding filtering in the information decoded in the entropy decoder 310 is provided to the filter 350. Meanwhile, a receiving unit (not shown) that receives a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300 or the receiving unit may be a component of the entropy decoder 310 .
[0085] Meanwhile, the decoding device according to this specification may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of an inverse quantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0086] The inverse quantizer 321 may inversely quantize the quantized transform coefficients and output the transform coefficients. The inverse quantizer 321 may rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The inverse quantizer 321 may inversely quantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0087] The inverse transformer 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0088] The predictor 320 may perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor 320 may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310, and determine a specific intra / inter prediction mode.
[0089] The predictor 320 can generate a prediction signal based on various prediction methods described later. For example, the predictor 320 can not only apply intra prediction or inter prediction to predict a block, but can also apply both intra prediction and inter prediction simultaneously. This can be referred to as a combined inter and intra prediction (CIIP) mode. Furthermore, the predictor can be based on an intra block copy (IBC) prediction mode or a palette mode for block prediction. The IBC prediction mode or palette mode can be used for content image / video encoding, such as screen content coding (SCC), for gaming, etc. IBC essentially performs prediction within the current picture, but it can be performed similarly to inter prediction in that it derives reference blocks within the current picture. In other words, IBC can use at least one of the inter prediction techniques described herein. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0090] The intra-frame predictor 331 can predict the current block by referencing samples within the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or at a certain distance away from the current block. In intra-frame prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The intra-frame predictor 331 can determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring blocks.
[0091] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and prediction information can include information indicating the inter-frame prediction mode used for the current block.
[0092] The adder 340 may add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331) to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the reconstructed block.
[0093] Adder 340 may be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering described later, or may be used for inter-frame prediction of the next picture. Luma mapping with chroma scaling (LMCS) may also be applied during picture decoding.
[0094] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and send the modified reconstructed picture to the memory 360, specifically the DPB of the memory 360. Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0095] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block derived (or decoded) from the motion information in its current picture and / or the motion information of the block in the pre-reconstructed picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of the spatially neighboring blocks or the motion information of the temporally neighboring blocks. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and send them to the intra-frame predictor 331.
[0096] Here, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may also be equally or correspondingly applied to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively.
[0097] Figure 4 An image decoding method performed by an encoding device according to the present disclosure is shown.
[0098] This disclosure proposes a method for defining and applying a compensation model to inter-frame prediction blocks. This compensation method utilizes a predetermined scaling factor (α) and offset (β) to configure a linear model for compensation between a current block and a reference block. In other words, the linear model can be configured by utilizing neighboring samples of the current block and neighboring samples of the reference block to satisfy Equation 1 below.
[0099] [Equation 1]
[0100]
[0101] The compensation model in the present disclosure may refer to an illumination compensation model between a current block and a reference block, but is not limited thereto. Hereinafter, a compensation method for an inter-frame prediction block will be described in detail.
[0102] refer to Figure 4 , a reference block of the current block may be determined based on the motion information of the current block S400 .
[0103] A reference block according to the present disclosure may belong to a reference picture. Here, the reference picture may be a different picture from the current picture to which the current block belongs. For example, the reference picture may have a different picture order count (POC) than the current picture. Alternatively, the reference picture may have a different decoding order than the current picture. The reference picture may be a picture that was encoded / decoded before the current picture.
[0104] The motion information of the current block may include at least one of a motion vector, a reference picture index, or prediction direction information. The motion vector may specify the position of the reference block within the reference picture. For example, the motion vector may refer to the difference between the position of the current block within the current picture and the position of the reference block within the reference picture. The reference picture index may specify any one of a plurality of reference pictures belonging to a reference picture list for the current block. The prediction direction information may indicate at least one of whether the current block is predicted based on a reference picture list in the L0 direction (List0) or whether the current block is predicted based on a reference picture list in the L1 direction (List1).
[0105] refer to Figure 4 , a filter may be applied to at least one of a reference sample belonging to a reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset to generate a prediction sample of the current block S410.
[0106] The adjacent samples used as input to the filter may be one or more samples adjacent to the reference sample. The adjacent samples may include at least one of a left adjacent sample, a right adjacent sample, an upper adjacent sample, a lower adjacent sample, a left upper adjacent sample, a left lower adjacent sample, a right upper adjacent sample, or a right lower adjacent sample adjacent to the reference sample. The offset used as input to the filter may be determined based on the bit depth of the current image to which the current block belongs. However, this is not limiting, and samples that are not adjacent to the reference sample but are adjacent to at least one of the above-mentioned adjacent samples may also be used. Here, the shape of the area to which the filter is applied within the reference block may be square, non-square, cross, or diamond.
[0107] Depending on the position of the reference samples to which the filter is applied, the filter length, or the filter shape, there may be a situation where at least one of the adjacent samples does not belong to the reference block. For example, the adjacent samples that do not belong to the reference block may belong to one or more sample lines adjacent to at least one of the left, top, right, or bottom directions of the reference block. Even when the adjacent samples used as filter input do not belong to the reference block, the prediction samples can be generated by using the corresponding adjacent samples as is. Alternatively, a restriction can be made so that only samples belonging to the reference block are used to perform compensation. When a filter is applied to a reference sample located at the boundary of a reference block, there may be unavailable adjacent samples. In this case, the unavailable adjacent samples can be replaced with the nearest available sample, and the filter can be applied.
[0108] The filter may be a filter for linear compensation. The filter may be a filter for a weighted sum of at least two of a reference sample belonging to a reference block, at least one adjacent sample adjacent to the reference sample, or a predetermined offset. The filter may be an M-tap convolution filter, where the filter length (or the number of filter coefficients) is M. M may be an integer greater than or equal to 2.
[0109] The filter coefficients of the filter can be derived based on one or more samples of the template belonging to the current block and one or more samples of the template belonging to the reference block. Hereinafter, for ease of description, the template of the current block and the template of the reference block are referred to as the current template and the reference template, respectively.
[0110] For example, a filter having a filter length of 6 may be used. In this case, as shown in Equation 2 below, the reference sample may be compensated based on a 6-tap convolution filter.
[0111] [Equation 2]
[0112]
[0113] According to Equation 2, the reference sample (pred x,y ) and adjacent samples to generate the predicted sample (pred') at the (x, y) position in the current block x,y Here, the adjacent samples may include the left adjacent samples (pred x-1,y ), right adjacent sample (pred x+1,y ), upper adjacent samples (pred x,y-1 ) and the next adjacent sample (pred x,y+1 ). When the width and height of the current block are W and H respectively, x may be greater than or equal to 0 and less than W, and y may be greater than or greater than 0 and less than H.
[0114] In Equation 2, ci (i=0..5) represents the filter coefficient, which can be calculated using the least squares method using samples belonging to the current template and samples belonging to the reference template. B can be defined as (1<<(BitDepth-1)). Here, BitDepth can refer to the bit depth of the current image.
[0115] A vector (A) including samples and offsets of a reference template, a vector (w) of filter coefficients, and a sample vector (p) of a current template may be defined as in Equation 3 below.
[0116] [Equation 3]
[0117]
[0118] When Equation 3 is expressed as a normal equation, it can be expressed as an equation consisting of the autocorrelation of A and the cross-correlation of A and P, as shown in Equation 4 below.
[0119] [Equation 4]
[0120]
[0121] Here, various methods can be used to derive the filter coefficients (w). For example, the filter coefficients can be derived by directly calculating the inverse matrix of the autocorrelation matrix. Alternatively, a linear system solution method can be used. For example, after the autocorrelation matrix is decomposed into a lower triangular matrix and an upper triangular matrix by using the Cholesky method, the filter coefficients can be derived by forward substitution and backward substitution. Alternatively, the autocorrelation matrix can be decomposed into a lower triangular matrix, a diagonal matrix, and an upper triangular matrix by using the LDL decomposition method to derive the filter coefficients. Alternatively, the filter coefficients can be derived by using the Gaussian elimination method. Alternatively, when the model in Equation 2 is defined for compensation of the inter-frame prediction block, and the filter coefficients in Equation 2 are derived by using a linear system solution of a specific method, it can be regarded as a method consistent with the present disclosure.
[0122] According to the present disclosure, the current template may refer to a pre-reconstructed adjacent region before the current block. For example, the current template may include at least one of the upper adjacent region, left adjacent region, upper left adjacent region, lower left adjacent region, or upper right adjacent region adjacent to the current block. Similarly, the reference template may include at least one of the upper adjacent region, left adjacent region, upper left adjacent region, lower left adjacent region, or upper right adjacent region adjacent to the reference block as the region corresponding to the current template.
[0123] The height of the upper adjacent region, the upper left adjacent region, and / or the upper right adjacent region may be N. The width of the left adjacent region, the upper left adjacent region, and / or the lower left adjacent region may be N, where N may be an integer greater than or equal to 1. The width of the upper adjacent region may be greater than or equal to the width of the current block (or reference block), and the height of the left adjacent region may be greater than or equal to the height of the current block (or reference block). However, this is not limiting, and the height of the upper adjacent region may be different from the width of the left adjacent region.
[0124] The filter coefficients can be derived by additionally using an area extending from the reference template by K sample lines. Alternatively, the filter coefficients can be derived by additionally using an area extending from the area consisting of the reference template and the reference block by K sample lines. Here, K can be an integer greater than or equal to 1. The extension direction can include at least one of the upper direction, the left direction, the lower direction, or the right direction. Whether to use the extension area can be adaptively determined based on at least one of the position of the reference sample to which the filter is applied, the filter length, or the filter shape.
[0125] Depending on at least one of the position of the reference sample to which the filter is applied, the filter length, or the filter shape, there may be cases where adjacent samples serving as input to the filter are unavailable. In such cases, the unavailable adjacent samples may be derived based on one or more samples within the reference template rather than using the extended template. The unavailable adjacent samples may be replaced with the sample closest to the unavailable adjacent sample among the samples belonging to the reference template.
[0126] Specifically, the reference template can be defined as a set of N sample lines adjacent to the top and left of the reference block. Here, the size (N) of the reference template can be an integer greater than or equal to 1. Assume that the width and height of the reference block are W and H respectively.
[0127] As an example, the reference template may consist of a WxN upper neighboring area, an NxN upper left neighboring area, and an NxH left neighboring area. Alternatively, the reference template may consist of a 2WxN upper neighboring area, an NxN upper left neighboring area, and an Nx2H left neighboring area. Alternatively, the reference template may consist of a WxN upper neighboring area, a WxN upper right neighboring area, an NxN upper left neighboring area, an NxH left neighboring area, and an NxH lower left neighboring area.
[0128] Alternatively, the reference template may be defined as a set of N sample lines adjacent to the top of the reference block. Here, the size (N) of the reference template may be an integer greater than or equal to 1. Assume that the width and height of the reference block are W and H, respectively.
[0129] As an example, the reference template may consist of a WxN upper adjacent region. Alternatively, the reference template may consist of 2 WxN upper adjacent regions. Alternatively, the reference template may consist of a WxN upper adjacent region and a WxN upper right adjacent region.
[0130] Alternatively, the reference template may be defined as a set of N sample lines adjacent to the left side of the reference block. Here, the size (N) of the reference template may be an integer greater than or equal to 1. Assume that the width and height of the reference block are W and H, respectively.
[0131] As an example, the reference template may consist of an NxH left-neighboring region. Alternatively, the reference template may consist of Nx2H left-neighboring regions. Alternatively, the reference template may consist of an NxH left-neighboring region and an NxH bottom-left-neighboring region.
[0132] Alternatively, the reference template can be defined as a set of N sample lines adjacent to the top and left sides of the reference block. Here, the size of the reference template (N) can be an integer greater than or equal to 1. Assume that the width and height of the reference block are W and H, respectively.
[0133] As an example, the reference template may consist of a WxN top-neighboring region and an NxH left-neighboring region. Alternatively, the reference template may consist of a 2WxN top-neighboring region and an Nx2H left-neighboring region. Alternatively, the reference template may consist of a WxN top-neighboring region, a WxN top-right neighboring region, an NxH left neighboring region, and an NxH bottom-left neighboring region.
[0134] The above-described method for configuring a reference template may be equally applied to the method for configuring a current template, and repeated descriptions will be omitted.
[0135] A plurality of template candidates may be defined equally for encoding devices and decoding devices. Here, the plurality of template candidates may include at least one of the template configuration examples regarding the above-mentioned template position. For example, the plurality of template candidates may include a first template candidate consisting of a WxN upper adjacent area, an NxN upper left adjacent area, and an NxH left adjacent area, and a second template candidate consisting of a 2WxN upper adjacent area, an NxN upper left adjacent area, and an Nx2H left adjacent area. Alternatively, the plurality of template candidates may include a first template candidate consisting of a WxN upper adjacent area and an NxH left adjacent area, a second template candidate consisting of a WxN upper adjacent area, and a third template candidate consisting of an NxH left adjacent area.
[0136] Alternatively, multiple template candidates can be configured according to any of the template configuration examples for the template position described above, but with different sizes (N). As an example, the multiple template candidates can include a first template candidate containing an N1xH left neighboring area and a second template candidate containing an N2xH left neighboring area. Here, N1 and N2 can be different integers.
[0137] Alternatively, multiple template candidates may be configured according to at least two of the template configuration examples regarding the above template position, and any one of the template candidates may have a different size (N) from another. Alternatively, the multiple template candidates may include a first template candidate consisting of a WxN1 upper neighboring area and an N1xH left neighboring area, a second template candidate consisting of a 2WxN2 upper neighboring area, and a third template candidate consisting of an N2x2H left neighboring area.
[0138] Any one of multiple predefined template candidates can be selectively used. Template type information can be used to select any one of the multiple template candidates. The template type information may include at least one of information indicating the location of the template or information indicating the size of the template. In this case, the template type information may be defined as an index to indicate the location and size of the template. Alternatively, the information indicating the location of the template and the information indicating the size of the template may be defined as separate indexes.
[0139] At least one of the template type information may be signaled via the bitstream. The template type information may be defined in region units such as a video sequence, picture, or slice, and may be signaled at higher levels such as a sequence parameter set (SPS), picture parameter set (PPS), picture header (PH), and slice header (SH). Alternatively, the template type information may be defined in block units such as coding tree units, coding units, or subblocks, and may be signaled at lower levels such as coding tree unit syntax and coding unit syntax. Information indicating the location of the template and information indicating the size of the template may be signaled in the same region unit or block unit. Information indicating the location of the template and information indicating the size of the template may be signaled in different region units or block units. Either the information indicating the location of the template or the information indicating the size of the template may be signaled in any of the aforementioned region units, and the other may be signaled in any of the aforementioned block units.
[0140] Alternatively, template type information (eg, information indicating the location of the template) may be derived based on the size of the reference block (or current block). In other words, any one of multiple template candidates may be selected based on the size of the reference block (or current block).
[0141] For example, when the width of the reference block is greater than its height, a template candidate including the upper neighboring region may be used. Conversely, when the width of the reference block is less than its height, a template candidate including the left neighboring region may be used. Otherwise (i.e., when the width and height of the reference block are the same), a template candidate including the left neighboring region, the upper neighboring region, and the upper-left neighboring region may be used, or a template candidate including the upper neighboring region and the left neighboring region but not the upper-left neighboring region may be used.
[0142] Alternatively, template type information (e.g., information indicating the size of the template) may be derived based on the size of the reference block (or current block). In other words, any one of multiple template candidates may be selected based on the size of the reference block (or current block). The size (N) of the template may be variably determined based on the size of the reference block (or current block).
[0143] For example, when the number of samples belonging to the reference block (or current block) is less than or equal to 64, a template candidate with a size of N1 may be selected, and otherwise, a template candidate with a size of N2 may be selected. Alternatively, when the number of samples belonging to the reference block (or current block) is less than or equal to 64, the size of the template may be determined to be N1, and otherwise, the size of the template may be determined to be N2. Here, N1 may be an integer smaller than N2. For example, N1 and N2 may be 4 and 6, respectively, but are not limited thereto.
[0144] As described above, by selectively using any one of multiple template candidates, it is possible to define a model that more accurately reflects the characteristics of the current block, thereby improving encoding performance. In addition, by adjusting the position / size of the template according to the block size, it is possible to derive the optimal trade-off between operation complexity and performance.
[0145] All samples belonging to the current and / or reference template can be used to derive the filter coefficients. Alternatively, the filter coefficients can be derived by using only some samples selected from among the samples belonging to the current and / or reference template. In other words, the compensation model can be derived based on only some samples without traversing all samples within the template.
[0146] Specifically, assume that sampling rate i and offset a are applied to the x-axis, and sampling rate j and offset b are applied to the y-axis. When the sample at the (x, y) position within the template is T x,y When , the samples required to derive the filter coefficients can be expressed as T i*x+a,j*y+b Here, the offset may be any integer including 0. a and b may be different integers, or may be the same integer.
[0147] As an example, the filter coefficients may be derived by using only samples belonging to at least one even row and at least one even column within the template. In this case, the samples used to derive the filter coefficients may be denoted as T 2x,2y The filter coefficients may also be derived by using only samples belonging to at least one odd row and at least one odd column within the template. The filter coefficients may also be derived by using only samples belonging to at least one odd row or at least one odd column within the template. The filter coefficients may also be derived by using only samples belonging to at least one even row or at least one even column within the template.
[0148] By selectively using samples within a template based on a predetermined sampling rate and offset, an optimal trade-off between operational complexity and performance can be derived. Furthermore, encoding performance can be improved by deriving a compensation model with high accuracy and generating prediction samples by referencing more neighboring samples in addition to the reference sample corresponding to the current sample.
[0149] For the current block, at least two compensation models can be derived / defined instead of a single compensation model. Any one of the multiple compensation models can be selected based on a specific condition. In this case, any one of the multiple compensation models can be selected for each reference sample within the reference block. Alternatively, any one of the multiple compensation models can be selected for each sample interval of one or more reference samples. Here, the specific condition can be the magnitude relationship between a specific reference sample within the reference block and a predetermined threshold. Hereinafter, the compensation model is referred to as a filter.
[0150] For example, a filter can be derived by traversing all or part of the samples in the reference template. In this case, the samples in the reference template can be classified into multiple sample intervals based on whether the value of the sample within the reference template is less than a predetermined threshold. A filter can be derived for each sampling interval.
[0151] The sample interval to which the corresponding reference sample belongs can be determined based on whether the value of the reference sample within the reference block is less than a predetermined threshold. The filter corresponding to the determined sampling interval can be applied to the corresponding reference sample. For example, when (L-1) thresholds are derived / defined, L filter types can be derived / defined for the current block. When the value of the reference sample within the reference block is greater than or equal to the first threshold, the first filter can be used. When the value of the reference sample within the reference block is less than the first threshold and greater than or equal to the second threshold, the second filter can be used. When the value of the reference sample within the reference block is less than the (L-2)th threshold and greater than or equal to the (L-1)th threshold, the (L-1)th filter can be used. In other cases, the Lth filter can be used. L can be an integer greater than or equal to 1.
[0152] The threshold according to the present disclosure may be determined based on an average value of samples within a reference block, at least one of a minimum value, a maximum value, or a median value of samples within a reference block, a bit depth of a current image, or information signaled to specify the threshold.
[0153] According to the present disclosure, prediction performance can be improved by defining multiple filter types and applying a more appropriate filter type to each reference sample.
[0154] Alternatively, at least two filters may be derived / defined for the current block. The at least two filters may be applied to reference samples to generate at least two prediction samples. The at least two prediction samples may be weighted and summed to generate a final prediction sample.
[0155] As an example, as in the above-described method, at least two filters can be derived / defined based on sample values within a reference template and a predetermined threshold. The at least two filters can be applied to samples of the current template, respectively, to generate at least two filtered current templates. A cost can be calculated between the filtered current template and the reference template. Here, the cost can be a value obtained by measuring the difference between the filtered current template and the reference template. The cost can be calculated based on the sum of absolute differences (SAD), the sum of absolute transformed differences, or the sum of squared differences.
[0156] At least two filters may be applied to reference samples within a reference block to generate at least two prediction samples. A weighted sum of the at least two prediction samples may be performed to generate a final prediction sample. In this case, a weight for the weighted sum may be determined based on a calculated cost. The sum of the costs calculated for the at least two filters may serve as the denominator of the weight, and a specific cost may serve as the numerator of the weight. In this case, a greater weight may be applied to the prediction sample generated based on the filter with a relatively small cost.
[0157] For example, a final prediction sample based on a weighted sum of two prediction samples may be generated as shown in Equation 5 below.
[0158] [Equation 5]
[0159]
[0160] In Equation 5, pred x,y f0(refBlock ) may represent the final prediction sample at the (x, y) position within the current block. Cost0 may refer to a first cost between the reference template and the current template to which a first filter among a plurality of filters is applied. Cost1 may refer to a second cost between the reference template and the current template to which a second filter among a plurality of filters is applied. x,y ) may refer to a first prediction sample generated by applying the first filter to a reference sample at the (x, y) position within the reference block. x,y ) may refer to a second predicted sample generated by applying the second filter to the reference sample at the (x, y) position within the reference block.
[0161] According to Equation 5, when Cost0 is less than Cost1, a greater weight may be applied to the first prediction sample generated by applying the first filter. In other words, the weight ratio applied to the first prediction sample and the second prediction sample may be Cost1:Cost0. Conversely, when Cost0 is greater than or equal to Cost1, a greater weight may be applied to the second prediction sample generated by applying the second filter. In other words, the weight ratio applied to the first prediction sample and the second prediction sample may be Cost1:Cost0.
[0162] According to the above method, the final prediction block can be configured by reflecting appropriate weights used for various equations between the current block and the reference block, thereby improving encoding performance.
[0163] Alternatively, at least two filters can be derived / defined for the current coding block instead of a single filter. These at least two filters can be derived based on at least two template candidates (or samples at specific locations within multiple template candidates). These at least two template candidates can be defined using the above-described method for configuring template candidates, with overlapping descriptions omitted. These at least two filters can be applied to reference samples of a reference block to generate at least two prediction samples. These at least two prediction samples can be weighted and summed to generate a final prediction sample.
[0164] For example, the at least two filters used for the current block may include at least one of a first filter derived from the left adjacent area, a second filter derived from the upper adjacent area, a third filter derived from the left and upper adjacent areas, or a fourth filter derived from the left, upper, and upper-left adjacent areas. Alternatively, the at least two filters may be derived by using different sampling rates, offsets, and / or thresholds within the same template candidate.
[0165] For example, a first filter may be derived from a first template candidate including a left adjacent region, and a second filter may be derived from a second template candidate including an upper adjacent region. These two filters may be applied to samples of the current template, respectively, to generate two filtered current templates. A cost may be calculated between the filtered current template and the reference template. Here, the cost may be a value obtained by measuring the difference between the filtered current template and the reference template. The cost may be calculated based on the sum of absolute differences (SAD), the sum of absolute transformed differences, or the sum of squared differences.
[0166] A first prediction sample can be generated by applying a first filter to a reference sample within a reference block, and a second prediction sample can be generated by applying a second filter to the reference sample. A final prediction sample can be generated by a weighted sum of the first prediction sample and the second prediction sample. The weight used for the weighted sum can be determined based on the calculated cost. The sum of the costs calculated for the two filters can become the denominator of the weight, and a specific cost can become the numerator of the weight. In this case, a greater weight can be applied to the prediction sample generated based on the filter with a relatively small cost.
[0167] For example, a final prediction sample based on a weighted sum of two prediction samples may be generated as shown in Equation 6 below.
[0168] [Equation 6]
[0169]
[0170] In Equation 6, pred x,y It can represent the final predicted sample at the (x, y) position in the current block. LeftCost may refer to a first cost between a reference template and a current template to which a first filter among a plurality of filters is applied. Here, the first filter may be derived from a second template candidate including a left adjacent region. Above It may refer to a second cost between the reference template and the current template to which the second filter among the plurality of filters is applied. Here, the second filter may be derived from the first template candidate including the upper adjacent region. Left (refBlock x,y ) may refer to a first prediction sample generated by applying the first filter to a reference sample at the (x, y) position within the reference block. Above (refBlock x,y ) may refer to a second predicted sample generated by applying the second filter to the reference sample at the (x, y) position within the reference block.
[0171] According to Equation 6, when Cost Left Less than Cost Above When , a greater weight may be applied to the first prediction sample generated by applying the first filter. In other words, the weight ratio applied to the first prediction sample and the second prediction sample may be Cost Above :Cost Left On the contrary, when Cost Left Greater than or equal to Cost Above When , a greater weight may be applied to the second prediction sample generated by applying the second filter. In other words, the weight ratio applied to the first prediction sample and the second prediction sample may be Cost Above :Cost Left .
[0172] According to the above method, the final prediction block can be configured by reflecting appropriate weights of various equations between the current block and the reference block, thereby improving encoding performance.
[0173] refer to Figure 4 , the current block may be reconstructed based on the prediction samples of the current block S420.
[0174] The current sample of the current block may be reconstructed based on the prediction sample and the residual sample of the current block.The residual sample may be derived by deriving a transform coefficient based on residual information obtained as a bitstream and performing dequantization / inverse transform on the derived transform coefficient.
[0175] At the same time, the above compensation method can also be applied in the same / similar manner to linear model-based intra-frame prediction between luminance component blocks and chrominance component blocks. In this case, the current block and reference block in the above compensation method can be understood by replacing them with the chrominance component block and luminance component block of the current block, respectively. In addition, the reference block in the above compensation method can be understood as belonging to the same current picture as the current block.
[0176] Figure 5 FIG. 3 shows a schematic configuration of the inter-frame predictor 332 that performs the image decoding method according to the present disclosure.
[0177] refer to Figure 5 , the inter predictor 332 may include a reference block determiner 500 and a reference sample compensator 510 .
[0178] The reference block determiner 500 may determine the reference block of the current block based on the motion information of the current block. Figure 4 Same as described.
[0179] The reference sample compensator 510 may generate a prediction sample of the current block by applying a filter to at least one of a reference sample belonging to a reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset.
[0180] The filter has a predetermined filter length and may be a filter for linear compensation. Alternatively, the filter may be a filter for a weighted sum of at least two of a reference sample belonging to a reference block, at least one adjacent sample adjacent to the reference sample, or a predetermined offset.
[0181] The filter coefficients of the filter can be derived based on one or more samples of the template belonging to the current block and one or more samples of the template belonging to the reference block. Figure 4 Same as described.
[0182] In addition, multiple filters can be derived / defined for the current block, and the method for generating prediction samples based on multiple filters is the same as the reference Figure 4 The method described is the same.
[0183] The prediction sample, which is the output of the reference sample compensator 510 , may be input to the adder 340 of the decoding apparatus 300 and used to reconstruct the current block.
[0184] Figure 6 An image encoding method performed by an encoding device according to the present disclosure is shown.
[0185] refer to Figure 6 , a reference block for inter-frame prediction of a current block may be determined at S600 .
[0186] A reference block according to the present disclosure may belong to a reference picture. Here, the reference picture may be a different picture from the current picture to which the current block belongs. For example, the reference picture may have a different picture order count (POC) than the current picture. Alternatively, the reference picture may have a different decoding order than the current picture. The reference picture may be a picture that was encoded / decoded before the current picture.
[0187] refer to Figure 6 , a filter may be applied to at least one of a reference sample belonging to a reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset to generate a prediction sample of the current block S610.
[0188] The filter has a predetermined filter length and may be a filter for linear compensation. Alternatively, the filter may be a filter for a weighted sum of at least two of a reference sample belonging to a reference block, at least one adjacent sample adjacent to the reference sample, or a predetermined offset.
[0189] The filter coefficients of the filter can be derived based on one or more samples of the template belonging to the current block and one or more samples of the template belonging to the reference block. Figure 4 Same as described.
[0190] In addition, multiple filters can be derived / defined for the current block, and the method for generating prediction samples based on multiple filters is the same as the reference Figure 4 The method described is the same.
[0191] refer to Figure 6 , residual samples of the current block may be derived based on the prediction samples S620 , and the residual samples may be encoded to generate a bitstream S630 .
[0192] Figure 7 A schematic configuration of the inter-frame predictor 221 that performs the image encoding method according to the present disclosure is shown.
[0193] refer to Figure 7 , the inter predictor 221 may include a reference block determiner 700 and a reference sample compensator 710 .
[0194] The reference block determiner 700 may determine a reference block for inter-frame prediction of a current block. Figure 6 Same as described.
[0195] The reference sample compensator 710 may generate a prediction sample of the current block by applying a filter to at least one of a reference sample belonging to a reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset.
[0196] The filter has a predetermined filter length and may be a filter for linear compensation. Alternatively, the filter may be a filter for a weighted sum of at least two of a reference sample belonging to a reference block, at least one adjacent sample adjacent to the reference sample, or a predetermined offset.
[0197] The filter coefficients of the filter can be derived based on one or more samples of the template belonging to the current block and one or more samples of the template belonging to the reference block. Figure 4 Same as described.
[0198] In addition, multiple filters can be derived / defined for the current block, and the method of generating prediction samples based on multiple filters is the same as that by referring to Figure 4 The method described is the same.
[0199] The prediction sample as the output of the reference sample compensator 710 may be input to the residual processor 230 of the encoding device 200. The residual processor 230 may derive a residual sample based on the prediction sample, perform transformation / quantization on the residual sample to derive a transform coefficient, and generate residual information about the derived transform coefficient. The residual information may be input to the entropy encoder 240, and the entropy encoder 240 may encode the residual information to generate a bitstream.
[0200] In the above embodiments, the methods are described as a series of steps or blocks based on the flowcharts, but the corresponding embodiments are not limited to the order of the steps, and some steps may occur simultaneously or in a different order than other steps described above. In addition, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included or one or more steps in the flowcharts may be deleted without affecting the scope of the embodiments of the present disclosure.
[0201] The above-mentioned method according to an embodiment of the present disclosure can be implemented in the form of software, and the encoding device and / or decoding device according to the present disclosure can be included in a device that performs image processing, such as a TV, a computer, a smart phone, a set-top box, a display device, etc.
[0202] In the present disclosure, when embodiments are implemented as software, the above methods can be implemented as modules (processes, functions, etc.) that perform the above functions. The modules can be stored in memory and executed by a processor. The memory can be located internally or externally to the processor and can be connected to the processor via various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, logic circuits, and / or a data processing device. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described herein can be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information regarding instructions) or algorithms used for implementation can be stored on a digital storage medium.
[0203] In addition, the decoding device and encoding device using the embodiments of the present disclosure can be included in multimedia broadcast transmission and reception devices, mobile communication terminals, home theater video equipment, digital theater video equipment, surveillance cameras, video conversation equipment, real-time communication equipment such as video communication, mobile streaming equipment, storage media, cameras, equipment for providing video on demand (VoD) services, over-the-top (OTT) equipment, equipment for providing Internet streaming services, three-dimensional (3D) video equipment, virtual reality (VR) equipment, augmented reality (AR) equipment, videophone video equipment, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, ship terminals, etc.), and medical video equipment, and can be used to process video signals or data signals. For example, over-the-top (OTT) devices may include game consoles, Blu-ray players, networked TVs, home theater systems, smartphones, tablet computers, digital video recorders (DVRs), and the like.
[0204] In addition, the processing method of the embodiment of the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment of the present disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media may include, for example, Blu-ray discs (BDs), universal serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical media storage devices. In addition, computer-readable recording media include media implemented in the form of carrier waves (for example, transmitted via the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or can be sent via a wired or wireless communication network.
[0205] In addition, the embodiments of the present disclosure may be implemented by a computer program product through program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0206] Figure 8 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown.
[0207] refer to Figure 8 The content streaming transmission system to which the embodiments of the present disclosure are applied may mainly include an encoding server, a streaming transmission server, a web server, a media storage, a user device, and a multimedia input device.
[0208] The encoding server compresses content input from a multimedia input device such as a smartphone, a camera, or a camcorder into digital data to generate a bitstream and transmits it to the streaming server. As another example, when a multimedia input device such as a smartphone, a camera, or a camcorder directly generates a bitstream, the encoding server may be omitted.
[0209] A bitstream may be generated by applying the encoding method or the bitstream generating method of the embodiment of the present disclosure, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0210] The streaming server transmits multimedia data to the user device via a web server based on the user's request, and the web server serves as a medium for notifying the user of available services. When the user requests the desired service from the web server, the web server delivers it to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, and in this case, the control server controls the commands and responses between each device in the content streaming system.
[0211] The streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0212] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smart watches, smart glasses, head-mounted displays (HMDs), digital TVs, desktop computers, digital signage, etc.).
[0213] Each server in the content streaming system may be operated as a distributed server, and in this case, data received from each server may be distributed and processed.
[0214] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of the present disclosure can be combined and implemented as a device, and the technical features of the device claims of the present disclosure can be combined and implemented as a method. In addition, the technical features of the method claims of the present disclosure and the technical features of the device claims of the present disclosure can be combined and implemented as a device, and the technical features of the method claims of the present disclosure and the technical features of the device claims of the present disclosure can be combined and implemented as a method.
Claims
1. An image decoding method, comprising: determining a reference block for the current block based on motion information of the current block; generating a prediction sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset; as well as The current block is reconstructed based on the prediction samples.
2. The method according to claim 1, wherein The adjacent sample includes at least one of a left adjacent sample, an upper adjacent sample, a left adjacent sample, or a lower adjacent sample.
3. The method according to claim 1, wherein A filter coefficient of the filter is determined based on at least one sample belonging to a first neighboring region of the current block and at least one sample belonging to a second neighboring region of the reference block.
4. The method according to claim 3, wherein: The second adjacent area includes the template of the reference block and an area extending from the template of the reference block through N sample lines.
5. The method according to claim 4, wherein The template includes at least one of a left adjacent region, an upper adjacent region, an upper left adjacent region, an upper right adjacent region, or a lower left adjacent region.
6. The method according to claim 4, wherein: At least one of a position or a size of the template is determined based on a size of the reference block.
7. The method according to claim 4, wherein: determining the template to be any one of a plurality of template candidates based on template type information signaled through a bitstream, and The template type information indicates at least one of the position or size of the template.
8. The method according to claim 3, wherein: deriving filter coefficients of the filter by using some samples of the second neighboring area belonging to the reference block, and The samples are specified based on a predetermined sampling rate and the predetermined offset.
9. The method according to claim 1, wherein The filter is determined to be any one of a plurality of filters based on a value of the reference sample and a predetermined threshold.
10. The method according to claim 9, wherein: classifying samples within the template of the reference block into a plurality of sample intervals based on the predetermined threshold, and The multiple filters are derived for the multiple sampling intervals respectively.
11. The method according to claim 1, wherein generating the prediction samples based on a plurality of filters including the filter, and The multiple filters are derived from multiple template candidates respectively.
12. An image encoding method, comprising: determining a reference block for inter prediction of a current block; generating a prediction sample of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset; deriving residual samples of the current block based on the prediction samples; as well as The residual samples are encoded.
13. A computer-readable storage medium storing a bit stream generated by the image encoding method according to claim 12.
14. A method for transmitting data, comprising: Obtaining a bitstream for image information, wherein the bitstream is generated by: determining a reference block of a current block used for inter-frame prediction of the current block, generating prediction samples of the current block by applying a filter to at least one of a reference sample belonging to the reference block, at least one neighboring sample adjacent to the reference sample, or a predetermined offset, deriving residual samples of the current block based on the prediction samples, and encoding the residual samples; and Data including the bitstream is transmitted.