Image coding device and method based on filtering-related information signaling
Through an image encoding method based on filter-related information signaling, the zero-order exponential Columbus scheme and adaptive loop filtering (ALF) are used to solve the problem of efficient compression and transmission of high-resolution images/videos, improving encoding efficiency and visual quality, and reducing operational complexity.
Patent Information
- Application Number
- CN202180021157.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-15
- Filing Date
- 2021-01-15
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-01-15
AI Technical Summary
The prior art is difficult to efficiently compress and transmit high-resolution, high-quality image/video data, especially in the transmission of virtual reality and artificial reality content, resulting in increased transmission and storage costs.
The image encoding method based on filter-related information signaling is adopted, and the absolute value information of the brightness/chromaticity ALF filter coefficients is analyzed using the zero-order exponential Columbus scheme, and the image/video encoding efficiency is improved through adaptive loop filtering (ALF).
Improves overall compression efficiency and subjective/objective visual quality of images/video, reduces operational overhead and complexity, while optimizing transmission and storage costs.
Smart Images

Figure CN115280771B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image coding method and apparatus based on filtering related information signaling. Background Art
[0002] Recently, the demand for high-resolution and high-quality images / videos such as 4K or 8K or higher ultra-high definition (UHD) images / videos has been increasing in various fields. As image / video data has high resolution and high quality, the amount of information or bits to be transmitted increases compared to existing image / video data. Therefore, transmitting image data using media such as existing wired / wireless broadband lines or existing storage media or storing image / video data using existing storage media increases the transmission cost and storage cost.
[0003] In addition, the attention and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms are increasing day by day, and the broadcasting of images / videos having characteristics different from real images (such as game images) is also increasing.
[0004] Therefore, highly efficient image / video compression techniques are needed to effectively compress, transmit, store, or reproduce information of high-resolution and high-quality images / videos having various characteristics as described above.
[0005] In addition, techniques such as adaptive loop filtering (ALF) are being discussed to improve compression efficiency and subjective / objective visual quality. To efficiently apply these techniques, a method for efficiently signaling related information is needed. Summary of the Invention
[0006] Technical Solution
[0007] According to an embodiment of this document, a method and apparatus for improving image / video coding efficiency are provided herein.
[0008] According to an embodiment of this document, a method and apparatus for applying efficient filtering are provided herein.
[0009] According to an embodiment of this document, a method and apparatus for efficiently applying adaptive loop filtering (ALF) are provided herein.
[0010] According to an embodiment of this document, a method and apparatus for improving image / video coding efficiency are provided herein.
[0011] According to an embodiment of this document, a method and apparatus for hierarchically signaling ALF related information are provided herein.
[0012] According to an embodiment of this document, the zero-order exponential Golomb scheme (ue(v)) can be used in the parsing process of information / syntax elements related to the absolute value of the luminance / chrominance ALF filter coefficients.
[0013] According to an embodiment of this document, the range of values of the information related to the absolute value of the luminance / chrominance ALF filter coefficients can be fixed.
[0014] According to an embodiment of this document, an encoding device for performing video / image encoding is provided.
[0015] According to an embodiment of this document, a computer-readable digital storage medium is provided, in which encoded video / image information generated according to at least one embodiment disclosed in the video / image encoding method according to the embodiment of this document is stored.
[0016] According to an embodiment of this document, a computer-readable digital storage medium is provided, in which encoded information or encoded video / image information that causes a decoding device to execute the video / image decoding method disclosed in at least one embodiment according to the embodiment of this document is stored.
[0017] Effects of the present invention
[0018] According to an embodiment of this document, the overall compression efficiency of an image / video can be improved.
[0019] According to an embodiment of this document, the subjective / objective visual quality can be improved through efficient filtering.
[0020] According to an embodiment of this document, ALF-related information can be signaled efficiently.
[0021] According to an embodiment of this document, by using the zero-order exponential Golomb scheme (ue(v)) in the parsing process of information / syntax elements related to the absolute value of the luminance / chrominance ALF filter coefficients, the operation (or calculation) overhead and complexity can be reduced.
[0022] According to an embodiment of this document, by fixing the range of values of the information related to the absolute value of the luminance / chrominance ALF filter coefficients, encoding using ue(v) can be performed. Brief description of the drawings
[0023] Figure 1 An example of a video / image encoding system to which the embodiments of the present disclosure can be applied is schematically shown.
[0024] Figure 2 It is a diagram schematically illustrating the configuration of a video / image encoding device to which the embodiments of the present disclosure can be applied.
[0025] Figure 3 FIG. is a diagram schematically illustrating the configuration of a video / image decoding apparatus that can be applied to an embodiment of the present disclosure.
[0026] Figure 4 Exemplarily shows the hierarchical structure of an encoded image / video.
[0027] Figure 5 Shows an example of the ALF filter shape.
[0028] Figure 6 and Figure 7 Respectively show general examples of a video / image encoding method and related components according to an embodiment of the present disclosure.
[0029] Figure 8 and Figure 9 Respectively show general examples of a video / image decoding method and related components according to an embodiment of the present disclosure.
[0030] Figure 10 Shows an example of a content stream transmission system to which an embodiment of the present disclosure can be applied. DETAILED DESCRIPTION
[0031] In this document, video may mean a collection of a series of images over time. A picture generally means a unit representing an image in a specific time region, and a slice / tile is a unit that forms part of a picture in encoding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles.
[0032] A pixel or pel may mean the smallest unit that constitutes a picture (or image). Additionally, "sample" may be used as a term corresponding to a pixel. A sample generally may represent a pixel or the value of a pixel, and may represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Further, a sample may mean a pixel value in the spatial domain, or in the case of transforming such a pixel value into the frequency domain, may mean a transform coefficient in the frequency domain.
[0033] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be related to the Versatile Video Coding (VVC) standard (ITU-T Rec. H.266), the next-generation video / image coding standard after VVC, or other video coding-related standards (e.g., the High Efficiency Video Coding (HEVC) standard (ITU-T Rec. H.265), the Essential Video Coding (EVC) standard, and the AVS2 standard).
[0034] This document presents various embodiments of video / image coding, and the embodiments can be executed in combination with each other unless otherwise specified.
[0035] This document can be modified in various ways and can have various embodiments, and specific embodiments will be illustrated and described in detail in the accompanying drawings. However, this is not intended to limit this document to a specific embodiment. The terms commonly used in this specification are used to describe specific embodiments and do not limit the technical spirit of this document. Singular expressions include plural expressions unless clearly expressed otherwise in the context. Terms such as "including" or "having" in this specification should be understood to indicate the existence of the features, quantities, steps, operations, elements, parts, or combinations thereof described in the specification, and thus should not be construed as excluding the possibility of the existence or addition of one or more different features, quantities, steps, operations, elements, parts, or combinations thereof.
[0036] In addition, each configuration of the drawings described in this document is an independent illustration for explaining the functions of different features from each other, and does not indicate that each configuration is implemented by different hardware or different software. For example, two or more configurations can be combined into one configuration, and one configuration can also be divided into multiple configurations. Embodiments of combining and / or separating configurations are included within the scope of the claims without departing from the gist of this document.
[0037] Hereinafter, the preferred embodiments of this document will be described more specifically with reference to the accompanying drawings. Hereinafter, in the drawings, the same reference numerals are used for the same elements, and repeated descriptions of the same elements are omitted.
[0038] A unit can represent a basic unit of image processing. A unit can include at least one of a specific area of a picture and information related to the area. One unit can include one luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the terms such as unit and terms like block, area, etc. can be used interchangeably. Generally, an M×N block can include a set (or array) of samples (or sample arrays) or transform coefficients composed of M columns and N rows. Alternatively, the sample can be a pixel value in the spatial domain, and when such a pixel value is transformed into the frequency domain, it can mean a transform coefficient in the frequency domain.
[0039] In this document, the terms " / " and "," are interpreted to indicate "and / or". For example, the expression "A / B" is interpreted to indicate "A and / or B", and "A, B" is interpreted to indicate "A and / or B". In addition, "A / B / C" can mean "at least one of A, B, and / or C". In addition, "A, B, C" can mean "at least one of A, B, and / or C".
[0040] In addition, in this document, the term "or" shall be construed as indicating "and / or". For example, the expression "A or B" may mean 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document may mean "additionally or alternatively".
[0041] In this specification, "at least one of A and B" may mean "only A", "only B", or "both A and B". In addition, in this specification, the expression "at least one of A or B" or "at least one of A and / or B" may be construed to be the same as "at least one of A and B".
[0042] In addition, in this specification, "at least one of A, B, and C" may mean "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0043] In addition, the parentheses used in this specification may mean "for example". Specifically, in the case of expressing "prediction (intra prediction)", it may indicate that "intra prediction" is proposed as an example of "prediction". In other words, the term "prediction" in this specification is not limited to "intra prediction", and it may indicate that "intra prediction" is proposed as an example of "prediction". In addition, even in the case of expressing "prediction (i.e., intra prediction)", it may also indicate that "intra prediction" is proposed as an example of "prediction".
[0044] In this specification, the technical features separately explained in one drawing may be implemented separately or may be implemented simultaneously.
[0045] Figure 1 Schematically illustrate an example of a video / image coding system to which the embodiments of this document can be applied.
[0046] Reference Figure 1 ., the video / image coding system may include a source device and a receiving device. The source device may deliver the encoded video / image information or data to the receiving device in the form of a file or a stream via a digital storage medium or a network.
[0047] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0048] The video source can obtain video / images through processes such as capturing, synthesizing, or generating video / images. The video source can include video / image capture devices and / or video / image generation devices. For example, the video / image capture device can include one or more cameras, video / image archives including previously captured video / images, etc. For example, the video / image generation device can include computers, tablet computers, and smart phones, and can (electronically) generate video / images. For example, virtual video / images can be generated by a computer or the like. In this case, the video / image capture process can be replaced by a process of generating relevant data.
[0049] The encoding device can encode the input video / image. For compression and encoding efficiency, the encoding device can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0050] The transmitter can send the encoded image / image information or data output in the form of a bitstream to the receiver of the receiving device in the form of a file or a stream through a digital storage medium or a network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating a media file in a predetermined file format and can include elements for transmission through a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to the decoding device.
[0051] The decoding device can decode the video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.
[0052] The renderer can render the decoded video / image. The rendered video / image can be displayed through a display.
[0053] Figure 2 is a diagram schematically explaining the configuration of the video / image encoding device to which this document is applicable. Hereinafter, the video encoding device can include an image encoding device.
[0054] Reference Figure 2, the encoding device 200 includes an image splitter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the image splitter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or a processor). Additionally, the memory 270 may include a decoded picture buffer (DPB), or may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.
[0055] The image splitter 210 may split an input image (or picture or frame) input to the encoding device 200 into one or more processors. For example, the processors may be referred to as coding units (CUs). In this case, the coding units may be recursively split from a coding tree unit (CTU) or a largest coding unit (LCU) according to a quadtree binary tree ternary tree (QTBTTT) structure. For example, one coding unit may be split into multiple coding units at a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, the quadtree structure may be applied first, and / or the binary tree structure and / or the ternary tree structure may be applied later. Alternatively, the binary tree structure may also be applied first. The encoding process according to the present disclosure may be performed based on the final coding units that are no longer splittable. In this case, based on the encoding efficiency according to the image features, the largest coding unit may be used as the final coding unit, or if necessary, the coding units may be recursively split into coding units at a deeper depth, and the coding unit with the optimal size may be used as the final coding unit. Here, the encoding process may include the processes of prediction, transformation, and reconstruction described later. As another example, the processor may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be partitioned or split from the aforementioned final coding unit. The prediction unit may be a unit for sample prediction, and the transformation unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from the transformation coefficients.
[0056] In some cases, a unit may be used interchangeably with terms such as a block or a region. In general, an M×N block may represent a set of samples or transform coefficients consisting of M columns and N rows. Samples may generally represent pixels or pixel values, which may represent only the luminance component of pixels / pixel values or only the chrominance component of pixels / pixel values. Samples may be used as a term corresponding to pixels or pels configuring a picture (or an image).
[0057] The subtractor 231 may generate a residual signal (residual block, residual samples, or residual sample array) by subtracting the prediction signal (prediction block, prediction samples, or prediction sample array) output from the predictor 220 from the input image signal (original block, original samples, or original sample array), and may send the generated residual signal to the transformer 232. The predictor 220 may perform prediction on a processing target block (hereinafter referred to as “current block”) and may generate a prediction block including prediction samples for the current block. The predictor 220 may determine whether to apply intra prediction or inter prediction in units of the current block or CU. The predictor may generate and transmit various information about the prediction, such as prediction mode information described later in the explanation of each prediction mode, to the entropy encoder 240. The information about the prediction may be encoded by the entropy encoder 240 and output in the form of a bit stream.
[0058] The intra predictor 222 may predict the current block by referring to samples in the current picture. Depending on the prediction mode, the samples referred to may be located near the current block or may be separated. In intra prediction, the prediction mode may include a plurality of non-directional modes and a plurality of directional modes. For example, the non-directional modes may include the DC mode and the planar mode. For example, depending on the level of detail of the prediction direction, the directional modes may include 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes may be used according to the settings. The intra predictor 222 may determine the prediction mode applied to the current block by using the prediction mode applied to neighboring blocks.
[0059] The inter - frame predictor 221 can derive a predicted block of a current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring block can be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block can be referred to as a collocated picture (colPic). For example, the inter - frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter - frame prediction can be performed based on various prediction modes. For example, in the skip mode and the merge mode, the inter - frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. In the skip mode, different from the merge mode, the residual signal may not be transmitted. The motion vector prediction (MVP) mode can use the motion vector of a neighboring block as a motion vector predictor and can indicate the motion vector of the current block by signaling a motion vector difference.
[0060] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra - frame prediction or inter - frame prediction to predict a block, but also apply intra - frame prediction and inter - frame prediction simultaneously. This can be referred to as combined inter - frame and intra - frame prediction (CIIP). In addition, the predictor can perform intra - block copy (IBC) for the prediction of a block. Intra - block copy can be used for content image / video coding such as games, for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter - frame prediction in terms of deriving a reference block in the current picture. That is, IBC can use at least one of the inter - frame prediction techniques described in this document.
[0061] The prediction signal generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditional non-linear transform (CNT). Here, GBT means a transform obtained from a graph when the relationship information between pixels is represented by the graph. CNT means a transform generated based on a prediction signal generated using all previously reconstructed pixels. Further, the transform processing can be applied to a square pixel block of the same size, or can be applied to a block having a variable size other than square.
[0062] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240 and the entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. The quantizer 233 can rearrange the block-based quantized transform coefficients into a one-dimensional vector form based on a coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoder 240 can perform various encoding methods, for example, exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode the information required for video / image reconstruction together with or separately from the quantized transform coefficients (for example, the value of a syntax element, etc.). The encoded information (for example, encoded video / image information) can be sent or stored in the form of a bitstream in units of NAL (network abstraction layer). The video / image information can further include information about various parameter sets, such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Further, the video / image information can further include general constraint information. In this document, the information and / or syntax elements signaled / sent later in this document can be encoded by the above encoding process and can be included in the bitstream. The bitstream can be sent through a network or can be stored in a digital storage medium. The network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as a USB, an SD, a CD, a DVD, a Blu-ray, an HDD, an SSD, etc. A transmitter (not shown) that sends the signal output from the entropy encoder 240 and / or a storage unit (not shown) that stores the signal can be included as an internal / external element of the encoding device 200, and alternatively, the transmitter can be included in the entropy encoder 240.
[0063] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the quantized transform coefficients can be dequantized and inverse-transformed by the dequantizer 234 and the inverse-transformer 235 to reconstruct a residual signal (residual block or residual sample). The adder 250 adds the reconstructed residual signal to the prediction signal output from the predictor 220 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array). If the block to be processed has no residual, such as in the case of applying the skip mode, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next block to be processed in the current picture and can be used for inter prediction of the next picture through filtering as described below.
[0064] Meanwhile, luminance mapping and chrominance scaling (LMCS) can be applied during picture encoding and / or reconstruction.
[0065] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0066] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter predictor 221. When inter prediction is applied by the encoding device, the prediction mismatch between the encoding device 200 and the decoding device 300 can be avoided, and the encoding efficiency can be improved.
[0067] The DPB of the memory 270 can store the modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 can store the motion information of the blocks from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be sent to the inter predictor 221 and used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the reconstructed samples of the reconstructed blocks in the current picture and can pass the reconstructed samples to the intra predictor 222.
[0068] Figure 3 is a diagram schematically explaining the configuration of a video / image decoding device to which this document is applicable.
[0069] ReferenceFigure 3 , the decoding device 300 may include and be configured with an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured by hardware components (e.g., a decoder chipset or a processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0070] When the input is a bitstream including video / image information, the decoding device 300 may reconstruct an image in response to the processing of the video / image information in the Figure 2 encoding device. For example, the decoding device 300 may derive units / blocks based on the block partition-related information obtained from the bitstream. The decoding device 300 may use the processor applied in the encoding device to perform decoding. Thus, the decoding processor may be, for example, an encoding unit, and the encoding unit may be partitioned from the coding tree unit or the largest coding unit according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Additionally, the reconstructed image signal decoded and output by the decoding device 300 may be reproduced by a reproduction device.
[0071] The decoding device 300 may receive, in the form of a bitstream, from Figure 2The signal output by the encoding device, and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information can further include information about various parameter sets, such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information can further include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. The information and / or syntax elements signaled / received in this document can be decoded through the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on coding methods such as exponential Golomb coding, CAVLC, or CABAC, and the output values of the syntax elements required for image reconstruction and the quantization values of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive the bins corresponding to each syntax element in the bitstream, use the decoding target syntax element information, the neighboring and decoded information of the decoding target block, or the information of the symbols / bins decoded in the previous stage to determine the context model, and perform arithmetic decoding on the bins by predicting the probability of the bin occurrence according to the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin after determining the context model. The information about prediction among the information decoded by the entropy decoder 310 can be provided to the predictor 330, and the information about the residuals for which entropy decoding has been performed in the entropy decoder 310, that is, the quantized transform coefficients and the related parameter information, can be input to the dequantizer 321. In addition, the information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiver (not shown) for receiving the signal output by the encoding device can be further configured as an internal / external component of the decoding device 300, or the receiver can be a component of the entropy decoder 310. Meanwhile, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoder 310, and the sample decoder can include at least one of the dequantizer 321, the inverse transformer 322, the predictor 330, the adder 340, the filter 350, and the memory 360.
[0072] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order executed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients by using quantization parameters (e.g., quantization step information) and obtain the transform coefficients.
[0073] The inverse transformer 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0074] The predictor 330 can perform prediction on the current block and generate a prediction block including prediction samples for the current block. The predictor can determine whether to apply intra prediction or inter prediction to the current block based on the information about prediction output from the entropy decoder 310 and can determine a specific intra / inter prediction mode.
[0075] The predictor can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can perform intra block copy (IBC) for prediction of a block. Intra block copy can be used for content image / video coding such as games, for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction because a reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document.
[0076] The intra predictor 332 can predict the current block by referring to samples in the current picture. The samples referred to can be located near the current block or can be located at separate positions according to the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra predictor 332 can determine the prediction mode to be applied to the current block by using the prediction mode applied to neighboring blocks.
[0077] The inter - frame predictor 331 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter - frame prediction mode, the motion information can be predicted in units of blocks, sub - blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter - frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter - frame prediction, neighboring blocks can include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter - frame predictor 331 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter - frame prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter - frame prediction for the current block.
[0078] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (330). If the block to be processed has no residual, such as when the skip mode is applied, the prediction block can be used as the reconstructed block.
[0079] The adder 340 can be referred to as a reconstructor or a reconstructed - block generator. The generated reconstructed signal can be used for intra - frame prediction of the next block to be processed in the current picture, can be output through filtering as described below, or can be used for inter - frame prediction of the next picture.
[0080] Meanwhile, luminance mapping and chrominance scaling (LMCS) can be applied during picture decoding.
[0081] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). Various filtering methods can include, for example, de - blocking filtering, sample - adaptive offset, adaptive loop filter, bilateral filter, etc.
[0082] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter - frame predictor 331. The memory 360 can store the motion information of the blocks from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks in the already reconstructed pictures. The stored motion information can be sent to the inter - frame predictor 331 so as to be used as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and pass the reconstructed samples to the intra - frame predictor 332.
[0083] In this specification, the embodiments explained in the predictor 330, de - quantizer 321, inverse transform device 322, and filter 350 of the decoding device 300 can be applied to or correspond to the predictor 220, de - quantizer 234, inverse transform device 235, and filter 260 of the encoding device 200 in the same way respectively.
[0084] In addition, as described above, in video coding, prediction is performed to improve the compression efficiency. Through this operation, a prediction block including prediction samples for the current block (i.e., the encoding target block) to be encoded can be generated. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block is derived in the same way in both the encoding device and the decoding device, and the encoding device decodes information (residual information) about the residual between the original block and the prediction block (rather than the original sample values of the original block itself). Signaling this to the device can improve the image coding efficiency. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed picture including the reconstructed block.
[0085] The residual information can be generated through transform processing and quantization processing. For example, the encoding device can derive a residual block between the original block and the prediction block, perform transform processing on the residual samples (residual sample array) included in the residual block to derive transform coefficients, and then derive quantized transform coefficients by performing quantization processing on the transform coefficients to signal the residual - related information (via the bitstream) to the decoding device. Here, the residual information can include value information, position information, transform technology, transform core, and quantization parameters of the quantized transform coefficients, etc. The decoding device can perform de - quantization / inverse transform processing based on the residual information and derive residual samples (or a residual block). The decoding device can generate a reconstructed picture based on the prediction block and the residual block. Additionally, for inter - frame prediction reference of subsequent pictures, the encoding device can also de - quantize / inverse - transform the quantized transform coefficients to derive a residual block and generate a reconstructed picture based on this.
[0086] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or, for the sake of consistency in expression, may still be referred to as transform coefficients.
[0087] In this document, the quantized transform coefficients and the transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In such a case, the residual information may include information about the transform coefficients, and the information about the transform coefficients may be signaled using the residual coding syntax. The transform coefficients may be derived based on the residual information (or the information about the transform coefficients), and the scaled transform coefficients may be derived by inverse-transforming (scaling) the transform coefficients. The residual samples may be derived based on the inverse-transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.
[0088] The predictor of an encoding device / decoding device may derive a predicted sample by performing inter-frame prediction in units of blocks. Inter-frame prediction may be prediction derived in a manner that depends on data elements (e.g., sample values or motion information) of pictures other than the current picture. When applying inter-frame prediction to a current block, a predicted block (predicted sample array) for the current block may be derived based on a reference block (reference sample array) specified by a motion vector on a reference picture indicated by a reference picture index. Here, in order to reduce the amount of motion information transmitted in the inter-frame prediction mode, the motion information of the current block may be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may also include inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be referred to as a collocated reference block, a collocated CU (colCU), etc., and the reference picture including the temporal neighboring block may be referred to as a collocated picture (colPic). For example, a motion information candidate list may be configured based on neighboring blocks of the current block, and a flag or index information indicating which candidate is selected (used) may be signaled to derive the motion vector and / or reference picture index of the current block. Inter-frame prediction may be performed based on various prediction modes. For example, in the case of the skip mode and the merge mode, the motion information of the current block may be the same as the motion information of the neighboring block. In the skip mode, unlike the merge mode, a residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected neighboring block may be used as a motion vector predictor, and the motion vector difference may be signaled. In this case, the sum of the motion vector predictor and the motion vector difference may be used to derive the motion vector of the current block.
[0089] Depending on the inter - frame prediction type (L0 prediction, L1 prediction, Bi - prediction, etc.), the motion information may include L0 motion information and / or L1 motion information. The motion vector in the L0 direction may be referred to as the L0 motion vector or MVL0, and the motion vector in the L1 direction may be referred to as the L1 motion vector or MVL1. The prediction based on the L0 motion vector may be referred to as L0 prediction, the prediction based on the L1 motion vector may be referred to as L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector may be referred to as bidirectional prediction. Here, the L0 motion vector may indicate a motion vector associated with the reference picture list L0 (L0), and the L1 motion vector may indicate a motion vector associated with the reference picture list L1 (L1). The reference picture list L0 may include pictures earlier than the current picture in the output order as reference pictures, and the reference picture list L1 may include pictures later than the current picture in the output order. The previous pictures may be referred to as forward (reference) pictures, and the subsequent pictures may be referred to as backward (reference) pictures. The reference picture list L0 may also include pictures later than the current picture in the output order as reference pictures. In this case, the previous pictures may be indexed first in the reference picture list L0, and the subsequent pictures may be indexed later. The reference picture list L1 may also include pictures earlier than the current picture in the output order as reference pictures. In this case, the subsequent pictures may be indexed first in the reference picture list 1, and the previous pictures may be indexed later. The output order may correspond to the picture order count (POC) order.
[0090] Figure 4 Exemplarily shows the hierarchical structure of the encoded image / video.
[0091] Reference Figure 4 , the encoded image / video is divided into a video coding layer (VCL) that processes the image / video and its own decoding process, a subsystem that transmits and stores the encoded information, and a NAL (network abstraction layer) that is responsible for the network adaptation function and exists between the VCL and the subsystem.
[0092] In the VCL, VCL data including compressed image data (slice data) may be generated, or parameter sets including picture parameter sets (PSP), sequence parameter sets (SPS), and video parameter sets (VPS), or supplementary enhancement information (SEI) messages additionally required for the image decoding process may be generated.
[0093] In the NAL, a NAL unit can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in the VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header can include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0094] As shown in the figure, the NAL unit can be divided into a VCL NAL unit and a non-VCL NAL unit according to the RBSP generated in the VCL. The VCL NAL unit can mean a NAL unit including information about an image (slice data), and the non-VCL NAL unit can mean a NAL unit including information required for decoding the image (parameter set or SEI message).
[0095] The aforementioned VCL NAL unit and non-VCL NAL unit can be sent over the network by attaching header information according to the data standard of the subsystem. For example, the NAL unit can be transformed into a data format of a predetermined standard such as the H.266 / VVC file format, Real-Time Transport Protocol (RTP), Transport Stream (TS), etc., and sent over various networks.
[0096] As described above, the NAL unit can be specified using the NAL unit type according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored and signaled in the NAL unit header.
[0097] For example, the NAL unit can be classified into a VCL NAL unit type and a non-VCL NAL unit type according to whether the NAL unit includes information about an image (slice data). The VCL NAL unit type can be classified according to the nature and type of the picture included in the VCL NAL unit, and the non-VCL NAL unit type can be classified according to the type of the parameter set.
[0098] The following are examples of NAL unit types specified according to the type of the parameter set included in the non-VCL NAL unit type.
[0099] - APS (Adaptive Parameter Set) NAL unit: The type for a NAL unit including APS
[0100] - DPS (Decoding Parameter Set) NAL unit: The type for a NAL unit including DPS
[0101] - VPS (Video Parameter Set) NAL unit: The type for a NAL unit including VPS
[0102] - SPS (Sequence Parameter Set) NAL unit: The type of NAL unit used to include SPS
[0103] - PPS (Picture Parameter Set) NAL unit: The type of NAL unit used to include PPS
[0104] - PH (Picture Header) NAL unit: The type of NAL unit used to include PH
[0105] The aforementioned NAL unit types may have syntax information for the NAL unit types and may store and signal the syntax information in the NAL unit header. For example, the syntax information may be nal_unit_type, and the NAL unit type may be specified by the nal_unit_type value.
[0106] Meanwhile, as described above, a picture may include multiple slices, and a slice may include a slice header and slice data. In this case, a picture header may be further added to the multiple slices (slice header and slice data set) in a picture. The picture header (picture header syntax) may include information / parameters that are commonly applicable to the picture. In this document, slices may be mixed or replaced with tile groups. Additionally, in this document, slice headers may be mixed or replaced with tile group headers.
[0107] The slice header (slice header syntax, slice header information) may include information / parameters that can be commonly applied to the slice. APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or pictures. SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. DPS may include information / parameters related to the concatenation of coded video sequences (CVS). The high-level syntax (HLS) in this document may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0108] In this document, the image / image information encoded by the encoding device and signaled to the decoding device in the form of a bitstream includes not only information related to partitions in a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but also information included in the slice header, information included in the picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. Additionally, the image / video information may also include information of the NAL unit header.
[0109] The following table shows a coding descriptor for the parsing process of coding-related information for the present disclosure. The coding descriptor can be used in the parsing process of syntax elements included in the syntax of the present disclosure.
[0110] [Table 1]
[0111]
[0112]
[0113] The following table shows the encoding of x based on zero-order, first-order, second-order, and third-order exponential Golomb coding. For example, x can be a decimal number, and the encoded x can be a binary number. k represents the order of the exponential Golomb coding (where k = 0, 1, 2, 3). According to the above coding descriptor ue(v), the zero-order exponential Golomb coding (where k = 0) can be used for the syntax element parsing process, and for the parsing process of syntax elements based on the coding descriptor ue(v), reference can be made to Table 2 (where k = 0).
[0114] [Table 2]
[0115]
[0116]
[0117]
[0118]
[0119] Meanwhile, in order to compensate for the difference between the original image and the reconstructed image caused by errors occurring during the compression coding process such as quantization, the in-loop filtering process can be performed on the reconstructed samples or the reconstructed picture as described above. As described above, the in-loop filtering can be performed by the filters of the encoding device and the decoding device, and a deblocking filter, SAO, and / or an adaptive loop filter (ALF) can be applied. For example, the ALF process can be performed after the deblocking filtering process and / or the SAO process. However, even in this case, the deblocking filtering process and / or the SAO process can be skipped.
[0120] In the following, picture reconstruction and filtering will be described in detail. In image / video coding, a reconstructed block can be generated based on intra prediction / inter prediction for each block unit, and a reconstructed picture including the reconstructed block can be generated. When the current picture / slice is an I picture / slice, the blocks included in the current picture / slice can be reconstructed only based on intra prediction. Meanwhile, when the current picture / slice is a P or B picture / slice, the blocks included in the current picture / slice can be reconstructed based on intra prediction or inter prediction. In this case, intra prediction can be applied to a part of the blocks in the current picture / slice, and inter prediction can be applied to the remaining blocks.
[0121] Intra prediction can represent a prediction for generating a prediction sample of a current block based on reference samples within a picture (hereinafter referred to as the current picture) to which the current block belongs. When applying intra prediction to the current block, neighboring reference samples to be used for the intra prediction of the current block can be derived. The neighboring reference samples of the current block can include samples adjacent to the left boundary of the current block having a size of nW×nH and a total of 2×nH samples adjacent to the lower left side, samples adjacent to the upper boundary of the current block and a total of 2×nW samples adjacent to the upper right side, and one sample adjacent to the upper left side of the current block. Alternatively, the neighboring reference samples of the current block can further include multiple upper neighboring samples in multiple columns and multiple left neighboring samples in multiple rows. Alternatively, the neighboring reference samples of the current block can further include a total of nH samples adjacent to the right boundary of the current block having a size of nW×nH, a total of nW samples adjacent to the lower boundary of the current block, and one sample adjacent to the lower right side of the current block.
[0122] However, among the neighboring reference samples of the current block, a part of the neighboring reference samples may still not be decoded or available. In this case, the decoder can configure the neighboring reference samples to be used for prediction by replacing the unavailable samples with available samples. Alternatively, the neighboring reference samples to be used for prediction can be configured by interpolation of available samples.
[0123] When deriving neighboring reference samples, a predicted sample can be derived based on the average or interpolation of neighboring reference samples of the current block, and (ii) a predicted sample can be derived based on reference samples among the neighboring reference samples of the current block that exist with respect to a specific (predicted) direction along the predicted sample. The case of (i) can be referred to as a non-directional mode or non-angle mode, and the case of (ii) can be referred to as a directional mode or angle mode. Additionally, based on the predicted sample among the neighboring reference samples of the current block, a predicted sample can be generated by interpolation between a second neighboring sample and a first neighboring sample located in a direction opposite to the prediction direction of the intra prediction mode of the current block. The above case can be referred to as linear interpolation intra prediction (LIP). Additionally, a chrominance prediction sample can be generated based on luminance samples using a linear model. This case can be referred to as the LM mode. Additionally, a temporary predicted sample of the current block can be derived based on filtered neighboring reference samples, and a weighted sum of at least one reference sample (i.e., unfiltered neighboring reference samples) derived according to the intra prediction mode among the existing neighboring reference samples and the temporary predicted sample can be performed to derive the predicted sample of the current block. The above case can be referred to as position-dependent intra prediction (PDPC). Additionally, a reference sample line with the highest prediction accuracy among the neighboring multi-reference sample lines of the current block can be selected, and a predicted sample can be derived by using the reference sample located in the prediction direction on the corresponding line. And, at this time, the reference sample line indication (signaled) used in this document can be given to the decoding device to perform intra prediction coding. The above case can be referred to as multi-reference line (MRL) intra prediction or MRL-based intra prediction. Additionally, the current block can be divided into vertical or horizontal sub-partitions, and then intra prediction can be performed based on the same intra prediction mode, and neighboring reference samples can be derived and used in the sub-partition unit. That is, in this case, the intra prediction mode for the current block is equally applied to the sub-partitions, and in some cases, the intra prediction performance can be improved by deriving and using neighboring reference samples in the sub-partition unit. Such a prediction method can be referred to as intra sub-partition (ISP) or ISP-based intra prediction. The foregoing intra prediction methods can be separately referred to as intra prediction types from the intra prediction modes described in Part 1.2. The intra prediction types can be referred to by various other terms such as intra prediction schemes or additional intra prediction modes. For example, the intra prediction type (or additional intra prediction mode, etc.) can include at least one of the foregoing LIP, PDPC, MRL, and ISP. A general intra prediction method other than a specific intra prediction type such as LIP, PDPC, MRL, or ISP can be referred to as a normal intra prediction type. When a specific intra prediction type cannot be applied, the normal intra prediction type can generally be applied, and prediction can be performed based on the above intra prediction mode. At the same time, post-filtering can be performed on the derived predicted sample as needed.
[0124] Specifically, the intra prediction process may include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and a prediction sample derivation step based on the intra prediction mode / type. Additionally, a post-filtering step may be performed on the derived prediction samples as needed.
[0125] A modified reconstructed picture may be generated through an in-loop filtering process, and the modified reconstructed picture may be output as a decoded picture from the decoding device. Also, the modified reconstructed picture may be stored in the decoded picture buffer or memory of the encoding device / decoding device for later use as a reference picture during the inter prediction process when encoding / decoding a picture. As described above, the in-loop filtering process may include a deblocking filtering process, a sample adaptive offset (SAO) process, and / or an adaptive loop filtering (ALF) process, etc. In this case, among the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filtering (ALF) process, and the bilateral filtering process, one process or some processes may be sequentially applied, or all processes may be sequentially applied. For example, after applying the deblocking filtering process to the reconstructed picture, the SAO process may be performed. Alternatively, for example, after applying the deblocking filtering process to the reconstructed picture, the ALF process may be performed. This may be similarly performed in the encoding device.
[0126] Deblocking filtering is a filtering scheme for removing any distortion occurring at the boundaries between blocks in the reconstructed picture. For example, the deblocking filtering process may derive a target boundary from the reconstructed picture, determine the boundary strength (bS) of the target boundary, and perform deblocking filtering on the target boundary based on bS. The bS may be determined based on the prediction mode of two blocks adjacent to the target boundary, the motion vector difference, whether the reference pictures are the same, the presence of non-zero significant coefficients, etc.
[0127] SAO is a method for compensating the offset difference between the reconstructed picture and the original picture. And, in this document, for example, SAO may be applied based on various types such as Band Offset, Edge Offset, etc. According to SAO, samples may be classified into different categories according to each SAO type, and an offset value may be added to each sample based on the category. The filtering information of SAO may include information on whether SAO is applied or not, SAO type information, SAO offset value information, etc. SAO may also be applied to the reconstructed picture after applying the deblocking filtering to the reconstructed picture.
[0128] Adaptive loop filtering (ALF) is a filtering scheme that is performed on a sample-by-sample basis based on filter coefficients according to the filter shape for reconstructing a picture. An encoding device can determine whether to apply or not apply ALF, the ALF shape, and / or the ALF filtering coefficients, etc. by comparing the reconstructed picture with the original picture, and then, the encoding device can signal the determined result to a decoding device. That is, the filtering information for ALF can include ALF filter shape information, ALF filtering coefficient information, etc. ALF can be applied to the reconstructed picture after applying deblocking filtering.
[0129] Figure 5 An example of the ALF filter shape is shown.
[0130] Figure 5 (a) of shows the shape of a 7x7 diamond filter, Figure 5 (b) of shows the shape of a 5x5 diamond filter. Figure 5 Cn within the shown filter shape represents a filter coefficient. When the value of n in Cn is the same, this indicates that the same filter coefficient can be assigned. In the present disclosure, the position and / or unit where filter coefficients are assigned according to the ALF filter shape may be referred to as a filter tab. In this case, one filter coefficient can be assigned to each filter tab, and the arrangement form of the filter tabs can correspond to the filter shape. The filter tab located at the center of the filter shape can be referred to as the center filter tab. The same filter coefficient can be assigned to two filter tabs existing in positions that are symmetric to each other based on the center filter tab. For example, in the case of a 7x7 diamond filter shape, since it includes 25 filter tabs and the filter coefficients C0 to C11 are assigned in a centrosymmetric structure, only 7 filter coefficients can be used to assign filter coefficients to 13 filter tabs. For example, in order to reduce the amount of data of the information related to the signalized filter coefficients, among the 13 filter coefficients of the 7x7 diamond filter shape, 12 filter coefficients can be (explicitly) signaled, and one filter coefficient can be (implicitly) derived. Additionally, for example, among the 7 filter coefficients of the 5x5 diamond filter shape, 6 coefficients can be (explicitly) signaled, and one filter coefficient can be (implicitly) derived.
[0131] In one example, before applying the filter, a geometric transformation can be applied to the filter coefficients and the corresponding clipping value based on the gradient value calculated for the corresponding block. The geometric transformation can include rotation, diagonal flipping, or vertical flipping.
[0132] The following equations show filter coefficients and clipping values with transformations for each direction (diagonal, vertical, rotation). In the equations shown below, K is the filter size, and k and l represent coefficient coordinates. For example, k can be greater than or equal to 0, and l can be less than or equal to K - 1. The position (0, 0) can be the upper left corner, and the position (K - 1, K - 1) can be the lower right corner. Based on the gradient value calculated for the corresponding block, the transformation can be applied to the filter coefficient f(k, l) and the clipping value c(k, l).
[0133] [Equation 1]
[0134] Diagonal: f_D(k, l) = f(l, k), c_D(k, l) = c(l, k)
[0135] [Equation 2]
[0136] Vertical flip: f_V(k, l) = f(k, K - l - 1), c_V(k, l) = c(k, K - l - 1)
[0137] [Equation 3]
[0138] Rotation: f_R(k, l) = f(K - l - 1, k), c_R(k, l) = c(K - l - 1, k)
[0139] The following table shows an exemplary relationship between the gradient values (g , , , , d2 , , v , Gradient value Transformation <![CDATA[g d2 <g d1 and g h <g v > No transformation <![CDATA[g d2 <g d1 and g v <g h > Diagonal <![CDATA[g d1 <g d2 and g h <g v > Vertical flip <![CDATA[g d1 <g d2 and g v <g h > Rotation , , , ,
[0141] ,
[0136] ,
[0139] , d1 ,
[0140] ,
[0137] , h ,
[0143] ,
[0138] ,
[0144] ,
[0142] ,g v ,g d1 ,g d2 ) and the transformation applied to the current block.
[0140] [Table 3]
[0141] Gradient value Transformation <![CDATA[g d2 <g d1 and g h <g v > No transformation <![CDATA[g d2 <g d1 and g v <g h > Diagonal <![CDATA[g d1 <g d2 and g h <g v > Vertical flip <![CDATA[g d1 <g d2 and g v <g h > Rotation
[0142] To reduce the bit overhead, it is necessary to combine the filter coefficients of different categories of luminance components. The ALF filter parameters can be signaled in the APS and / or slice header. For example, in one APS, up to 25 groups of luminance filter coefficients and clipping value indices can be signaled, and up to 8 groups of chrominance filter coefficients and clipping value indices can be signaled. In the slice header, the index of the APS for the current slice can be signaled.
[0143] The clipping value index decoded from the APS can be used together with the luminance table of clipping values and the chrominance table of clipping values to determine the clipping value. Such a clipping value can be based on the internal bit depth.
[0144] In one example, a luminance table of clipping values and a chrominance table of clipping values can be derived based on the following equations. In the equations shown below, B represents the internal bit depth, and N can be the number of clipping values. For example, N can be equal to 4.
[0145] [Equation 4]
[0146] AlfClipL = {round(2^(B(N - n + 1) / N)) where n ∈ [1..N]}
[0147] [Equation 5]
[0148] AlfClipC = {round(2^((B - 8)+8((N - n)) / (N - 1))) where n ∈ [1..N]}
[0149] As an example, to indicate the luminance filter set used in the current slice, the slice header can signal up to 7 APS indices. The filtering process can be further controlled at the CTB level. A flag indicating whether ALF is applied to the luminance CTB can always be signaled. The luminance CTB can select a filter set from 16 fixed filter sets and the filter sets of the APS. To indicate which filter set is being applied, a filter set index for the CTB can be signaled. The 16 fixed filter sets can be predefined and hard - coded in both the encoder and the decoder.
[0150] In the case of the chrominance configuration element, the APS index can be signaled in the slice header to indicate the chrominance filter set used for the current slice. At the CTB level, when two or more chrominance filter sets are present in the APS, a filter index can be signaled for each chrominance CTB.
[0151] To further limit (or constrain) the multiplication complexity, bitstream conformance is applied, allowing coefficient values at non - central positions to be in the range of 0 to 2^8, and allowing coefficient values at the remaining positions to be in the range of - 2^7 to 2^7 - 1. The central position coefficient is not signaled in the bitstream and can be inferred to be equal to 128.
[0152] When the ALF is available for the current CTB, each R(i,j) within the CU can be filtered so that R′(i,j) can be calculated. For example, R′(i,j) can be calculated based on the following equation. f(k,l) can be filter coefficients, and K(x,y) can be a clipping function. Additionally, c(k,l) can be the decoded clipping parameter. k and l can vary from -L / 2 to L / 2, and in this document, L can be the filter length. The clipping function K(x,y) = min(y,max(-y,x)) can also be expressed as Clip3(-y,y,x).
[0153] [Equation 6]
[0154] R′(i,j) = R(i,j) + ((∑ k≠0 ∑ l≠0 f(k,l) × K(R(i + k,j + l) - R(i,j),c(k,l)) + 64) >> 7)
[0155] As described above, the in-loop filtering process can be applied to the reconstructed picture. In this case, a virtual boundary is defined to improve the subjective / objective visual quality of the reconstructed picture, and the in-loop filtering process can be applied across the virtual boundary. The virtual boundary can include, for example, discontinuous edges, such as 360-degree images, VR images, or picture-in-picture (PIP), etc. For example, the virtual boundary can be present at a predetermined position, or its presence and / or its position can be signaled. For example, the virtual boundary can be located in the first 4 sample lines of the CTU row (more specifically, for example, in the upper part of the first 4 sample lines of the CTU row). As another example, information related to the presence and / or position of the virtual boundary can be signaled via HLS. As described above, HLS can include SPS, PPS, picture headers, slice headers, etc.
[0156] Hereinafter, the high-level syntax signaling and semantics related to the embodiments of this specification will be described.
[0157] The embodiments of this specification can include a method for controlling a loop filter. The method for controlling a loop filter can be applied to the reconstructed picture. The in-loop filter (loop filter) can be used to decode the encoded bitstream. The loop filter can include the above-mentioned deblocking, SAO, and ALF. The SPS can include flags related to each of deblocking, SAO, and ALF. These flags can indicate whether each tool is available for encoding the coded layer video sequence (CLVS) and the coded video sequence (CVS) of the reference SPS.
[0158] In one example, when a loop filter is available for encoding pictures within a CVS, the application of the loop filter can be controlled such that the loop filter is not applied across a particular boundary. For example, the loop filter can be controlled not to cross a sub-picture boundary, the loop filter can be controlled not to cross a tile boundary, the loop filter can be controlled not to cross a slice boundary, and / or the loop filter can be controlled not to cross a virtual boundary.
[0159] Information related to in-loop filtering can include the information, syntax, syntax elements, and / or semantics described in this specification (or the embodiments included in this specification). Information related to in-loop filtering can include information related to whether the in-loop filtering process is (fully or partially) available for use across a particular boundary (e.g., a virtual boundary, a sub-picture boundary, a slice boundary, and / or a tile boundary). The image information included in the bitstream can include High-Level Syntax (HLS), and the HLS can include information related to loop filtering. Modified (or filtered) reconstructed samples (reconstructed pictures) can be generated based on a determination of whether in-loop filtering is applied across a particular boundary. In one example, if the in-loop filtering process is disabled for all blocks / boundaries, the modified reconstructed samples may be the same as the reconstructed samples. In another example, the modified reconstructed samples can include modified reconstructed samples derived based on in-loop filtering. However, in this case, based on the determined result, among the reconstructed samples, a part (e.g., the reconstructed samples across the virtual boundary) can be not processed by in-loop filtering. For example, although the reconstructed samples across a particular boundary (including at least one of the virtual boundary, sub-picture boundary, slice boundary, and / or tile boundary that enables in-loop filtering to be performed) can be processed within in-loop filtering, the reconstructed samples across other boundaries (including at least one of the virtual boundary, sub-picture boundary, slice boundary, and / or tile boundary that is disabled to perform in-loop filtering) can be not processed within in-loop filtering.
[0160] In one example, regarding whether to perform the in-loop filtering process across a virtual boundary, the information related to in-loop filtering can include the SPS virtual boundary presence flag, the picture header virtual boundary presence flag, the information related to the number of virtual boundaries, the information about the position of the virtual boundary, etc.
[0161] In the embodiments included in this specification, the information related to the position of the virtual boundary may include information about the x - coordinate of the vertical virtual boundary and / or the y - coordinate of the horizontal virtual boundary. More specifically, the information related to the position of the virtual boundary may include the x - coordinate of the vertical virtual boundary in units of luminance samples and / or the y - coordinate of the horizontal virtual boundary in units of luminance samples. Additionally, the information related to the position of the virtual boundary may include information about the number of information (syntax elements) related to the x - coordinate of the vertical virtual boundary present in the SPS. Additionally, the information related to the position of the virtual boundary may include information about the number of information (syntax elements) related to the y - coordinate of the horizontal virtual boundary present in the SPS. Alternatively, the information related to the position of the virtual boundary may include information about the number of information (syntax elements) related to the x - coordinate of the vertical virtual boundary present in the picture header. Additionally, the information related to the position of the virtual boundary may include information about the number of information (syntax elements) related to the y - coordinate of the horizontal virtual boundary present in the picture header.
[0162] The following table shows the exemplary syntax and semantics of the sequence parameter set (SPS) according to this embodiment.
[0163] [Table 4]
[0164]
[0165] [Table 5]
[0166]
[0167]
[0168]
[0169] The following table shows the exemplary syntax and semantics of the picture parameter set (PPS) according to this embodiment.
[0170] [Table 6]
[0171]
[0172] [Table 7]
[0173]
[0174]
[0175] The following table shows the exemplary syntax and semantics of the picture header according to this embodiment.
[0176] [Table 8]
[0177]
[0178]
[0179] [Table 9]
[0180]
[0181]
[0182]
[0183]
[0184]
[0185] The following table shows the exemplary syntax and semantics of the slice header according to the present embodiment.
[0186] [Table 10]
[0187]
[0188] [Table 11]
[0189]
[0190]
[0191]
[0192] In the following, the signaling of information related to the ALF filter coefficients will be described.
[0193] In a conventional ALF process, the k-th order exponential Golomb code (where k = 3) is used to signal the absolute values of the luminance and chrominance ALF coefficients. However, the disadvantage of the k-th order exponential Golomb coding is that it causes a considerable amount of operation overhead and complexity.
[0194] The embodiments to be described in the following paragraphs can propose solutions for solving the above problems. The embodiments can be applied independently. Alternatively, at least two or more embodiments can be applied in combination.
[0195] The following table shows the exemplary syntax of the adaptive parameter set (APS) according to the embodiments of the present specification.
[0196] [Table 12]
[0197]
[0198] The following table shows the exemplary syntax of the ALF data according to the present embodiment.
[0199] [Table 13]
[0200]
[0201] The following table shows exemplary semantics related to the syntax elements included in the syntax.
[0202] [Table 14]
[0203]
[0204]
[0205]
[0206] According to another embodiment of the present specification, information (alf_luma_coeff_abs[sfIdx][j], alf_chroma_coeff_abs[altIdx][j]) regarding the absolute values of the luminance / chrominance ALF filter coefficients can be parsed based on the zero-order exponential Golomb coding scheme (ue(v)).
[0207] The following table shows an exemplary syntax of the ALF data according to this embodiment.
[0208] [Table 15]
[0209]
[0210] The following table shows exemplary semantics related to the syntax elements included in the syntax.
[0211] [Table 16]
[0212]
[0213]
[0214]
[0215]
[0216] According to the embodiment of the present specification described with reference to the above tables, by applying the zero-order exponential Golomb coding scheme (ue(v)) to the parsing process of information related to the absolute values of the luminance / chrominance ALF filter coefficients (alf_luma_coeff_abs[sfIdx][j], alf_chroma_coeff_abs[altIdx][j]), the operation overhead and complexity can be reduced. Additionally, by fixing the value range of the information related to the absolute values of the luminance / chrominance ALF filter coefficients (e.g., 0 to 128), the coding using (ue(v)) can be efficiently performed.
[0217] Figure 6 and Figure 7 respectively show general examples of a video / image encoding method and related components according to an embodiment of the present disclosure.
[0218] Figure 6 The method disclosed in Figure 2 or Figure 7 can be executed by the encoding device shown. More specifically, for example, Figure 6 S600 and S610 of Figure 7 can be executed by the predictor 220 of the encoding device of Figure 6 S620 to S640 of Figure 7 can be executed by the residual processor 230 of the encoding device of Figure 6 S650 of Figure 7 can be executed by the adder 250 of the encoding device of Figure 6 S660 and / or S670 of Figure 7 can be executed by the filter 260 of the encoding device of Figure 6 S680 of Figure 7 can be executed by the entropy encoder 240 of the encoding device. Additionally, although not shown in Figure 6 , the predictor 220 of the encoding device can derive a prediction sample or prediction-related information, and the entropy encoder 240 of the encoding device can generate a bitstream based on the residual information or prediction-related information. Figure 6 The method disclosed in
[0219] can include the above-described embodiments in this specification. Figure 6 Referring to
[0220] , the encoding device can derive a prediction sample (S600). The encoding device can derive a prediction sample of the current block based on a prediction mode. The encoding device can derive a prediction sample of the current block based on a prediction mode. In this case, various prediction methods disclosed in this specification can be applied, for example, inter prediction or intra prediction.
[0221] The encoding device can derive residual samples (S620). The encoding device can derive the residual samples of the current block and can derive the residual samples of the current block based on the original samples and the predicted samples of the current block. More specifically, the encoding device can derive the predicted samples of the current block based on the prediction mode. In this case, various prediction methods disclosed in this specification can be applied, such as inter-frame prediction or intra-frame prediction, etc. The residual samples can be derived based on the predicted samples and the original samples. For inter-frame prediction, the encoding device can derive at least one reference picture and can perform inter-frame prediction based on the at least one reference picture. The predicted samples can be generated based on the inter-frame prediction. The encoding device can generate reference picture related information based on the at least one reference picture.
[0222] The encoding device can derive transform coefficients (S630). The encoding device can derive transform coefficients based on the transformation process of the residual samples. For example, the transformation process can include at least one of DCT, DST, GBT, or CNT.
[0223] The encoding device can derive quantized transform coefficients. The encoding device can derive quantized transform coefficients based on the quantization process of the transform coefficients. The quantized transform coefficients can have a one-dimensional vector form based on the coefficient scanning order.
[0224] The encoding device can generate residual information (S640). The encoding device can generate residual information based on the residual samples of the current block. The encoding device can generate residual information indicating the quantized transform coefficients. The residual information can be generated by various coding methods, such as exponential Golomb, CAVLC, CABAC, etc.
[0225] The encoding device can generate reconstructed samples (S650). The encoding device can generate reconstructed samples based on the residual information. The reconstructed samples can be generated by adding the residual samples based on the residual information and the predicted samples. More specifically, the encoding device can perform prediction (intra-frame or inter-frame prediction) on the current block and then can generate reconstructed samples based on the predicted samples, which are generated from the prediction using the original samples.
[0226] The reconstructed samples can include reconstructed luminance samples and reconstructed chrominance samples. More specifically, the residual samples can include residual luminance samples and residual chrominance samples. The residual luminance samples can be generated based on the original luminance samples and the predicted luminance samples. The residual chrominance samples can be generated based on the original chrominance samples and the predicted chrominance samples. The encoding device can derive transform coefficients for the residual luminance samples (luminance transform coefficients) and / or derive transform coefficients for the residual chrominance samples (chrominance transform coefficients). The quantized transform coefficients can include quantized luminance transform coefficients and / or quantized chrominance transform coefficients.
[0227] The encoding device can derive ALF filter coefficients (S660). The ALF filter coefficients can include luminance ALF filter coefficients and / or chrominance ALF filter coefficients. Also, modified reconstructed samples can be generated based on the ALF filter coefficients.
[0228] The encoding device can generate ALF-related information (S670). The ALF-related information can include information related to the ALF filter coefficients, information related to ALF-related clipping, etc. Additionally, the ALF-related information can include information related to the availability flag, information related to the presence flag for specifying the location of the high-level syntax (e.g., PPS, SPS, APS, picture header, etc.). For example, the information related to the ALF filter coefficients can include information related to the absolute value of the ALF filter coefficients and information related to the sign of the ALF filter coefficients.
[0229] The encoding device can encode video / image information (S680). The image information can include residual information, prediction-related information, reference picture-related information, sub-picture-related information, in-loop filtering-related information, and / or virtual boundary-related information (and / or additional virtual boundary-related information). The encoded video / image information can be output in bitstream format. The bitstream can be transmitted to the decoding device via a network or a storage medium.
[0230] The video / image information can include various information according to the embodiments of this specification. For example, the video / image information can include the information disclosed in at least one of Tables 1 to 16 presented above.
[0231] According to an embodiment, the image information can include information about the absolute value of the luminance filter coefficients for the ALF process and information about the absolute value of the chrominance filter coefficients for the ALF process. The absolute value of the luminance filter coefficients for the ALF process and the absolute value of the chrominance filter coefficients for the ALF process can be within a predetermined range.
[0232] According to an embodiment, the image information can include a parameter set and ALF data. At least one of the parameter sets can include an Adaptive Parameter Set (APS). The ALF data can include information about the absolute value of the luminance filter coefficients for the ALF process and information about the absolute value of the chrominance filter coefficients for the ALF process. The ALF data can be included in the APS.
[0233] According to an embodiment, the picture information may include header information and an Adaptive Parameter Set (APS) related to ALF. The header information may include information related to the number of ALF-related APS IDs. The number of ALF-related APS IDs may be derived based on the value of the information related to the number of ALF-related APS IDs. The number of ALF-related APS ID syntax elements is equal to the number of ALF-related APS IDs that may be included in the header information.
[0234] According to an embodiment, the picture information may include header information and an Adaptive Parameter Set (APS) related to ALF. The header information may include an ALF availability flag indicating the availability of ALF within a picture or slice and information related to the number of ALF-related APS IDs. When the value of the ALF availability flag is equal to 1, the header information may include information related to the number of ALF-related APS IDs. The value of the information related to the number of ALF-related APS IDs plus 1 may be the same as the number of ALF-related APS IDs.
[0235] According to an embodiment, the absolute value of the luminance filter coefficients for the ALF process may be within a predetermined range.
[0236] According to an embodiment, the absolute value of the chrominance filter coefficients for the ALF process may be within a predetermined range.
[0237] According to an embodiment, the predetermined range may be a range from 0 to 128.
[0238] Figure 8 and Figure 9 respectively show general examples of a video / image decoding method and related components according to an embodiment of the present disclosure.
[0239] Figure 8 The method disclosed in Figure 3 or Figure 9 may be performed by the decoding device shown. More specifically, for example, Figure 8 S800 of Figure 8 may be performed by the entropy decoder 310 of the decoding device, S810 and S820 may be performed by the residual processor 320 of the decoding device, S830 may be performed by the predictor 330 of the decoding device, S840 may be performed by the adder 340 of the decoding device, and S850 and / or S860 may be performed by the filter 350 of the decoding device.
[0240] Refer to Figure 8, the decoding device can receive / acquire video / image information (S800). The video / image information may include residual information, prediction-related information, reference picture-related information, sub-picture-related information, in-loop filtering-related information, and / or ALF-related information. The decoding device can receive / acquire the video / image information through a bitstream.
[0241] The video / image information may include various information according to the embodiments of this specification. For example, the video / image information may include the information disclosed in at least one of Tables 1 to 16 presented above.
[0242] The decoding device can derive quantized transform coefficients. The decoding device can derive quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on the coefficient scan order. The quantized transform coefficients may include quantized luminance transform coefficients and / or quantized chrominance transform coefficients.
[0243] The decoding device can derive transform coefficients (S810). The decoding device can derive transform coefficients based on the dequantization process of the quantized transform coefficients. The decoding device can derive luminance transform coefficients by dequantization based on the quantized luminance transform coefficients. The decoding device can derive chrominance transform coefficients by dequantization based on the quantized chrominance transform coefficients.
[0244] The decoding device can generate / derive residual samples (S820). The decoding device can derive residual samples based on the inverse transform process of the transform coefficients. The decoding device can derive residual luminance samples by the inverse transform process based on the residual luminance samples. The decoding device can derive residual chrominance samples by the inverse transform process based on the residual chrominance samples.
[0245] The decoding device can derive at least one reference picture based on the reference picture-related information. The decoding device can perform a prediction process on the at least one reference picture.
[0246] The decoding device can generate a prediction sample for the current block based on the prediction mode (S830). In this case, various prediction methods disclosed in this specification can be applied, such as inter-frame prediction or intra-frame prediction, etc. The decoding device can generate a prediction sample for the current block within the current picture based on the prediction process. For example, the decoding device can perform an inter-frame prediction process based on the at least one reference picture and can generate a prediction sample based on the inter-frame prediction process.
[0247] The decoding device may generate / derive reconstructed samples (S840). For example, the decoding device may generate / derive reconstructed luma samples and / or reconstructed chroma samples. The decoding device may generate the reconstructed luma samples and / or the reconstructed chroma samples based on the residual information. The decoding device may generate the reconstructed samples based on the residual information. The reconstructed samples may include the reconstructed luma samples and / or the reconstructed chroma samples. The luma component of the reconstructed samples may correspond to the reconstructed luma samples, and the chroma component of the reconstructed samples may correspond to the reconstructed chroma samples. The decoding device may generate predicted luma samples and / or predicted chroma samples through a prediction process. The decoding device may generate the reconstructed luma samples based on the predicted luma samples and the residual luma samples. The decoding device may generate the reconstructed chroma samples based on the predicted chroma samples and the residual chroma samples.
[0248] The decoding device may derive ALF filter coefficients (S850). The ALF filter coefficients may include luma ALF filter coefficients and / or chroma ALF filter coefficients. The ALF filter coefficients may be derived based on the information about the absolute values of the luma filter coefficients and the information about the absolute values of the chroma filter coefficients.
[0249] The decoding device may generate modified (filtered) reconstructed samples (S860). The decoding device may generate the modified reconstructed samples based on an in-loop filtering process for the reconstructed samples. The decoding device may generate the modified reconstructed samples based on the in-loop filtering related information. The decoding device may use a deblocking process, a SAO process, and / or an ALF process to generate the modified reconstructed samples.
[0250] According to an embodiment, the image information may include the information about the absolute values of the luma filter coefficients for the ALF process and the information about the absolute values of the chroma filter coefficients for the ALF process. The absolute values of the luma filter coefficients for the ALF process and the absolute values of the chroma filter coefficients for the ALF process may be within a predetermined range.
[0251] According to an embodiment, the image information may include a parameter set and ALF data. At least one of the parameter sets may include an adaptive parameter set (APS). The ALF data may include the information about the absolute values of the luma filter coefficients for the ALF process and the information about the absolute values of the chroma filter coefficients for the ALF process. The ALF data may be included in the APS.
[0252] According to an embodiment, the picture information may include header information and an ALF-related Adaptive Parameter Set (APS). The header information may include information related to the number of ALF-related APS IDs. The number of ALF-related APS IDs may be derived based on the value of the information related to the number of ALF-related APS IDs. The number of ALF-related APS ID syntax elements is equal to the number of ALF-related APS IDs that may be included in the header information.
[0253] According to an embodiment, the picture information may include header information and an ALF-related Adaptive Parameter Set (APS). The header information may include an ALF availability flag indicating the availability of ALF within a picture or slice and information related to the number of ALF-related APS IDs. When the value of the ALF availability flag is equal to 1, the header information may include information related to the number of ALF-related APS IDs. The value of the information related to the number of ALF-related APS IDs plus 1 may be the same as the number of ALF-related APS IDs.
[0254] According to an embodiment, the absolute value of the luminance filter coefficients for the ALF process may be within a predetermined range.
[0255] According to an embodiment, the absolute value of the chrominance filter coefficients for the ALF process may be within a predetermined range.
[0256] According to an embodiment, the predetermined range may be a range from 0 to 128.
[0257] In the presence of residual samples for a current block, the decoding device may receive information regarding the residual for the current block. The information regarding the residual may include transform coefficients for the residual samples. The decoding device may derive the residual samples (or an array of residual samples) for the current block based on the residual information. Specifically, the decoding device may derive quantized transform coefficients based on the residual information. The quantized transform coefficients may have a one-dimensional vector form based on the coefficient scan order. The decoding device may derive transform coefficients based on an inverse quantization process for the quantized transform coefficients. The decoding device may derive residual samples based on the transform coefficients.
[0258] The decoding device may generate reconstructed samples based on (intra) prediction samples and residual samples, and may derive a reconstructed block or a reconstructed picture based on the reconstructed samples. Specifically, the decoding device may generate reconstructed samples based on the sum of the (intra) prediction samples and the residual samples. Thereafter, as described above, if necessary, the decoding device may apply loop filter processing (e.g., deblocking filtering and / or SAO processing) to the reconstructed picture to improve the subjective / objective picture quality.
[0259] For example, the decoding device can obtain image information including all or some of the above information (or syntax elements) by decoding a bitstream or encoded information. In addition, the bitstream or encoded information can be stored in a computer-readable storage medium, or the above decoding method can be executed.
[0260] In the above embodiments, the method is described based on a flowchart having a series of steps or blocks. The present disclosure is not limited to the order of the above steps or blocks. Some steps or blocks can occur simultaneously with other steps or blocks as described above or in an order different from other steps or blocks as described above. In addition, those skilled in the art will understand that the steps shown in the above flowchart are not exclusive and can include additional steps, or one or more steps in the flowchart can be deleted without affecting the scope of this document.
[0261] The method according to the above embodiments of this document can be implemented in software form, and the encoding device and / or decoding device according to this document can be included, for example, in devices that perform image processing such as TVs, computers, smartphones, set-top boxes, and display devices.
[0262] When the embodiments in this document are implemented in software, the above method can be implemented as a module (processing, function, etc.) that performs the above functions. The module can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), other chip sets, logic circuits, and / or data processing devices. The memory can include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be implemented and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information about the instructions or algorithms for the implementation can be stored in a digital storage medium.
[0263] In addition, the decoding apparatus and the encoding apparatus applying this document may be included in a multimedia broadcast transmission / reception apparatus, a mobile communication terminal, a home theater video apparatus, a digital cinema video apparatus, a surveillance camera, a video chat apparatus, a real-time communication apparatus such as video communication, a mobile streaming apparatus, a storage medium, a camera, a VoD service providing apparatus, an over-the-top (OTT) video apparatus, an Internet streaming service providing apparatus, a three-dimensional (3D) video apparatus, a virtual reality (VR) apparatus, an augmented reality (AR) apparatus, a videoconference video apparatus, a transportation user equipment (i.e., in-vehicle (including autonomous vehicle) user equipment, aircraft user equipment, ship user equipment, etc.), and a medical video apparatus, and may be used to process video signals and data signals. For example, an over-the-top (OTT) video apparatus may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smart phone, a tablet computer, a digital video recorder (DVR), etc.
[0264] In addition, the processing method applying this document may be generated in the form of a program executable by a computer and may be stored in a computer-readable recording medium. Multimedia data having a data structure according to the present disclosure may also be stored in the computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices storing data readable by a computer system. For example, the computer-readable recording medium may include a BD, a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (i.e., transmission through the Internet). In addition, a bitstream generated by an encoding method may be stored in the computer-readable recording medium or may be transmitted through a wired or wireless communication network.
[0265] In addition, embodiments of this document may be implemented using a computer program product according to program code, and the program code may be executed in a computer according to the embodiments of this document. The program code may be stored on a computer-readable carrier.
[0266] Figure 10 An example of a content streaming system to which embodiments of this document may be applied is shown.
[0267] Reference Figure 10 , a content streaming system to which embodiments of this document are applied may generally include an encoding server, a streaming server, a web server, a media storage unit, a user equipment, and a multimedia input device.
[0268] The encoding server is used to compress the content input from a multimedia input device (e.g., a smart phone, a camera, a video camera, etc.) into digital data to generate a bitstream, and send the bitstream to the streaming server. As another example, when a multimedia input device such as a smart phone, a camera, a video camera, etc. directly generates a bitstream, the encoding server can be omitted.
[0269] A bitstream can be generated by applying the encoding method or the bitstream generation method of the embodiments of the present disclosure, and the streaming server can temporarily store the bitstream in the process of sending or receiving the bitstream.
[0270] The streaming server sends multimedia data to the user device via a web server based on the user's request, and the web server serves as a medium for notifying the user of the service. When the user requests a desired service from the web server, the web server passes the request to the streaming server, and the streaming server sends the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between the devices in the content streaming system.
[0271] The streaming server can receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0272] Examples of the user device may include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smart watch, smart glasses, a head-mounted display), a digital TV, a desktop computer, a digital sign, etc.
[0273] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received from each server can be processed in a distributed manner.
[0274] The claims described herein can be combined in various ways. For example, the technical features of the method claims in this document can be combined and implemented as a device, and the technical features of the device claims in this document can be combined and implemented as a method. Additionally, the technical features of the method claims in this document and the technical features of the device claims in this document can be combined to be implemented as a device, and the technical features of the method claims in this document and the technical features of the device claims in this document can be combined and implemented as a method.
Claims
1. An image decoding method performed by a decoding apparatus, the image decoding method comprising the following steps: Obtaining image information including prediction mode information and residual information from a bitstream; Deriving transform coefficients based on the residual information; Deriving residual samples based on the transform coefficients; Deriving prediction samples based on the prediction mode information; Generating reconstructed samples based on the prediction samples and the residual samples; Deriving filter coefficients for an adaptive loop filtering (ALF) process for the reconstructed samples; And Generating modified reconstructed samples based on the reconstructed samples and the filter coefficients, wherein the image information includes information related to the absolute value of the luminance filter coefficients for the ALF process and information related to the absolute value of the chrominance filter coefficients for the ALF process, wherein the information related to the absolute value of the luminance filter coefficients for the ALF process includes syntax elements for the absolute value of the luminance filter coefficients, wherein the information related to the absolute value of the chrominance filter coefficients for the ALF process includes syntax elements for the absolute value of the chrominance filter coefficients, wherein the syntax elements for the absolute value of the luminance filter coefficients and the syntax elements for the absolute value of the chrominance filter coefficients are encoded based on zero - order exponential Golomb coding, and wherein the values of the syntax elements for the absolute value of the luminance filter coefficients and the values of the syntax elements for the absolute value of the chrominance filter coefficients are within a predetermined range.
2. The image decoding method according to claim 1, wherein, The image information includes a parameter set and ALF data, wherein at least one of the parameter sets includes an adaptive parameter set (APS), wherein the ALF data includes information related to the absolute value of the luminance filter coefficients for the ALF process and information related to the absolute value of the chrominance filter coefficients for the ALF process, and wherein the ALF data is included in the APS.
3. The image decoding method according to claim 1, wherein, The image information includes header information and an ALF - related adaptive parameter set (APS), wherein the header information includes information related to the number of ALF - related APS IDs, wherein the number of ALF - related APS IDs is derived based on the value of the information related to the number of ALF - related APS IDs, and wherein the number of syntax elements of the ALF - related APS ID is equal to the number of ALF - related APS IDs included in the header information.
4. The image decoding method according to claim 1, wherein, The image information includes header information and an ALF - related adaptive parameter set (APS), wherein the header information includes an ALF availability flag indicating whether the ALF is available for use within a picture or slice and information related to the number of ALF - related APS IDs, wherein when the value of the ALF availability flag is equal to 1, the header information includes the information related to the number of ALF - related APS IDs, and wherein the value of the information related to the number of ALF - related APS IDs is equal to the number of ALF - related APS IDs.
5. The image decoding method according to claim 1, wherein, The predetermined range is from 0 to 128, including 0 and 128.
6. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Derive prediction samples for a current block; Generate prediction mode information based on the prediction samples; Derive residual samples for the current block; Derive transform coefficients based on the residual samples; Generate residual information based on the transform coefficients; Generate reconstructed samples based on the residual information; Derive filter coefficients for an Adaptive Loop Filter (ALF) process for the reconstructed samples; Generate ALF-related information based on the filter coefficients; And Encode image information including the residual information and the ALF-related information, wherein the image information includes information related to the absolute values of the luminance filter coefficients used for the ALF process and information related to the absolute values of the chrominance filter coefficients used for the ALF process, wherein the information related to the absolute values of the luminance filter coefficients used for the ALF process includes syntax elements for the absolute values of the luminance filter coefficients, wherein the information related to the absolute values of the chrominance filter coefficients used for the ALF process includes syntax elements for the absolute values of the chrominance filter coefficients, wherein the syntax elements for the absolute values of the luminance filter coefficients and the syntax elements for the absolute values of the chrominance filter coefficients are encoded based on zero-order exponential Golomb coding, and wherein the values of the syntax elements for the absolute values of the luminance filter coefficients and the values of the syntax elements for the absolute values of the chrominance filter coefficients are within a predetermined range.
7. The image encoding method according to claim 6, wherein, The image information includes a parameter set and ALF data, wherein at least one of the parameter sets includes an Adaptive Parameter Set (APS), wherein the ALF data includes information related to the absolute values of the luminance filter coefficients used for the ALF process and information related to the absolute values of the chrominance filter coefficients used for the ALF process, and wherein the ALF data is included in the APS.
8. The image encoding method according to claim 6, wherein, The image information includes header information and an ALF-related Adaptive Parameter Set (APS), wherein the header information includes information related to the number of ALF-related APS IDs, wherein the number of ALF-related APS IDs is derived based on the value of the information related to the number of ALF-related APS IDs, and wherein the number of syntax elements of the ALF-related APS IDs is equal to the number of ALF-related APS IDs included in the header information.
9. The image encoding method according to claim 6, wherein, The image information includes header information and an ALF-related Adaptive Parameter Set (APS), wherein the header information includes an ALF availability flag indicating whether ALF is available for use within a picture or slice and information related to the number of ALF-related APS IDs, wherein when the value of the ALF availability flag is equal to 1, the header information includes the information related to the number of ALF-related APS IDs, and wherein the value of the information related to the number of ALF-related APS IDs is equal to the number of ALF-related APS IDs.
10. The image encoding method according to claim 6, wherein, The predetermined range is from 0 to 128, inclusive of 0 and 128.
11. A method for transmitting a bitstream, the method comprising the steps of: Deriving prediction samples for a current block; Generating prediction mode information based on the prediction samples; Deriving residual samples for the current block; Deriving transform coefficients based on the residual samples; Generating residual information based on the transform coefficients; Generating reconstructed samples based on the residual information; Deriving filter coefficients for an adaptive loop filtering (ALF) process for the reconstructed samples; Generating ALF-related information based on the filter coefficients; Encoding image information to generate the bitstream, wherein the image information includes the residual information and the ALF-related information; and Transmitting the bitstream, wherein the image information includes information related to the absolute values of the luminance filter coefficients for the ALF process and information related to the absolute values of the chrominance filter coefficients for the ALF process, wherein the information related to the absolute values of the luminance filter coefficients for the ALF process includes syntax elements for the absolute values of the luminance filter coefficients, wherein the information related to the absolute values of the chrominance filter coefficients for the ALF process includes syntax elements for the absolute values of the chrominance filter coefficients, wherein the syntax elements for the absolute values of the luminance filter coefficients and the syntax elements for the absolute values of the chrominance filter coefficients are encoded based on Golomb-Rice coding of order zero, and wherein the values of the syntax elements for the absolute values of the luminance filter coefficients and the values of the syntax elements for the absolute values of the chrominance filter coefficients are within a predetermined range.
12. A method for transmitting data of an image, the transmitting method comprising the steps of: Obtaining a bitstream of the image, wherein the bitstream is generated based on the steps of: deriving prediction samples for a current block, generating prediction mode information based on the prediction samples, deriving residual samples for the current block, deriving transform coefficients based on the residual samples, generating residual information based on the transform coefficients, generating reconstructed samples based on the residual information, deriving filter coefficients for an adaptive loop filtering (ALF) process for the reconstructed samples, generating ALF-related information based on the filter coefficients, and encoding image information including the residual information and the ALF-related information; and Transmitting the data including the bitstream, wherein the image information includes information related to the absolute values of the luminance filter coefficients for the ALF process and information related to the absolute values of the chrominance filter coefficients for the ALF process, wherein the information related to the absolute values of the luminance filter coefficients for the ALF process includes syntax elements for the absolute values of the luminance filter coefficients, wherein the information related to the absolute values of the chrominance filter coefficients for the ALF process includes syntax elements for the absolute values of the chrominance filter coefficients, Among them, the syntax element for the absolute value of the luminance filter coefficient and the syntax element for the absolute value of the chrominance filter coefficient are encoded based on zero-order exponential Golomb coding, and Among them, the value of the syntax element for the absolute value of the luminance filter coefficient and the value of the syntax element for the absolute value of the chrominance filter coefficient are within a predetermined range.