In-loop filtering-based image coding device and method
Patent Information
- Application Number
- KR1020227023882
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-15
- Filing Date
- 2021-01-15
- Publication Date
- 2026-09-23
- Estimated Expiration
- 2041-01-15
Smart Images

Figure 112022072307423-PCT00030_ABST
Abstract
Description
Technology Field
[0001] This document relates to an in-loop filtering-based image coding device and method. Background Technology
[0002] Recently, the demand for high-resolution, high-quality video, such as 4K or 8K or higher UHD (Ultra High Definition) video, is increasing across various fields. As video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases relative to existing video data; therefore, when transmitting video data using media such as existing wired or wireless broadband lines or storing video data using existing storage media, transmission and storage costs increase.
[0003] In addition, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have recently been increasing, and the broadcasting of video content with characteristics different from reality, such as game footage, is on the rise.
[0004] Accordingly, high-efficiency image / video compression technology is required to effectively compress, transmit, store, and play back high-resolution, high-quality image / video information having various characteristics as described above.
[0005] Furthermore, there is discussion regarding technologies such as ALF (adaptive loop filtering) to improve compression efficiency and enhance subjective and objective visual quality. To apply these technologies efficiently, a method for efficiently signaling related information is required. means of solving the problem
[0006] According to one embodiment of the present document, a method and apparatus for increasing image / video coding efficiency are provided.
[0007] According to one embodiment of the present document, an efficient filtering application method and apparatus are provided.
[0008] According to one embodiment of the present document, a method and apparatus for efficiently applying adaptive loop filtering (ALF) are provided.
[0009] According to one embodiment of the present document, a method and apparatus for increasing image / video coding efficiency are provided.
[0010] According to one embodiment of the present document, a method and apparatus for hierarchically signaling ALF-related information are provided.
[0011] According to one embodiment of the present document, a zero-order exponent Golomb coding scheme (ue(v)) may be used for a parsing procedure of information / syntax elements regarding absolute values of luminance / chroma ALF filter coefficients.
[0012] According to one embodiment of the present document, the range of values of information regarding the absolute values of the luminance / chroma ALF filter coefficients can be fixed.
[0013] According to one embodiment of the present document, an encoding device for performing video / image encoding is provided.
[0014] According to one embodiment of the present document, a computer-readable digital storage medium is provided that stores encoded video / image information generated according to a video / image encoding method disclosed in at least one of the embodiments of the present document.
[0015] According to one embodiment of the present document, a computer-readable digital storage medium is provided that stores encoded information or encoded video / image information, which causes a video / image decoding method disclosed in at least one of the embodiments of the present document to be performed by a decoding device. Effects of the invention
[0016] According to one embodiment of the present document, overall image / video compression efficiency can be increased.
[0017] According to one embodiment of the present document, subjective / objective visual quality can be improved through efficient filtering.
[0018] According to one embodiment of the present document, ALF-related information can be efficiently signaled.
[0019] According to one embodiment of the present document, computational overhead and complexity can be reduced by using a zero-order exponential Golomb coding scheme (ue(v)) for the parsing procedure of information / syntax elements regarding absolute values of luminance / chroma ALF filter coefficients.
[0020] According to one embodiment of the present document, coding using ue(v) can be performed efficiently by fixing the range of values of information regarding the absolute values of the luminance / chroma ALF filter coefficients. Brief explanation of the drawing
[0021] Figure 1 schematically illustrates an example of a video / image coding system that can be applied to embodiments of the present document. FIG. 2 is a diagram schematically illustrating the configuration of a video / image encoding device that can be applied to embodiments of the present document. FIG. 3 is a diagram schematically illustrating the configuration of a video / image decoding device that can be applied to embodiments of the present document. Figure 4 illustrates an exemplary hierarchical structure for a coded image / video. Figure 5 shows an example of an ALF filter shape. FIGS. 6 and 7 schematically illustrate an example of a video / image encoding method and related components according to the embodiment(s) of the present document. FIGS. 8 and 9 schematically illustrate an example of an image / video decoding method and related components according to the embodiment(s) of this document. FIG. 10 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied. Specific details for implementing the invention
[0022] As this document is subject to various modifications and may have various embodiments, specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. Terms used in this specification are used merely to describe specific embodiments and are not intended to limit the technical scope of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "comprising" or "having" in this specification are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0023] Meanwhile, each component in the drawings described in this document is depicted independently for the convenience of explaining different characteristic functions and does not imply that each component is implemented in separate hardware or separate software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of this document, provided that they do not deviate from the essence of this document.
[0024] Hereinafter, preferred embodiments of the present document will be described in more detail with reference to the attached drawings. Hereinafter, the same reference numerals are used for identical components in the drawings, and redundant descriptions of identical components are omitted.
[0025] This document relates to video / image coding. For example, the methods / executions disclosed in this document may relate to the VVC (Versatile Video Coding) standard (ITU-T Rec. H.266), next-generation video / image coding standards following VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec. H.265), EVC (essential video coding) standard, AVS2 standard, etc.).
[0026] This document presents various embodiments regarding video / image coding, and unless otherwise noted, the embodiments may be performed in combination with one another.
[0027] In this document, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image at a specific time, and "slice" or "tile" are units that constitute a part of a picture in coding. A slice or tile may contain one or more CTUs (coding tree units). A single picture may consist of one or more slices or tiles. A single picture may consist of one or more tile groups. A tile group may contain one or more tiles.
[0028] A pixel or pel can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. Generally, a sample can represent a pixel or its value; it may represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chroma component. Alternatively, a sample may refer to a pixel value in the spatial domain, or it may refer to the transformation coefficient in the frequency domain when such pixel values are converted to the frequency domain.
[0029] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0030] In this document, " / " and "," are interpreted as "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B and / or C." Furthermore, "A, B, C" also means "at least one of A, B and / or C."
[0031] Additionally, in this document, "or" is interpreted as "and / or". For example, "A or B" may mean 1) "A" only, 2) "B" only, or 3) "A and B". Alternatively, "or" in this document may mean "additionally or alternatively".
[0032] In this specification, "at least one of A and B" may mean "only A," "only B," or "both A and B." Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as synonymous with "at least one of A and B."
[0033] Additionally, in this specification, "at least one of A, B and C" may mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."
[0034] Additionally, parentheses used in this specification may mean "for example." Specifically, where indicated as "prediction (intra-prediction)," "intra-prediction" may be proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be proposed as an example of "prediction." Furthermore, even where indicated as "prediction (i.e., intra-prediction)," "intra-prediction" may be proposed as an example of "prediction."
[0035] Technical features described individually within a single drawing in this specification may be implemented individually or simultaneously.
[0036] Figure 1 schematically illustrates an example of a video / image coding system to which this document can be applied.
[0037] Referring to FIG. 1, a video / image coding system may include a source device and a receiving device. The source device may transmit encoded video / image information or data in the form of a file or streaming to the receiving device via a digital storage medium or a network.
[0038] The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0039] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include, for example, a computer, a tablet, and a smartphone, etc., and may generate video / images (electronically). For example, virtual video / images may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.
[0040] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0041] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0042] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.
[0043] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.
[0044] FIG. 2 is a diagram schematically illustrating the configuration of a video / image encoding device to which the present document may be applied. Hereinafter, the term "video encoding device" may include an image encoding device.
[0045] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoder chipset or processor) according to the embodiment. Additionally, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0046] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit may be divided into multiple coding units of a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first and the binary-tree structure and / or ternary structure may be applied later. Or the binary-tree structure may be applied first. A coding procedure according to this document may be performed based on the final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into lower-depth coding units so that a coding unit of the optimal size is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may each be divided or partitioned from the final coding unit described above.The above prediction unit may be a unit of sample prediction, and the above transformation unit may be a unit that derives transformation coefficients and / or a unit that derives a residual signal from transformation coefficients.
[0047] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.
[0048] The subtraction unit (231) can generate a residual signal (residual block, residual samples, or residual sample array) by subtracting the predicted signal (predicted block, predicted samples, or predicted sample array) output from the prediction unit (220) from the input image signal (original block, original samples, or original sample array), and the generated residual signal is transmitted to the conversion unit (232). The prediction unit (220) can perform a prediction for a block to be processed (hereinafter referred to as the current block) and generate a predicted block (predicted block) containing predicted samples for the current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied at the current block or CU unit. The prediction unit can generate various information regarding prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit it to the entropy encoding unit (240). Information regarding the prediction can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0049] The intra prediction unit (222) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, a DC mode and a Planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional prediction modes may be used. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0050] The inter prediction unit (221) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the reference blocks and the reference picture containing the temporal neighboring blocks may be the same or different. The above temporal surrounding blocks may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above temporal surrounding blocks may be referred to as a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on surrounding blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of surrounding blocks as motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of surrounding blocks is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0051] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called combined inter and intra prediction (CIIP). Additionally, the prediction unit may perform intra block copy (IBC) for a block prediction. The intra block copy may be used for content video / video coding, such as in screen content coding (SCC), for example, games. IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document.
[0052] The prediction signal generated through the inter prediction unit (221) and / or the intra prediction unit (222) can be used to generate a restoration signal or to generate a residual signal. The transformation unit (232) can generate transform coefficients by applying a transformation technique to the residual signal. For example, the transformation technique may include a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on generating a prediction signal using all previously reconstructed pixels. Additionally, the transformation process may be applied to a square pixel block of the same size, or to a non-square block of variable size.
[0053] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients. The entropy encoding unit (240) can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information necessary for video / image restoration (e.g., values of syntax elements) together or separately, in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The signaling / transmitting information and / or syntax elements described below in this document may be encoded through the encoding procedure described above and included in the bitstream.The above bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).
[0054] Quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transform to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235). An adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed samples, or reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit (220). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra-prediction of the next block to be processed within the current picture, and can also be used for inter-prediction of the next picture after undergoing filtering as described below.
[0055] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0056] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (270). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (290), as described below in the description of each filtering method. The information regarding filtering can be encoded in the entropy encoding unit (290) and output in the form of a bitstream.
[0057] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (280). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (200) and the decoding device, and can also improve encoding efficiency.
[0058] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter-prediction unit (221). The memory (270) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (221) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (270) can store restoration samples of the blocks restored within the current picture and transmit them to the intra-prediction unit (222).
[0059] Figure 3 is a diagram schematically illustrating the configuration of a video / image decoding device to which this document can be applied.
[0060] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (331) and an intra-predictor (332). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321). The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoder chipset or a processor) according to an embodiment. Additionally, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0061] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed by the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied by the encoding device. Accordingly, the processing unit for decoding may be, for example, a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.
[0062] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based on information regarding the parameter sets and / or the general constraint information. The signaling / receiving information and / or syntax elements described below in this document can be obtained from the bitstream by being decoded through the decoding procedure. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential chording, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information of the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate symbols corresponding to the values of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (330), and information regarding the residual on which entropy decoding was performed in the entropy decoding unit (310), namely quantized transformation coefficients and related parameter information, can be input to the inverse quantization unit (321). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310). Meanwhile, the decoding device according to the present document may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit (310), and the sample decoder may include at least one of the inverse quantization unit (321), inverse transform unit (322), prediction unit (330), addition unit (340), filtering unit (350), and memory (360).
[0063] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.
[0064] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0065] The prediction unit performs a prediction for the current block and can generate a predicted block containing prediction samples for the current block. Based on information regarding the prediction output from the entropy decoding unit (310), the prediction unit can determine whether an intra prediction or an inter prediction is applied to the current block and can determine a specific intra / inter prediction mode.
[0066] The prediction unit may generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for the prediction of a single block, as well as apply intra prediction and inter prediction simultaneously. This may be referred to as combined inter and intra prediction (CIIP). Additionally, the prediction unit may perform intra block copy (IBC) for the prediction of a block. The intra block copy may be used for content video / video coding, such as in games, for example, screen content coding (SCC). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this document.
[0067] The intra prediction unit (332) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or away from it, depending on the prediction mode. In intra prediction, the prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit (332) may determine the prediction mode applied to the current block by using the prediction mode applied to the surrounding blocks.
[0068] The inter prediction unit (331) can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit (331) may construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the mode of inter-prediction for the current block.
[0069] The adder (340) can generate a restoration signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the prediction signal (predicted block, predicted sample array) output from the prediction unit (330). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the restoration block.
[0070] The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture.
[0071] Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0072] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to the DPB of the memory (60), specifically the memory (360). The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0073] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter-prediction unit (331). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (331) to be used as motion information of spatially surrounding blocks or motion information of temporally surrounding blocks. The memory (360) can store restoration samples of the blocks restored within the current picture and transmit them to the intra-prediction unit (332).
[0074] In this specification, the embodiments described in the prediction unit (330), inverse quantization unit (321), inverse transformation unit (322), and filtering unit (350), etc. of the decoding device (300) may be applied to the prediction unit (220), inverse quantization unit (234), inverse transformation unit (235), and filtering unit (260), etc. of the encoding device (200) in the same or corresponding manner.
[0075] As described above, prediction is performed to increase compression efficiency during video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes predicted samples in the spatial domain (or pixel domain). The predicted block is derived identically by both the encoding device and the decoding device, and the encoding device can increase video coding efficiency by signaling information regarding the residual between the original block and the predicted block (residual information) to the decoding device, rather than the original sample value of the original block itself. The decoding device derives a residual block containing residual samples based on the residual information, can generate a restored block containing restored samples by combining the residual block and the predicted block, and can generate a restored picture containing the restored blocks.
[0076] The above residual information can be generated through transformation and quantization procedures. For example, an encoding device may derive a residual block between the original block and the predicted block, perform a transformation procedure on residual samples (residual sample array) included in the residual block to derive transformation coefficients, perform a quantization procedure on the transformation coefficients to derive quantized transformation coefficients, and signal the related residual information to a decoding device (via a bitstream). Here, the residual information may include information such as value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device may perform an inverse quantization / inverse transformation procedure based on the residual information and derive residual samples (or residual blocks). The decoding device may generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inversely quantizing / inversely transforming the quantized transform coefficients for reference to inter-predicting of the picture, and generate a restored picture based thereon.
[0077] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If the quantization / inverse quantization is omitted, the quantized transformation coefficients may be called transformation coefficients. If the transformation / inverse transformation is omitted, the transformation coefficients may be called coefficients or residual coefficients, or for the sake of consistency of expression, they may still be called transformation coefficients.
[0078] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information regarding the transform coefficient(s), and information regarding said transform coefficient(s) may be signaled through residual coding syntax. Transform coefficients may be derived based on said residual information (or information regarding said transform coefficient(s), and scaled transform coefficients may be derived through inverse transform (scaling) of said transform coefficients. Residual samples may be derived based on the inverse transform (transform) of said scaled transform coefficients. This may be similarly applied / expressed in other parts of this document.
[0079] The prediction unit of the encoding / decoding device can derive prediction samples by performing inter-prediction on a block-by-block basis. Inter-prediction may represent a prediction derived in a manner dependent on data elements (e.g., sample values, or motion information, etc.) of picture(s) other than the current picture. When inter-prediction is applied to the current block, a predicted block (prediction sample array) for the current block can be derived based on a reference block (reference sample array) specified by a motion vector on the reference picture pointed to by the reference picture index. At this time, to reduce the amount of motion information transmitted in the inter-prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample basis, based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on the inter-prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter-prediction is applied, neighboring blocks may include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture containing the above reference block and the reference picture containing the above temporal surrounding block may be the same or different. The above temporal surrounding block may be called by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above temporal surrounding block may be called collocated picture (colPic). For example, a list of motion information candidates may be constructed based on the surrounding blocks of the current block, and a flag or index information indicating which candidate is selected (used) to derive the motion vector and / or reference picture index of the current block may be signaled.Inter-prediction can be performed based on various prediction modes; for example, in the case of skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected surrounding block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the selected surrounding block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.
[0080] The above motion information may include L0 motion information and / or L1 motion information depending on the inter-prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be called an L0 prediction, a prediction based on an L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a pair (Bi) prediction. Here, the L0 motion vector may represent a motion vector associated with reference picture list L0 (L0), and the L1 motion vector may represent a motion vector associated with reference picture list L1 (L1). Reference picture list L0 may include pictures that are prior to the current picture in output order as reference pictures, and reference picture list L1 may include pictures that are subsequent to the current picture in output order. The aforementioned previous pictures may be referred to as forward (reference) pictures, and the aforementioned subsequent pictures may be referred to as reverse (reference) pictures. The reference picture list L0 may include additional pictures that are output later than the current picture as reference pictures. In this case, within the reference picture list L0, the previous pictures may be indexed first, and the subsequent pictures may be indexed next. The reference picture list L1 may include additional pictures that are output earlier than the current picture as reference pictures. In this case, within the reference picture list L1, the subsequent pictures may be indexed first, and the previous pictures may be indexed next. Here, the output order may correspond to the POC (picture order count) order.
[0081] Figure 4 illustrates an exemplary hierarchical structure for a coded image / video.
[0082] Referring to Figure 4, the coded image / video is divided into a VCL (video coding layer) that handles the decoding processing of the image / video and the image / video itself, a subsystem that transmits and stores the encoded information, and a NAL (network abstraction layer) that exists between the VCL and the subsystem and is responsible for network adaptation functions.
[0083] In VCL, VCL data containing compressed image data (slice data) can be generated, or parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally required in the decoding process of the image can be generated.
[0084] In NAL, a NAL unit can be created by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated in VCL. In this case, the RBSP refers to slice data, parameter sets, SEI messages, etc. generated in VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the NAL unit.
[0085] As illustrated in the figure above, NAL units can be classified into VCL NAL units and Non-VCL NAL units depending on the RBSP generated in VCL. A VCL NAL unit may refer to a NAL unit containing information about an image (slice data), and a Non-VCL NAL unit may refer to a NAL unit containing information necessary to decode an image (parameter set or SEI message).
[0086] The aforementioned VCL NAL unit and Non-VCL NAL unit can be transmitted over a network by attaching header information according to the data specifications of the underlying system. For example, the NAL unit can be transformed into a data format of a specified specification, such as H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted over various networks.
[0087] As described above, the NAL unit type can be determined according to the RBSP data structure included in the NAL unit, and information about this NAL unit type can be stored in the NAL unit header and signaled.
[0088] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether they contain information about the image (slice data). VCL NAL unit types can be classified according to the properties and types of the picture included in the VCL NAL unit, while Non-VCL NAL unit types can be classified according to the types of parameter sets.
[0089] The following is an example of a NAL unit type specified according to the type of parameter set included in the Non-VCL NAL unit type.
[0090] - APS (Adaptation Parameter Set) NAL unit: Type for the NAL unit containing the APS
[0091] - DPS (Decoding Parameter Set) NAL unit: Type for the NAL unit containing the DPS
[0092] - VPS (Video Parameter Set) NAL unit: Type for the NAL unit containing the VPS
[0093] - SPS (Sequence Parameter Set) NAL unit: Type for the NAL unit containing the SPS
[0094] - PPS(Picture Parameter Set) NAL unit: Type for the NAL unit containing the PPS
[0095] - PH (Picture header) NAL unit: Type for NAL unit containing PH
[0096] The above-described NAL unit types have syntax information for the NAL unit type, and said syntax information can be stored in the NAL unit header and signaled. For example, said syntax information may be nal_unit_type, and NAL unit types may be specified by the nal_unit_type value.
[0097] Meanwhile, as described above, a single picture may include multiple slices, and a single slice may include a slice header and slice data. In this case, a picture header may be additionally added for multiple slices (slice header and slice data set) within a single picture. The picture header (picture header syntax) may include information / parameters that can be commonly applied to the picture. In this document, a slice may be used interchangeably or replaced with a tile group. Additionally, in this document, a slice header may be used interchangeably or replaced with a type group header.
[0098] The slice header (slice header syntax, slice header information) may include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. The DPS may include information / parameters related to the concatenation of a CVS (coded video sequence). In this document, the term High level syntax (HLS) may include at least one of the above APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0099] In this document, the video information encoded from an encoding device to a decoding device and signaled in the form of a bitstream includes not only information related to picture partitioning, intra / inter prediction information, residual information, and in-loop filtering information, but may also include information included in the slice header, information included in the picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. Additionally, the video information may further include information from the NAL unit header.
[0100] The following table shows coding descriptors for the parsing procedure of coding-related information in this document. Coding descriptors can be used for the parsing procedure of syntax elements included in the syntax of this document.
[0101]
[0102] The following table shows x coded based on zero-order, first-order, second-order, and third-order exponential Colomb coding. For example, x can be a decimal number and the coded x can be a binary number. k represents the orders for exponential Colomb coding (k=0, 1, 2, 3). Zero-order exponential Colomb coding (k=0) can be used for the syntax element parsing procedure according to the coding descriptor of ue(v) described above, and Table 2 (k=0) can be referenced for the syntax element parsing procedure based on the coding descriptor of ue(v).
[0103]
[0104]
[0105]
[0106]
[0107] Meanwhile, to compensate for the difference between the original image and the restored image caused by errors occurring during compression encoding processes such as quantization, an in-loop filtering procedure may be performed on the restored samples or restored pictures as described above. As described above, in-loop filtering may be performed in the filter section of the encoding device and the filter section of the decoding device, and a deblocking filter, SAO, and / or an adaptive loop filter (ALF) may be applied. For example, the ALF procedure may be performed after the deblocking filtering procedure and / or SAO procedure are completed. However, even in this case, the deblocking filtering procedure and / or SAO procedure may be omitted.
[0108] A detailed explanation regarding picture restoration and filtering will be provided below. In video coding, restoration blocks may be generated based on intra prediction and inter prediction on a block-by-block basis, and a restored picture containing the restoration blocks may be generated. If the current picture / slice is I picture / slice, the blocks included in the current picture / slice may be restored based solely on intra prediction. Meanwhile, if the current picture / slice is P or B picture / slice, the blocks included in the current picture / slice may be restored based on either intra prediction or inter prediction. In this case, intra prediction may be applied to some blocks within the current picture / slice, while inter prediction may be applied to the remaining blocks.
[0109] Intra prediction may represent a prediction that generates prediction samples for the current block based on reference samples within the picture to which the current block belongs (hereinafter, the current picture). When intra prediction is applied to the current block, surrounding reference samples to be used for the intra prediction of the current block may be derived. The surrounding reference samples of the current block may include a sample adjacent to the left boundary of the current block of size nWxnH and a total of 2xnH samples adjacent to the bottom-left, a sample adjacent to the top boundary of the current block and a total of 2xnW samples adjacent to the top-right, and one sample adjacent to the top-left of the current block. Alternatively, the surrounding reference samples of the current block may include multiple columns of upper surrounding samples and multiple rows of left surrounding samples. Additionally, the surrounding reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of size nWxnH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.
[0110] However, some of the surrounding reference samples of the current block may not yet be decoded or may not be available. In this case, the decoder may construct the surrounding reference samples to be used for prediction by substituting the unavailable samples with available samples. Alternatively, the surrounding reference samples to be used for prediction may be constructed through the interpolation of available samples.
[0111] When neighboring reference samples are derived, a prediction sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the prediction sample can also be derived based on a reference sample existing in a specific (prediction) direction with respect to the prediction sample among the neighboring reference samples of the current block. Case (i) may be called a non-directional mode or non-angular mode, and case (ii) may be called a directional mode or angular mode. Additionally, the prediction sample may be generated through interpolation between the first neighboring sample and the second neighboring sample located in the opposite direction of the prediction direction of the intra prediction mode of the current block relative to the prediction sample of the current block among the neighboring reference samples. The above case may be called Linear Interpolation Intra Prediction (LIP). Furthermore, chroma prediction samples may be generated based on luminance samples using a linear model. This case may be called LM mode. In addition, a provisional prediction sample of the current block may be derived based on filtered surrounding reference samples, and a prediction sample of the current block may be derived by performing a weighted sum of the provisional prediction sample and at least one reference sample derived according to the intra prediction mode among the existing surrounding reference samples, i.e., unfiltered surrounding reference samples. The above case may be called PDPC (Position dependent intra prediction).In addition, intra-prediction coding can be performed by selecting the reference sample line with the highest prediction accuracy among the surrounding multiple reference sample lines of the current block, deriving a prediction sample using a reference sample located in the prediction direction from that line, and signaling the used reference sample line to a decoding device. The above-described case may be referred to as multi-reference line (MRL) intra prediction or MRL-based intra prediction. Furthermore, the current block may be divided into vertical or horizontal subpartitions to perform intra prediction based on the same intra prediction mode, while deriving and utilizing surrounding reference samples at the subpartition level. That is, in this case, the intra prediction mode for the current block is applied identically to the subpartitions, but intra prediction performance can be improved depending on the circumstances by deriving and utilizing surrounding reference samples at the subpartition level. This prediction method may be referred to as intra sub-partitions (ISP) or ISP-based intra prediction. The above-described intra prediction methods may be referred to as intra prediction types to distinguish them from the intra prediction mode in Section 1.2. The above-mentioned intra-prediction type may be referred to by various terms, such as intra-prediction technique or additional intra-prediction mode. For example, the above-mentioned intra-prediction type (or additional intra-prediction mode, etc.) may include at least one of the aforementioned LIP, PDPC, MRL, and ISP. A general intra-prediction method excluding specific intra-prediction types such as LIP, PDPC, MRL, and ISP may be referred to as a normal intra-prediction type. The normal intra-prediction type may be generally applied when specific intra-prediction types such as the above are not applied, and prediction may be performed based on the aforementioned intra-prediction mode. Meanwhile, post-processing filtering may be performed on the derived prediction samples as necessary.
[0112] Specifically, the intra-prediction procedure may include an intra-prediction mode / type determination step, a peripheral reference sample derivation step, and an intra-prediction mode / type-based prediction sample derivation step. Additionally, a post-filtering step for the derived prediction samples may be performed as needed.
[0113] A modified restored picture may be generated through an in-loop filtering procedure, and the modified restored picture may be output as a decoded picture from a decoding device, or stored in the decoded picture buffer or memory of an encoding / decoding device to be used as a reference picture in an inter-prediction procedure during subsequent encoding / decoding of the picture. As described above, the in-loop filtering procedure may include a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, and / or an adaptive loop filter (ALF) procedure. In this case, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, the adaptive loop filter (ALF) procedure, and the bilateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture. Or, for example, the ALF procedure may be performed after the deblocking filtering procedure is applied to the restored picture. This can be done in the same way in the encoding device.
[0114] Deblocking filtering is a filtering technique that removes distortion occurring at the boundaries between blocks in a restored picture. The deblocking filtering procedure may, for example, derive a target boundary in the restored picture, determine a boundary strength (bS) for the target boundary, and perform deblocking filtering on the target boundary based on the bS. The bS may be determined based on the prediction mode of two blocks adjacent to the target boundary, the difference in motion vectors, whether the reference picture is identical, the existence of non-zero valid coefficients, etc.
[0115] SAO is a method that compensates for the offset difference between the restored picture and the original picture on a sample-by-sample basis, and can be applied based on types such as band offset and edge offset. According to SAO, samples are classified into different categories based on each SAO type, and an offset value can be added to each sample based on the category. Filtering information for SAO may include information regarding whether SAO is applied, SAO type information, and SAO offset value information. SAO may also be applied to the restored picture after the deblocking filtering described above has been applied.
[0116] Adaptive Loop Filter (ALF) is a technique that filters a restored picture on a sample-by-sample basis using filter coefficients determined by the filter shape. An encoding device can determine whether to apply ALF, the ALF shape, and / or ALF filtering coefficients by comparing the restored picture with the original picture, and can signal this information to a decoding device. In other words, the filtering information for ALF may include information regarding whether to apply ALF, information on the ALF filter shape, and information on the ALF filtering coefficients. ALF may also be applied to the restored picture after the aforementioned deblocking filtering has been applied.
[0117] Figure 5 shows an example of an ALF filter shape.
[0118] Figure 5(a) shows a 7x7 diamond filter shape, and Figure 5(b) shows a 5x5 diamond filter shape. In Figure 5, Cn within the filter shape represents a filter coefficient. If n is the same in Cn, it indicates that the same filter coefficient can be assigned. In this document, the location and / or unit where filter coefficients are assigned according to the filter shape of the ALF may be called a filter tab. In this case, one filter coefficient may be assigned to each filter tab, and the arrangement of the filter tabs may correspond to the filter shape. A filter tab located at the center of the filter shape may be called a center filter tab. Two filter tabs with the same n value located at mutually corresponding positions relative to the center filter tab may be assigned the same filter coefficient. For example, in the case of a 7x7 diamond filter shape, it includes 25 filter tabs, and since filter coefficients C0 to C11 are assigned in a centrally symmetrical form, filter coefficients can be assigned to the 25 filter tabs using only 13 filter coefficients. Additionally, for example, in the case of a 5x5 diamond filter shape, it includes 13 filter tabs, and since filter coefficients C0 through C5 are assigned in a centrally symmetrical form, filter coefficients can be assigned to the 13 filter tabs using only 7 filter coefficients. For example, to reduce the amount of data regarding information on the signaled filter coefficients, 12 of the 13 filter coefficients for a 7x7 diamond filter shape may be (explicitly) signaled, and 1 filter coefficient may be (implicitly) derived. Also, for example, 6 of the 7 filter coefficients for a 5x5 diamond filter shape may be (explicitly) signaled, and 1 filter coefficient may be (implicitly) derived.
[0119] In one example, prior to the application of filtering, a geometric transformation may be applied to the filter coefficients and the corresponding filter clipping values based on the gradient values calculated for the block. The geometric transformation may include rotation, diagonal flipping, or vertical flipping.
[0120] The following equations represent the filter coefficients and clipping values to which transformations are applied for each direction (diagonal, vertical, rotation). In the equations below, K is the filter size, and k and l represent the coefficient coordinates. For example, k can be greater than or equal to 0, and l can be less than or equal to K-1. Position (0, 0) can be the top-left corner, and position (K-1, K-1) can be the bottom-right corner. Transformations can be applied to the filter coefficients f(k, l) and clipping values c(k, l) based on the gradient values calculated for the corresponding block.
[0121]
[0122]
[0123]
[0124] The following table shows the slope values for the four directions (g h , g v , g d1 , g d2 It shows an exemplary relationship between ) and the transformation applied to the current block.
[0125]
[0126] Filter parameters are signaled to derive ALF filter coefficients at the decoding stage. ALF filter parameters can be signaled in the APS and / or slice header. For example, in a single APS, up to 25 sets of luminance filter coefficients and clipping value indices can be signaled, and up to 8 sets of chroma filter coefficients and clipping value indices can be signaled. To reduce bit overhead, filter coefficients of different classes of luminance components need to be combined. In the slice header, the indices of the APS used for the current slice can be signaled.
[0127] Clipping value indices decoded from APS can be used to determine clipping values along with the luminance table and chroma table of clipping values. These clipping values can be based on internal bit depth.
[0128] In one example, the luminance table of the clipping values and the chroma table of the clipping values can be derived based on the following mathematical formulas. In the following formulas, B represents the internal bit depth, and N may be the number of clipping values. For example, N may be 4.
[0129]
[0130]
[0131] As an example, up to seven APS indices can be signaled in the slice header to indicate the sets of luminance filters currently used for the slice. Filtering procedures can be further controlled at the CTB level. A flag indicating whether an ALF is applied to the luminance CTB can always be signaled. The luminance CTB can select a filter set from 16 fixed filter sets and the filter sets of the APSs. Filter set indices can be signaled for the luminance CTB to indicate which filter set is applied. The 16 fixed filter sets can be predefined and hard-coded in both the encoder and the decoder.
[0132] For chroma components, the APS index can be signaled in the slice header to indicate the sets of chroma filters used for the current slice. At the CTB level, filter indices can be signaled for each chroma CTB if there are two or more sets of chroma filters in the APS.
[0133] To further limit multiplication complexity, bitstream conformance is applied so that the coefficient values at non-central positions are in the range of 0 to 28, and the coefficient values at other positions can be in the range of -27 to 27-1. The coefficient at the central position is not signaled in the bitstream and can be considered equal to 128.
[0134] If the ALF is currently available for the CTB, each R(i,j) within the CU can be filtered to calculate R'(i,j). For example, R'(i,j) can be calculated based on the following mathematical formula. f(k,l) can be filter coefficients, and K(x,y) can be a clipping function. Additionally, c(k,l) can be decoded clipping parameters. k and l can vary between -L / 2 and L / 2, where L can be the filter length. The clipping function K(x,y)=min(y,max(-y,x)) can also be expressed as Clip3(-y,y,x).
[0135]
[0136] As described above, an in-loop filtering procedure may be applied to the restored picture. In this case, to further enhance the subjective / objective visual quality of the restored picture, a virtual boundary may be defined, and the in-loop filtering procedure may be applied across said virtual boundary. The said virtual boundary may include discontinuous edges, such as those in 360-degree images, VR images, or PIP (picture in picture). For example, said virtual boundary may exist at a predetermined agreed-upon location, and its existence and / or location may be signaled. As an example, said virtual boundary may be located at the upper 4th sample line of the CTU row (specifically, for example, above the upper 4th sample line of said CTU row). As another example, information regarding the existence and / or location of said virtual boundary may be signaled via HLS. As described above, said HLS may include SPS, PPS, Pick the header, Slice header, etc.
[0137] High-level syntax signaling and semantics regarding the embodiments of this document will be described below.
[0138] One embodiment of this document may include a method for controlling loop filters. This method for controlling loop filters may be applied to a restored picture. In-loop filters (loop filters) may be used for decoding encoded bitstreams. Loop filters may include the deblocking, SAO, and ALF described above. SPS may include flags associated with deblocking, SAO, and ALF, respectively. The flags may indicate whether each tool is available for coding CLVS (coded layer video sequence) and CVS (coded video sequence) that reference the SPS.
[0139] In one example, when loop filters are available for coding pictures within CVS, the application of said loop filters can be controlled so as not to cross specific boundaries. For example, loop filters can be controlled so as not to cross subpicture boundaries, loop filters can be controlled so as not to cross tile boundaries, loop filters can be controlled so as not to cross slice boundaries, and / or loop filters can be controlled so as not to cross virtual boundaries.
[0140] Information related to in-loop filtering may include information, syntax, syntax elements, and / or semantics described in this document (or embodiments included therein). Information related to in-loop filtering may include information regarding whether an in-loop filtering procedure (in whole or in part) is available across specific boundaries (e.g., virtual boundaries, subpicture boundaries, slice boundaries, and / or tile boundaries). Image information included in a bitstream may include high-level syntax (HLS), and said HLS may include information related to in-loop filtering. Based on a determination of whether the in-loop filtering procedure is applied across specific boundaries, modified (or filtered) reconstructed samples (reconstructed pictures) may be generated. In one example, if the in-loop filtering procedure is disabled for all blocks / boundaries, the modified reconstructed samples may be identical to the reconstructed samples. In another example, the modified reconstructed samples may include modified reconstructed samples derived based on the in-loop filtering. However, in this case, based on the above decision, some of the reconstructed samples (e.g., reconstructed samples across virtual boundary) may not be in-loop filtered. For example, reconstructed samples crossing a specific boundary (including at least one of a virtual boundary, subpicture boundary, slice boundary, and / or tile boundary for which in-loop filtering is enabled) may be in-loop filtered, but reconstructed samples crossing other boundaries (including at least one of a virtual boundary, subpicture boundary, slice boundary, and / or tile boundary for which in-loop filtering is disabled) may not be in-loop filtered.
[0141] In one example, regarding whether an in-loop filtering procedure is performed across a virtual boundary, the in-loop filtering information may include an SPS virtual boundary existence flag, a picture header virtual boundary existence flag, information regarding the number of virtual boundaries, information regarding the locations of virtual boundaries, etc.
[0142] In the embodiments included in this document, information regarding the location of a virtual boundary may include information regarding the x-coordinate of a vertical virtual boundary and / or information regarding the y-coordinate of a horizontal virtual boundary. Specifically, information regarding the location of a virtual boundary may include information regarding the x-coordinate of a vertical virtual boundary and / or the y-coordinate of a horizontal virtual boundary in units of luma samples. Additionally, information regarding the location of a virtual boundary may include information regarding the number of information (syntax elements) regarding the x-coordinate of a vertical virtual boundary existing in the SPS. Additionally, information regarding the location of a virtual boundary may include information regarding the number of information (syntax elements) regarding the y-coordinate of a horizontal virtual boundary existing in the SPS. Alternatively, information regarding the location of a virtual boundary may include information regarding the number of information (syntax elements) regarding the x-coordinate of a vertical virtual boundary existing in the picture header. Additionally, information regarding the location of a virtual boundary may include information regarding the number of information (syntax elements) regarding the y-coordinate of a horizontal virtual boundary existing in the picture header.
[0143] The following tables show exemplary syntax and semantics of the sequence parameter set (SPS) according to the present embodiment.
[0144]
[0145]
[0146] The following tables show exemplary syntax and semantics of the picture parameter set (PPS) according to the present embodiment.
[0147]
[0148]
[0149] The following tables show exemplary syntax and semantics of a picture header according to the present embodiment.
[0150]
[0151]
[0152]
[0153]
[0154] The following tables show exemplary syntax and semantics of a slice header according to the present embodiment.
[0155]
[0156]
[0157] The signaling of information related to ALF filter coefficients will be explained below.
[0158] In conventional ALF procedures, k-order exponential Colomb coding with k=3 is used to signal the absolute values of luminance and chroma ALF coefficients. However, k-order exponential Colomb coding is problematic because it causes significant computational overhead and complexity.
[0159] The embodiments described in the following paragraphs may propose solutions for solving the aforementioned problems. The embodiments may be applied independently. Alternatively, at least two or more embodiments may be applied in combination.
[0160] The following table shows exemplary syntax of an adaptation parameter set (APS) according to the embodiments of this document.
[0161]
[0162] The following table shows exemplary syntax of ALF data according to the present embodiment.
[0163]
[0164] The following table shows exemplary semantics regarding the syntax elements included in the above syntax.
[0165]
[0166]
[0167] According to another embodiment of the present document, information regarding the absolute values of the lumina / chroma ALF filter coefficients (alf_luma_coeff_abs[sfIdx][j], alf_chroma_coeff_abs[altIdx][j]) can be parsed based on a zero-order exponential Columbus coding scheme (ue(v)).
[0168] The following table shows exemplary syntax of ALF data according to the present embodiment.
[0169]
[0170] The following table shows exemplary semantics regarding the syntax elements included in the above syntax.
[0171]
[0172]
[0173] According to the embodiments of this document described together with the tables above, computational overhead and complexity can be reduced by using a zero-order exponential Golomb coding scheme (ue(v)) for the parsing procedure of information / syntax elements (alf_luma_coeff_abs[sfIdx][j], alf_chroma_coeff_abs[altIdx][j]) regarding absolute values of lumina / chroma ALF filter coefficients. Additionally, coding using ue(v) can be performed efficiently by fixing the range of values (e.g., 0 to 128) of the information regarding absolute values of lumina / chroma ALF filter coefficients.
[0174] FIGS. 6 and 7 schematically illustrate an example of a video / image encoding method and related components according to the embodiment(s) of the present document.
[0175] The method disclosed in FIG. 6 may be performed by the encoding device disclosed in FIG. 2 or FIG. 7. Specifically, for example, S600 and S610 of FIG. 6 may be performed by the residual processing unit (230) of the encoding device in FIG. 7, S620 of FIG. 6 may be performed by the addition unit (250) of the encoding device in FIG. 7, S630 and / or S640 of FIG. 6 may be performed by the filtering unit (260) of the encoding device in FIG. 7, and S650 of FIG. 6 may be performed by the entropy encoding unit (240) of the encoding device in FIG. 7. Additionally, although not shown in FIG. 6, prediction samples or prediction-related information may be derived by the prediction unit (220) of the encoding device in FIG. 6, and a bitstream may be generated from residual information or prediction-related information by the entropy encoding unit (240) of the encoding device. The method disclosed in FIG. 6 may include the embodiments described above in this document.
[0176] Referring to FIG. 6, the encoding device can derive residual samples (S600). The encoding device can derive residual samples for the current block, and the residual samples for the current block can be derived based on the original samples and predicted samples of the current block. Specifically, the encoding device can derive predicted samples of the current block based on a prediction mode. In this case, various prediction methods disclosed in this document, such as inter-prediction or intra-prediction, may be applied. Residual samples can be derived based on the predicted samples and original samples. For inter-prediction, the encoding device can derive at least one reference picture, and inter-prediction can be performed based on the at least one reference picture. Predicted samples can be generated based on the inter-prediction. The encoding device can generate reference picture-related information based on the at least one reference picture.
[0177] The encoding device can derive transformation coefficients. The encoding device can derive transformation coefficients based on a transformation procedure for the residual samples. For example, the transformation procedure may include at least one of DCT, DST, GBT, or CNT.
[0178] The encoding device can derive quantized transformation coefficients. The encoding device can derive quantized transformation coefficients based on a quantization procedure for the transformation coefficients. The quantized transformation coefficients may have a one-dimensional vector form based on the coefficient scan order.
[0179] The encoding device can generate residual information (S610). The encoding device can generate residual information based on the residual samples for the current block. The encoding device can generate residual information representing the quantized transform coefficients. Residual information can be generated through various encoding methods such as exponential Golomb, CAVLC, CABAC, etc.
[0180] The encoding device can generate restoration samples (S620). The encoding device can generate restoration samples based on the residual information. The restoration samples can be generated by adding the residual samples based on the residual information with the prediction samples. Specifically, the encoding device performs a prediction (intra or inter prediction) for the current block and can generate restoration samples based on the original samples and the prediction samples generated from the prediction.
[0181] Reconstructed samples may include reconstructed luminance samples and reconstructed chroma samples. Specifically, residual samples may include residual luminance samples and residual chroma samples. Residual luminance samples may be generated based on original luminance samples and predicted luminance samples. Residual chroma samples may be generated based on original chroma samples and predicted chroma samples. An encoding device may derive transformation coefficients (luma transformation coefficients) for the residual luminance samples and / or transformation coefficients (chroma transformation coefficients) for the residual chroma samples. Quantized transformation coefficients may include quantized luminance transformation coefficients and / or quantized chroma transformation coefficients.
[0182] The encoding device can derive ALF filter coefficients (S630). The ALF filter coefficients may include luminance ALF filter coefficients and / or chroma ALF filter coefficients. Modified reconstructed samples can be generated based on the ALF filter coefficients.
[0183] The encoding device can generate ALF-related information (S640). The ALF-related information may include information regarding ALF filter coefficients, information regarding ALF-related clipping, etc. In addition, the ALF-related information may include information regarding available flags, information regarding presence flags for positioning high-level syntax (e.g., PPS, SPS, APS, picture header, etc.). For example, the information regarding ALF filter coefficients may include information regarding the absolute values of the ALF filter coefficients and information regarding the signs of the ALF filter coefficients.
[0184] An encoding device can encode video / image information (S650). The image information may include residual information, prediction-related information, reference picture-related information, subpicture-related information, in-loop filtering-related information and / or virtual boundary-related information (and / or additional virtual boundary-related information). The encoded video / image information may be output in the form of a bitstream. The bitstream may be transmitted to a decoding device via a network or a storage medium.
[0185] The above image / video information may include various information according to embodiments of the present document. For example, the above image / video information may include information disclosed in at least one of Tables 1 to 16 described above.
[0186] In one embodiment, the image information may include sets of parameters. At least one of the sets of parameters may include ALF data. The filter coefficients may be derived from the ALF data based on 0th-order Exponential Golomb coding.
[0187] In one embodiment, at least one of the parameter sets may include an adaptation parameter set (APS). The ALF data may be included in the APS.
[0188] In one embodiment, the image information may include header information and an ALF-related adaptation parameter set (APS). The header information may include information regarding the number of ALF-related APS IDs. The number of ALF-related APS IDs may be derived based on the value of the information regarding the number of ALF-related APS IDs. A number of ALF-related APS ID syntax elements equal to the number of ALF-related APS IDs may be included in the header information.
[0189] In one embodiment, the image information may include header information and an ALF-related adaptation parameter set (APS). The header information may include an ALF availability flag indicating whether the ALF is available in a picture or slice, and information regarding the number of ALF-related APS IDs. When the value of the ALF availability flag is 1, the header information may include information regarding the number of ALF-related APS IDs. The value of the information regarding the number of ALF-related APS IDs plus 1 may be equal to the number of ALF-related APS IDs.
[0190] In one embodiment, the ALF data may include information regarding the absolute values of the luminance filter coefficients for the ALF procedure. The information regarding the absolute values of the luminance filter coefficients for the ALF procedure may be coded based on ue(v).
[0191] In one embodiment, the ALF data may include information regarding the absolute values of the chroma filter coefficients for the ALF procedure. The information regarding the absolute values of the chroma filter coefficients for the ALF procedure may be coded based on ue(v).
[0192] In one embodiment, the image information may include information regarding the absolute values of the luminance filter coefficients for the ALF procedure and information regarding the absolute values of the chroma filter coefficients for the ALF procedure. The absolute values of the luminance filter coefficients for the ALF procedure and the absolute values of the chroma filter coefficients for the ALF procedure may be within a predetermined range.
[0193] FIGS. 8 and 9 schematically illustrate an example of a video / image decoding method and related components according to the embodiment(s) of the present document.
[0194] The method disclosed in FIG. 8 may be performed by the decoding device disclosed in FIG. 3 or FIG. 9. Specifically, for example, S800 of FIG. 8 may be performed by the entropy decoding unit (310) of the decoding device, S810 may be performed by the residual processing unit (320) and / or addition unit (340) of the decoding device, and S820 and / or S830 may be performed by the filtering unit (350) of the decoding device. The method disclosed in FIG. 8 may include the embodiments described above in this document.
[0195] Referring to FIG. 8, a decoding device can receive / acquire video / image information (S800). The video / image information may include residual information, prediction-related information, reference picture-related information, subpicture-related information, in-loop filtering-related information and / or ALF-related information. The decoding device can receive / acquire the video / image information through a bitstream.
[0196] The above image / video information may include various information according to embodiments of the present document. For example, the above image / video information may include information disclosed in at least one of Tables 1 to 16 described above.
[0197] The decoding device can derive quantized transformation coefficients. The decoding device can derive quantized transformation coefficients based on the residual information. The quantized transformation coefficients may have a one-dimensional vector form based on the coefficient scan order. The quantized transformation coefficients may include quantized luminance transformation coefficients and / or quantized chroma transformation coefficients.
[0198] The decoding device can derive transformation coefficients. The decoding device can derive transformation coefficients based on an inverse quantization procedure for the quantized transformation coefficients. The decoding device can derive luminance transformation coefficients through inverse quantization based on the quantized luminance transformation coefficients. The decoding device can derive chroma transformation coefficients through inverse quantization based on the quantized chroma transformation coefficients.
[0199] The decoding device can generate / derive residual samples. The decoding device can derive residual samples based on an inverse transform procedure for the transformation coefficients. The decoding device can derive residual luminance samples through an inverse transform procedure based on the luminance transformation coefficients. The decoding device can derive residual chroma samples through an inverse transform procedure based on the chroma transformation coefficients.
[0200] The decoding device can derive at least one reference picture based on information related to the reference picture. The decoding device can perform a prediction procedure based on the at least one reference picture. Specifically, the decoding device can generate prediction samples of the current block based on a prediction mode. In this case, various prediction methods disclosed in this document, such as inter-prediction or intra-prediction, may be applied. The decoding device can generate prediction samples for the current block within the current picture based on the prediction procedure. For example, the decoding device can perform an inter-prediction procedure based on the at least one reference picture and generate prediction samples based on the inter-prediction procedure.
[0201] The decoding device can generate / derive reconstructed samples (S810). For example, the decoding device can generate / derive reconstructed luminance samples and / or reconstructed chroma samples. The decoding device can generate reconstructed luminance samples and / or reconstructed chroma samples based on the residual information. The decoding device can generate reconstructed samples based on the residual information. The reconstructed samples may include reconstructed luminance samples and / or reconstructed chroma samples. The luminance component of the reconstructed samples may correspond to the reconstructed luminance samples, and the chroma component of the reconstructed samples may correspond to the reconstructed chroma samples. The decoding device can generate predicted luminance samples and / or predicted chroma samples through a prediction procedure. The decoding device can generate reconstructed luminance samples based on the predicted luminance samples and the residual luminance samples. The decoding device can generate reconstructed chroma samples based on the predicted chroma samples and the residual chroma samples.
[0202] The decoding device can derive ALF filter coefficients (S820). The ALF filter coefficients may include luminance ALF filter coefficients and / or chroma ALF filter coefficients. The ALF filter coefficients can be derived based on information regarding the absolute values of the ALF filter coefficients and information regarding the signs of the ALF filter coefficients.
[0203] The decoding device can generate modified (filtered) restored samples (S830). The decoding device can generate modified restored samples based on an in-loop filtering procedure for the restored samples. The decoding device can generate modified restored samples based on in-loop filtering related information. The decoding device may use a deblocking procedure, an SAO procedure, and / or an ALF procedure to generate modified restored samples.
[0204] In one embodiment, the image information may include sets of parameters. At least one of the sets of parameters may include ALF data. The filter coefficients may be derived from the ALF data based on 0th-order Exponential Golomb coding.
[0205] In one embodiment, at least one of the parameter sets may include an adaptation parameter set (APS). The ALF data may be included in the APS.
[0206] In one embodiment, the image information may include header information and an ALF-related adaptation parameter set (APS). The header information may include information regarding the number of ALF-related APS IDs. The number of ALF-related APS IDs may be derived based on the value of the information regarding the number of ALF-related APS IDs. A number of ALF-related APS ID syntax elements equal to the number of ALF-related APS IDs may be included in the header information.
[0207] In one embodiment, the image information may include header information and an ALF-related adaptation parameter set (APS). The header information may include an ALF availability flag indicating whether the ALF is available in a picture or slice, and information regarding the number of ALF-related APS IDs. When the value of the ALF availability flag is 1, the header information may include information regarding the number of ALF-related APS IDs. The value of the information regarding the number of ALF-related APS IDs plus 1 may be equal to the number of ALF-related APS IDs.
[0208] In one embodiment, the ALF data may include information regarding the absolute values of the luminance filter coefficients for the ALF procedure. The information regarding the absolute values of the luminance filter coefficients for the ALF procedure may be coded based on ue(v).
[0209] In one embodiment, the ALF data may include information regarding the absolute values of the chroma filter coefficients for the ALF procedure. The information regarding the absolute values of the chroma filter coefficients for the ALF procedure may be coded based on ue(v).
[0210] In one embodiment, the image information may include information regarding the absolute values of the luminance filter coefficients for the ALF procedure and information regarding the absolute values of the chroma filter coefficients for the ALF procedure. The absolute values of the luminance filter coefficients for the ALF procedure and the absolute values of the chroma filter coefficients for the ALF procedure may be within a predetermined range.
[0211] If residual samples for the current block exist, the decoding device may receive information regarding residuals for the current block. The information regarding residuals may include transformation coefficients regarding residual samples. The decoding device may derive residual samples (or a residual sample array) for the current block based on the residual information. Specifically, the decoding device may derive quantized transformation coefficients based on the residual information. The quantized transformation coefficients may have a one-dimensional vector form based on the coefficient scan order. The decoding device may derive transformation coefficients based on an inverse quantization procedure for the quantized transformation coefficients. The decoding device may derive residual samples based on the transformation coefficients.
[0212] The decoding device can generate reconstructed samples based on (intra) prediction samples and residual samples, and can derive a reconstructed block or a reconstructed picture based on said reconstructed samples. Specifically, the decoding device can generate reconstructed samples based on the sum between the (intra) prediction samples and the residual samples. As previously described, the decoding device may then apply in-loop filtering procedures, such as deblocking filtering and / or SAO procedures, to said reconstructed picture to improve subjective / objective image quality as needed.
[0213] For example, a decoding device can decode a bitstream or encoded information to obtain image information containing all or part of the information (or syntax elements) described above. Additionally, the bitstream or encoded information may be stored in a computer-readable storage medium and may cause the decoding method described above to be performed.
[0214] In the embodiments described above, methods are described based on flowcharts as a series of steps or blocks; however, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps as described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps of the flowcharts may be omitted without affecting the scope of the embodiments of this document.
[0215] The method according to the embodiments of the present document described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present document may be included in a device that performs image processing, such as a TV, computer, smartphone, set-top box, display device, etc.
[0216] When the embodiments described in this document are implemented in software, the method described above may be implemented as a module (process, function, etc.) that performs the function described above. The module may be stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.
[0217] In addition, the decoding device and encoding device to which the embodiment(s) of this document apply may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.
[0218] Additionally, the processing method to which the embodiment(s) of this document are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this document may also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Additionally, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission over the Internet). Additionally, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0219] Additionally, the embodiment(s) of this document may be implemented as a computer program product by program code, and said program code may be executed on a computer by the embodiment(s) of this document. said program code may be stored on a computer-readable carrier.
[0220] FIG. 10 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.
[0221] Referring to FIG. 10, a content streaming system to which embodiments of the present document are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0222] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.
[0223] The bitstream above may be generated by an encoding method or a bitstream generation method to which the embodiments of the present document are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0224] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays the role of controlling commands and responses between each device within the content streaming system.
[0225] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.
[0226] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.
[0227] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.
[0228] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined to be implemented as a device, and the technical features of the device claims in this specification may be combined to be implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a device, and the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a method.
Claims
Claim 1 A video decoding method performed by a decoding device, comprising: a step of acquiring video information including residual information through a bitstream; a step of generating reconstructed samples based on the residual information; a step of deriving filter coefficients for an adaptive loop filter (ALF) procedure of the reconstructed samples; and a step of generating modified reconstructed samples based on the reconstructed samples and the filter coefficients, wherein the video information includes parameter sets, at least one of the parameter sets includes ALF data, and the ALF data includes information representing the absolute value of the filter coefficient for the ALF procedure and information representing the sign of the filter coefficient for the ALF procedure, wherein the information representing the absolute value of the filter coefficient for the ALF procedure is coded based on 0th-order Exponential Golomb coding, and the information representing the sign of the filter coefficient for the ALF procedure is based on an unsigned integer using 1 bit. Claim 2 A video decoding method according to claim 1, wherein at least one of the parameter sets includes an adaptation parameter set (APS), and the ALF data is included in the APS. Claim 3 A method for decoding images according to claim 1, wherein the image information includes header information and an ALF-related APS (adaptation parameter set), the header information includes information regarding the number of ALF-related APS IDs, the number of ALF-related APS IDs is derived based on the value of the information regarding the number of ALF-related APS IDs, and the header information includes ALF-related APS ID syntax elements equal to the number of ALF-related APS IDs. Claim 4 A video decoding method according to claim 1, wherein the video information includes header information and an ALF-related APS (adaptation parameter set), the header information includes an ALF availability flag indicating whether the ALF is available in a picture or slice and information regarding the number of ALF-related APS IDs, and based on the value of the ALF availability flag being 1, the header information includes information regarding the number of ALF-related APS IDs, and the value of the information regarding the number of ALF-related APS IDs is equal to the number of ALF-related APS IDs. Claim 5 delete Claim 6 delete Claim 7 delete Claim 8 A video encoding method performed by an encoding device, comprising: a step of deriving residual samples for a current block; a step of generating residual information based on the residual samples; a step of generating restored samples based on the residual information; a step of deriving filter coefficients for an adaptive loop filter (ALF) procedure for the restored samples; a step of generating ALF-related information based on the filter coefficients; and a step of encoding video information including the residual information and the ALF-related information, wherein the video information includes parameter sets, at least one of the parameter sets includes ALF data, information representing the absolute value of the filter coefficient for the ALF procedure is coded based on 0th-order Exponential Golomb coding, and information representing the sign of the filter coefficient for the ALF procedure is based on an unsigned integer using 1 bit. Claim 9 A video encoding method according to claim 8, wherein at least one of the parameter sets includes an adaptation parameter set (APS), and the ALF data is included in the APS. Claim 10 A video encoding method according to claim 8, wherein the video information includes header information and an ALF-related APS (adaptation parameter set), the header information includes information regarding the number of ALF-related APS IDs, the number of ALF-related APS IDs is derived based on the value of the information regarding the number of ALF-related APS IDs, and ALF-related APS ID syntax elements equal to the number of ALF-related APS IDs are included in the header information. Claim 11 A video encoding method according to claim 8, wherein the video information includes header information and an ALF-related APS (adaptation parameter set), the header information includes an ALF availability flag indicating whether the ALF is available in a picture or slice and information regarding the number of ALF-related APS IDs, and based on the value of the ALF availability flag being 1, the header information includes information regarding the number of ALF-related APS IDs, and the value of the information regarding the number of ALF-related APS IDs plus 1 is equal to the number of ALF-related APS IDs. Claim 12 delete Claim 13 delete Claim 14 delete Claim 15 A computer-readable storage medium that stores a bitstream generated by an image encoding method, wherein the image encoding method comprises: a step of deriving residual samples for a current block; a step of generating residual information based on the residual samples; a step of generating restored samples based on the residual information; a step of deriving filter coefficients for an adaptive loop filter (ALF) procedure for the restored samples; a step of generating ALF-related information based on the filter coefficients; and a step of encoding image information including the residual information and the ALF-related information to generate the bitstream, wherein the image information includes parameter sets, at least one of the parameter sets includes ALF data, information representing the absolute value of the filter coefficient for the ALF procedure is coded based on 0th-order Exponential Golomb coding, and information representing the sign of the filter coefficient for the ALF procedure is based on an unsigned integer using 1 bit. Claim 16 A method for transmitting data for an image, comprising: acquiring a bitstream for the image, wherein the bitstream is generated based on the steps of: deriving residual samples for a current block; generating residual information based on the residual samples; generating restored samples based on the residual information; deriving filter coefficients for an adaptive loop filter (ALF) procedure for the restored samples; generating ALF-related information based on the filter coefficients; and encoding image information including the residual information and the ALF-related information; and transmitting the data including the bitstream, wherein the image information includes parameter sets, at least one of the parameter sets includes ALF data, information representing the absolute value of the filter coefficient for the ALF procedure is coded based on 0th-order Exponential Golomb coding, and information representing the sign of the filter coefficient for the ALF procedure is based on an unsigned integer using 1 bit.
Citation Information
Patent Citations
Coding parameter sets and nal unit headers for video coding
KR1020140120336A
Image decoding method and apparatus in image coding system
KR1020180054695A
A picture encoder, a picture decoder and corresponding methods
WO2020007489A1