Method and apparatus for signaling video information applicable at picture or slice level

By defining the application level of tools in the image coding system, image decoding efficiency is improved, the problem of increased transmission and storage costs for high-resolution and high-quality images is solved, and more efficient image compression is achieved.

CN115104313BActive Publication Date: 2026-05-26LG ELECTRONICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2020-12-10
Publication Date
2026-05-26

Smart Images

  • Figure CN115104313B_ABST
    Figure CN115104313B_ABST
Patent Text Reader

Abstract

A video decoding method performed by a decoding apparatus according to this disclosure includes the following steps: obtaining indication information indicating whether one or more tools for a current block can be applied at the picture level or the slice level; determining, based on the indication information, whether information related to the one or more tools exists in the picture header or the slice header; parsing the information related to the one or more tools from the picture header or the slice header based on the determination result; and decoding the current block based on the information related to the one or more tools.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to image coding technology, and most specifically, to a method and apparatus for signaling image (or video) information applicable at the picture or slice level in an image coding system. Background Technology

[0002] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, has been growing across various fields. As the resolution and quality of image data increase, the size of the information transmitted, or the number of bits, increases compared to existing image data. Therefore, transmission and storage costs increase when transmitting image data using the same media such as conventional (or existing) wired / wireless broadband lines, or when storing image data using conventional (or existing) storage media.

[0003] Therefore, there is a need for efficient image compression techniques for effectively sending, storing, and reproducing (or playing back) information about high-resolution and high-quality images. Summary of the Invention

[0004] Technical Purpose

[0005] The technical objective of this disclosure is to provide methods and apparatus for increasing image coding efficiency.

[0006] Another technical objective of this disclosure is to provide a method and apparatus for signaling image (or video) information applicable at the picture or slice level.

[0007] Another technical objective of this disclosure is to provide a method and apparatus for performing decoding on a current block based on image (or video) information applicable at the picture level or slice level.

[0008] Technical solution

[0009] According to embodiments of this disclosure, an image decoding method performed by a decoding device is provided herein. The method may include the following steps: obtaining indication information indicating whether at least one tool for a current block is applied at the picture level or the slice level; determining, based on the indication information, whether information related to the at least one tool exists in the picture header or the slice header; resolving, based on the determination, the information related to the at least one tool from the picture header or the slice header; and decoding the current block based on the information related to the at least one tool.

[0010] According to another embodiment of this disclosure, an image encoding method performed by an encoding device is provided herein. The method may include the following steps: generating indication information indicating whether at least one tool to be applied to the current block is applied at the picture level or the slice level; generating information related to said at least one tool; and encoding image information including said indication information and the information related to said at least one tool. Here, the indication information indicates whether the information related to said at least one tool is present in the picture header or the slice header.

[0011] According to another embodiment of this disclosure, a computer-readable digital storage medium is provided herein in which encoded image information allowing an image decoding method to be executed by a decoding device is stored. The image decoding method according to the embodiment may include the steps of: obtaining indication information indicating whether at least one tool to be applied to a current block is applied at the picture level or the slice level; determining, based on the indication information, whether information related to the at least one tool is present in a picture header or a slice header; resolving, based on the determination, the information related to the at least one tool from the picture header or the slice header; and decoding the current block based on the information related to the at least one tool.

[0012] The effect of this disclosure

[0013] According to this manual, overall image / video compression efficiency can be improved.

[0014] According to this specification, the efficiency of image decoding can be improved based on indication information indicating whether at least one tool used for the current block is applied at the picture level or the slice level. Attached Figure Description

[0015] Figure 1 Examples of video / image coding systems applicable to this disclosure are illustrated schematically.

[0016] Figure 2 This is a schematic illustration of the configuration of a video / image encoding device applicable to this disclosure.

[0017] Figure 3 This is a schematic illustration of the configuration of a video / image decoding device applicable to this disclosure.

[0018] Figure 4 An exemplary hierarchical structure for encoded data is shown.

[0019] Figure 5 This is a flowchart illustrating a method for performing deblocking filtering according to an embodiment.

[0020] Figure 6 This is a flowchart that schematically illustrates an example of an ALF procedure.

[0021] Figure 7 An example of a filter shape for ALF is shown.

[0022] Figure 8 This is a flowchart illustrating the operation of an image encoding device according to an embodiment.

[0023] Figure 9 This is a block diagram illustrating the configuration of an image encoding device according to an embodiment.

[0024] Figure 10 This is a flowchart illustrating the operation of an image decoding device according to an embodiment.

[0025] Figure 11 This is a block diagram illustrating the configuration of an image decoding device according to an embodiment.

[0026] Figure 12 An example of a content streaming system to which the embodiments disclosed in this specification can be applied is shown. Detailed Implementation

[0027] This document may be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit this document. The terminology used in the following description is used only to describe specific embodiments and is not intended to limit this document. Singular expressions include plural expressions, provided that different interpretations are clear. Terms such as “comprising” and “having” are intended to indicate the presence of the features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, and therefore it should be understood that the possibility of having or adding one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.

[0028] Furthermore, the various configurations described in the accompanying drawings are illustrated independently to illustrate functions that are distinct from each other, and do not imply that the configurations are implemented using different hardware or different software. For example, two or more configurations may be combined to form one configuration, and a configuration may be divided into multiple configurations. Without departing from the spirit of this document, embodiments in which configurations are combined and / or separated are included within the scope of the claims.

[0029] In this disclosure, the term "A or B" may mean "A only", "B only", or "both A and B". In other words, in this disclosure, the term "A or B" may be interpreted as indicating "A and / or B". For example, in this disclosure, the term "A, B or C" may mean "A only", "B only", "C only", or "any combination of A, B, and C".

[0030] The forward slash ( / ) or comma used in this disclosure can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".

[0031] In this disclosure, "at least one of A and B" may mean "only A", "only B" or "both A and B". Furthermore, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as the same as "at least one of A and B".

[0032] Additionally, in this disclosure, "at least one of A, B, and C" may mean "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".

[0033] Additionally, the parentheses used in this disclosure may mean "for example". Specifically, when expressing "prediction (intra-frame prediction)", it may indicate that "intra-frame prediction" is proposed as an example of "prediction". In other words, the term "prediction" in this disclosure is not limited to "intra-frame prediction", and may indicate that "intra-frame prediction" is proposed as an example of "prediction". Furthermore, even when expressing "prediction (i.e., intra-frame prediction)", it may indicate that "intra-frame prediction" is proposed as an example of "prediction".

[0034] The technical features described separately in the accompanying drawings of this specification can be implemented individually or simultaneously.

[0035] In the following, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used to indicate the same configured elements in the drawings, and for simplicity, overlapping (or repetitive) descriptions of the same configured elements will be omitted.

[0036] Figure 1 Examples of video / image coding systems applicable to this disclosure are illustrated schematically.

[0037] refer to Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.

[0038] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0039] Video sources can acquire video / images through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. For example, a video / image capture device may include one or more cameras, a video / image archive including previously captured video / images, etc. For example, a video / image generation device may include a computer, tablet computer, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, etc. In this case, the video / image capture process may be replaced by a process that generates related data.

[0040] Encoding devices can encode input video / images. For compression and encoding efficiency, encoding devices can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output as a bitstream.

[0041] The transmitter can transmit encoded images / image information or data, output as a bitstream, to the receiver of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received bitstream to a decoding device.

[0042] Decoding devices can decode video / images by performing a series of processes such as inverse quantization, inverse transform, and prediction, which correspond to the operations of encoding devices.

[0043] The renderer can render decoded video / images. The rendered video / images can be displayed on a monitor.

[0044] This specification relates to video / image coding. For example, the methods / examples disclosed in this specification can be applied to methods disclosed in the Universal Video Coding (VVC) standard, the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding Standard 2 (AVS2) standard, or other next-generation video / image coding standards (e.g., H.267, H.268, etc.).

[0045] This document proposes various implementations of video / image coding, and unless otherwise specified, the above implementations can also be combined with each other.

[0046] In this document, video can refer to a series of images over time. An image typically refers to a unit representing an image within a specific time frame, and a slice / tile refers to a unit that, from a coding perspective, constitutes part of an image. A slice / tile may include one or more coding tree units (CTUs). An image can consist of one or more slices / tiles.

[0047] A tile is a rectangular area of ​​CTUs within a specific tile column and a specific tile row in an image. A tile column is a rectangular area of ​​CTUs with a height equal to the height of the image and a width specified by a syntax element in the image parameter set. A tile row is a rectangular area of ​​CTUs with a width specified by a syntax element in the image parameter set and a height equal to the height of the image. A tile scan is a specific ordering of the CTUs in a segmented image, where CTUs are ordered consecutively by a CTU raster scan within a tile, and tiles in the image are ordered consecutively by a raster scan of the tile within the image. A slice may include multiple whole (complete) tiles or multiple consecutive (or adjacent) CTU matrices that can be included within a tile of an image in a single NAL unit. In this specification, tile groups and slices may be used interchangeably. For example, in this specification, a tile group / tile group header may be referred to as a slice / slice header.

[0048] Furthermore, an image can be divided into two or more sub-images. A sub-image can be a rectangular region of one or more slices within the image.

[0049] A pixel, or image unit, can refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or a pixel value, and can represent pixel / pixel values ​​for either the luminance component or the chrominance component.

[0050] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients in M ​​columns and N rows. Alternatively, a sample may refer to a pixel value in the spatial domain, and when such a pixel value is transformed to the frequency domain, it may refer to a transform coefficient in the frequency domain.

[0051] Figure 2This is a schematic illustration of the configuration of a video / image coding apparatus applicable to this disclosure. Hereinafter, the term "video coding apparatus" may include an image coding apparatus.

[0052] Reference Figure 2 The encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transform 232, a quantizer 233, an inverse quantizer 234, and an inverse transform 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.

[0053] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into one or more processors. For example, a processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively segmented from coding tree unit (CTU) or maximum coding unit (LCU) according to a quadtree-binary-trinary tree (QTBTTT) structure. For example, a coding unit may be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure may be applied first, followed by a binary tree structure and / or a ternary structure. Alternatively, a binary tree structure may be applied first. The encoding process according to this disclosure may be performed based on the final coding unit that is no longer segmented. In this case, the maximum coding unit may be used as the final coding unit based on image characteristics, coding efficiency, etc., or if necessary, the coding unit may be recursively segmented into deeper coding units, and the coding unit with the optimal size may be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes (described later). As another example, the processor may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or separated from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.

[0054] In some cases, a unit can be used interchangeably with terms such as block or region. Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples typically represent pixels or pixel values, and can represent pixel / pixel values ​​of only the luminance component or only the chrominance component. Samples can be used as a term corresponding to a picture (or image) of pixels or cells.

[0055] In encoding device 200, a residual signal (residual block, residual sample array) is generated by subtracting the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is sent to converter 232. In this case, as shown, the portion in encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block that includes the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As described later in the description of each prediction mode, the predictor can generate various prediction-related information such as prediction mode information and send the generated information to entropy encoder 240. The prediction information can be encoded in entropy encoder 240 and output as a bitstream.

[0056] Intra-predictor 222 can refer to samples in the current image to predict the current block. Depending on the prediction mode, the referenced samples may be located near or separated from the current block. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, non-directional modes may include DC mode and planar mode. For example, depending on the level of detail in the prediction direction, the directional modes may include 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-predictor 222 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.

[0057] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to deduce the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0058] Predictor 220 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as Intra-Frame and Inter-Frame Prediction Combination (CIIP). Alternatively, the predictor can predict blocks based on Intra-Block Copy (IBC) prediction mode or Palette mode. IBC prediction mode or Palette mode can be used for content image / video coding such as games, for example, Screen Content Coding (SCC). IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction in terms of deriving reference blocks in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure. Palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying Palette mode, sample values ​​within the frame can be signaled based on information about the palette table and palette index.

[0059] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform technique can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transform generated based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or to blocks of variable size that are not square.

[0060] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scan order, and generate information about the quantized transform coefficients based on this one-dimensional vector form. Entropy encoder 240 can perform various encoding methods such as (for example) exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can encode information necessary for video / image reconstruction (e.g., values ​​of syntax elements) other than the quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be sent or stored in NAL (Network Abstraction Layer) units as a bitstream. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may also include general constraint information. In this disclosure, information and / or syntax elements sent / signed from the encoding device to the decoding device may be included in the video / image information. The video / image information can be encoded using the encoding process described above and included in a bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 or a storage unit (not shown) that stores the signal may be included as an internal / external element of the encoding device 200, and alternatively, the transmitter may be included in the entropy encoder 240.

[0061] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients using inverse quantizer 234 and inverse transform 235. Adder 250 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If the target block to be processed has no residual (e.g., in the case of applying a skip mode), the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be filtered for inter-frame prediction of the next image, as described below.

[0062] In addition, a luminance mapping with chroma scaling (LMCS) can be applied during image encoding and / or reconstruction processing.

[0063] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a corrected reconstructed image by applying various filtering methods to the reconstructed image and store the corrected reconstructed image in memory 270, specifically in the DPB of memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various types of filtering-related information and transmit the generated information to entropy encoder 290, as described in the subsequent descriptions of each filtering method. The filtering-related information can be encoded by entropy encoder 290 and output as a bitstream.

[0064] The corrected reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. When inter-frame prediction is applied by the encoding device, prediction mismatch between the encoding device 200 and the decoding device can be avoided, and encoding efficiency can be improved.

[0065] The DPB of memory 270 can store the corrected reconstructed image for use as a reference image in inter-frame predictor 221. Memory 270 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of blocks in already reconstructed images. The stored motion information can be transmitted to inter-frame predictor 221 to be used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and can transmit these reconstructed samples to intra-frame predictor 222.

[0066] Figure 3This is a schematic diagram illustrating the configuration of a video / image decoding device applicable to this disclosure.

[0067] Reference Figure 3 The decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 may be configured by hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded image buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0068] When the input includes a bitstream containing video / image information, the decoding device 300 can reconstruct and... Figure 2 The encoding device processes video / image information corresponding to the image. For example, the decoding device 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can use a processor applied in the encoding device to perform decoding. Therefore, for example, the processor for decoding can be an encoding unit, and the encoding unit can be segmented from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units can be derived from the encoding units. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.

[0069] Decoding device 300 can receive data from... in the form of a bitstream. Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The signaling / receiving information and / or syntax elements described subsequently in this disclosure can be decoded by the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CABAC, or CAVLC, and output the syntax elements required for image reconstruction and the quantized values ​​of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine the context model using information about the target syntax element, the decoding information of the target block, or information about the symbols / bins decoded in the previous stage, and perform arithmetic decoding on the bins by predicting the probability of bin occurrence based on the determined context model, generating symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin after determining the context model. Prediction-related information from the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values ​​(i.e., quantized transform coefficients and related parameter information) from the entropy decoding performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive the residual signals (residual blocks, residual samples, residual sample arrays). In addition, filtering information from the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) for receiving signals output from the encoding device can be additionally configured as an internal / external component of the decoding device 300, or the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this disclosure can be referred to as a video / image / picture decoding device, and the decoding device can be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of an inverse quantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0070] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. The dequantizer 321 can use quantization parameters (e.g., quantization step size information) to perform dequantization on the quantized transform coefficients and obtain the transform coefficients.

[0071] The inverse transformer 322 performs inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0072] The predictor can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 310, and can determine a specific intra-frame / inter-frame prediction mode.

[0073] Predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as Intra-Frame and Inter-Frame Prediction Combination (CIIP). Alternatively, the predictor can predict blocks based on Intra-Block Copy (IBC) prediction mode or Palette mode. IBC prediction mode or Palette mode can be used for content image / video coding such as games, for example, Screen Content Coding (SCC). IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction in terms of deriving reference blocks in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure. Palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying Palette mode, sample values ​​within the frame can be signaled based on information about the palette table and palette index.

[0074] Intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples may be located among the neighbors of the current block, or their location may be separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Intra-predictor 331 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.

[0075] Inter-frame predictor 332 can deduce the predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the motion information correlation between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can construct a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.

[0076] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, or reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from predictor 330. If there is no residual to process the target block (e.g., in the case of applying a jump mode), the prediction block can be used as the reconstruction block.

[0077] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and as described later, it can also be output by filtering or used for inter-frame prediction of the next image.

[0078] In addition, Luminance Mapping with Chroma Scaling (LMCS) can also be applied to image decoding processing.

[0079] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a corrected reconstructed image by applying various filtering methods to the reconstructed image and store the corrected reconstructed image in memory 360, specifically in the DPB of memory 360. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.

[0080] The (corrected) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks from which motion information in the current image is derived (or decoded) and / or motion information of blocks in already reconstructed images. The stored motion information can be transmitted to inter-frame predictor 332 to be used as motion information for spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and transmit the reconstructed samples to intra-frame predictor 331.

[0081] In this document, the implementation methods described in the filter 260, inter-frame predictor 221, and intra-frame predictor 222 of the encoding device 200 can be the same as or respectively applied to the filter 350, inter-frame predictor 332, and intra-frame predictor 331 of the decoding device 300. The same applies to the inter-frame predictor 332 and the intra-frame predictor 331.

[0082] Furthermore, as mentioned above, prediction is performed during video encoding to enhance compression efficiency. A prediction block, including predicted samples of the current block (i.e., the target coding block), can be generated through prediction. In this case, the prediction block includes predicted samples in the spatial domain (or pixel domain). The prediction block is derived identically in both the encoding and decoding devices. The encoding device can enhance image coding efficiency by signaling information (residual information) about the residual between the original block (rather than the original sample values ​​of the original block) and the prediction block. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed image including the reconstructed block.

[0083] Residual information can be generated through transform and quantization processes. For example, the encoding device can derive the residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, derive quantized transform coefficients by performing a quantization process on the transform coefficients, and signal the relevant residual information (via bitstream) to the decoding device. In this case, the residual information may include the value information, position information, transform technique, transform kernel, and quantization parameters of the quantized transform coefficients. The decoding device can perform inverse quantization / inverse transform processes based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the prediction block and the residual block. Furthermore, for inter-frame prediction reference of subsequent images, the encoding device can derive the residual block by performing inverse quantization / inverse transform on the quantized transform coefficients and generate a reconstructed image based on this.

[0084] Figure 4 An example is shown of the hierarchical structure of encoded data.

[0085] refer to Figure 4 Encoded data can be divided into the Video Coding Layer (VCL), which manipulates the encoding processing of video / images and the video / images themselves, and the Network Abstraction Layer (NAL), which exists between the VCL and the subsystem that stores and transmits the encoded video / images.

[0086] VCL can generate parameter sets corresponding to the headers of sequences and images (Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc.) as well as Supplementary Enhancement Information (SEI) messages needed for additional video / image encoding processing. SEI messages are separate from information about the video / image (slice data). The VCL, including information about the video / image, consists of slice data and slice headers. Furthermore, the slice header can be called a tile group header, and the slice data can be called tile group data.

[0087] In NAL, NAL cells can be generated by adding header information (NAL cell header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc., generated in VCL. The NAL cell header can include NAL cell type information specified according to the RBSP data included in the corresponding NAL cell.

[0088] The NAL unit plays the role of mapping encoded images to bit sequences of subsystems such as file formats, Real-time Transport Protocol (RTP), and Transport Stream (TS).

[0089] As shown in the figure, based on the RBSP generated in the VCL, NAL units can be classified into VCL NAL units and non-VCL NAL units. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data), while a non-VCL NAL unit can refer to a NAL unit that includes information (parameter set or SEI message) required to decode the image.

[0090] The aforementioned VCL NAL units and non-VCL NAL units can be transmitted over a network according to the data standard appender information of the subsystem. For example, NAL units can be converted into predetermined standard data formats such as H.266 / VVC file format, Real-time Transport Protocol (RTP), Transport Stream (TS), etc., and transmitted over various networks.

[0091] As described above, a NAL unit can be specified by the NAL unit type according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header for concurrent signal notification.

[0092] For example, NAL units can be classified into VCLNAL unit types and non-VCL NAL unit types based on whether they include information about the image (slice data). VCL NAL unit types can be classified according to the nature and type of the image included in the VCL NAL unit, while non-VCL NAL unit types can be classified according to the type of parameter set.

[0093] The following are examples of NAL unit types specified based on the parameter set types included in non-VCL NAL unit types. For example, a NAL unit type can be specified as an APS NAL unit as a NAL unit type including an Adaptive Parameter Set (APS), a DPS NAL unit as a NAL unit type including a Decoding Parameter Set (DPS), a VPS NAL unit as a NAL unit type including a Video Parameter Set (VPS), an SPS NAL unit as a NAL unit type including a Sequence Parameter Set (SPS), and a PPS NAL unit as a NAL unit type including a Picture Parameter Set (PPS).

[0094] The NAL unit types mentioned above can have syntax information specific to the NAL unit type, and this syntax information can be stored in the NAL unit header for concurrent signaling. For example, the syntax information can be `nal_unit_type`, and the NAL unit type can be specified through the `nal_unit_type` value.

[0095] Furthermore, as mentioned above, an image can include multiple slices, and a slice can include a slice header and slice data. In this case, an image header can be further added to the multiple slices (slice headers and slice datasets) in an image. The image header (image header syntax) can include information / parameters that can be commonly applied to the image. The slice header (slice header syntax) can include information / parameters that can be commonly applied to the slice. APS (APS syntax) or PPS (PPS syntax) can include information / parameters that can be commonly applied to one or more slices or images. SPS (SPS syntax) can include information / parameters that can be commonly applied to one or more sequences. VPS (VPS syntax) can include information / parameters that can be commonly applied to multiple layers. DPS (DPS syntax) can include information / parameters that can be commonly applied to the entire video. DPS can include information / parameters related to the concatenation of encoded video sequences (CVS). The High-Level Syntax (HLS) in this document can include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.

[0096] In this document, the image / image information encoded by the encoding device and transmitted to the decoding device in the form of a bitstream includes not only segmentation-related information, intra / inter-frame prediction information, residual information, and loop filtering information in the image, but also information included in the slice header, APS, PPS, SPS, VPS, and / or DPS. Additionally, the image / image information may also include NAL unit header information.

[0097] Furthermore, as mentioned above, to enhance the quality of the subjective / objective images, the encoding / decoding device can perform a loop filtering process on the reconstructed image. A modified reconstructed image can be generated through the loop filtering process, and this modified reconstructed image can be output from the decoding device as a decoded image and can also be stored in the decoding image buffer or memory of the encoding / decoding device. Additionally, in subsequent processing, the modified reconstructed image can be used as a reference image in the inter-frame prediction process during encoding / decoding. As mentioned above, the loop filtering process can include a deblocking filtering process, a sample adaptive offset (SAO) process, and / or an adaptive loop filter (ALF) process. In this case, one or a portion of the deblocking filtering process, the sample adaptive offset (SAO) process, and the adaptive loop filter (ALF) process can be applied sequentially, or all processes can be applied sequentially. For example, the SAO process can be performed after applying the deblocking filtering process to the reconstructed image. Alternatively, for example, the ALF process can be performed after applying the deblocking filtering process to the reconstructed image. This can be performed similarly in the encoding device.

[0098] Deblocking filtering is the process of removing any distortions that occur at the boundaries between blocks in a reconstructed image. For example, deblocking filtering can derive the target boundary from the reconstructed image, determine the boundary strength (bS) of the target boundary, and perform deblocking filtering on the target boundary based on the determined bS. ​​bS can be determined based on factors such as the prediction patterns of two adjacent blocks, the difference in motion vectors, whether the reference image is the same, and the presence of non-zero effective coefficients.

[0099] The SAO (Side Array Optimization) process is a method of compensating for the offset difference between the reconstructed image and the original image on a sample-by-sample basis. For example, the SAO process can be applied based on offset types such as frequency band offset or edge offset. Based on the SAO, samples can be categorized according to the SAO type, and an offset value can be added to each sample according to the category. The filtering information for SAO can include information about whether SAO is applied or not, SAO type information, and SAO offset value information. For example, SAO can be applied to the reconstructed image after applying deblocking filtering.

[0100] The Adaptive Loop Filter (ALF) process is a process of filtering the reconstructed image on a sample-by-sample basis based on the filter coefficients according to the filter shape. The encoding device can compare the reconstructed image with the original image to determine whether to apply ALF, the ALF shape, and / or the ALF filter coefficients, and can signal the reconstructed image back to the encoding device. That is, filtering information about the ALF process can include information about whether ALF is applied, ALF shape information, ALF filter coefficient information, etc. The ALF process can be applied to the reconstructed image after deblocking filtering.

[0101] Figure 5 This is a flowchart illustrating a method for performing deblocking filtering according to an embodiment.

[0102] As mentioned above, encoding / decoding devices can reconstruct an image block by block. When performing such image reconstruction block by block, block distortion may occur at the boundaries between blocks within the reconstructed image. Therefore, to remove the block distortion that occurs at the boundaries between blocks within the reconstructed image, encoding and decoding devices can use deblocking filters.

[0103] Therefore, the encoding / decoding device can deduce the boundaries between blocks where deblocking filtering is performed within the reconstructed image. Furthermore, the boundaries where deblocking filtering is performed can be referred to as edges. Additionally, the boundaries where deblocking filtering is performed can include two different types, and these two different types of boundaries can be vertical boundaries and horizontal boundaries. Vertical boundaries can also be referred to as vertical edges, and horizontal boundaries can also be referred to as horizontal edges. The encoding / decoding device can perform deblocking filtering on vertical edges, and it can also perform deblocking filtering on horizontal edges.

[0104] For example, the encoding / decoding device can deduce the target boundary that will be processed by filtering from the reconstructed image (S510).

[0105] Additionally, the encoding / decoding device can determine the boundary strength (bS) of the boundary to which deblocking filtering is performed (S520). bS can also be indicated as the boundary filtering strength. For example, it can be assumed that the bS value of the boundary (block edge) between block P and block Q is obtained. In this case, the encoding / decoding device can obtain the bS value of the boundary (block edge) between block P and block Q based on block P and block Q. For example, bS can be determined according to the table shown below.

[0106] [Table 1]

[0107]

[0108]

[0109] In this paper, p can indicate a sample of block P adjacent to the deblocking filter target boundary, and q can indicate a sample of block Q adjacent to the deblocking filter target boundary.

[0110] Additionally, for example, p0 can indicate samples of blocks adjacent to the left or top side of the deblocking filter target boundary, and q0 can indicate samples of blocks adjacent to the right or bottom side of the deblocking filter target boundary. For example, when the direction of the target boundary is vertical (i.e., when the target boundary is a vertical boundary), p0 can indicate samples of blocks adjacent to the left side of the deblocking filter target boundary, and q0 can indicate samples of blocks adjacent to the right side of the deblocking filter target boundary. Alternatively, for example, when the direction of the target boundary is horizontal (i.e., when the target boundary is a horizontal boundary), p0 can indicate samples of blocks adjacent to the top side of the deblocking filter target boundary, and q0 can indicate samples of blocks adjacent to the bottom side of the deblocking filter target boundary.

[0111] Return to reference Figure 5 The encoding / decoding device can perform block filtering based on bS (S530). For example, when the bS value is equal to 0, deblocking filtering is not applied to the target boundary. Furthermore, based on the determined bS value, a filter applied to the boundaries between blocks can be determined. Filters can be classified as strong filters and weak filters. By using different filters to perform filtering on each of the boundaries within the reconstructed image where the probability of block distortion is high and the boundaries where the probability of block distortion is low, the encoding / decoding device can improve encoding efficiency.

[0112] Figure 6 This is a flowchart that schematically illustrates an example of an ALF procedure. Figure 6 The ALF procedures (or processing) disclosed herein can be executed in both encoding and decoding devices. In this document, the encoding device may include both encoding and / or decoding devices.

[0113] refer to Figure 6 The encoding device derives a filter for ALF (S610). The filter may include filter coefficients. The encoding device can determine whether to apply ALF, and when it is determined that ALF is applied, it can derive a filter including filter coefficients for ALF. The information used to derive the filter (coefficients) for ALF, or the filter (coefficients) for ALF, may be referred to as ALF parameters. Information regarding whether ALF is applied (i.e., the ALF enable flag) and ALF data used to derive the filter can be signaled from the encoding device to the decoding device. The ALF data may include information used to derive the filter for ALF. Additionally, for example, for ALF hierarchical control, the ALF enable flag can be signaled at the SPS, picture header, slice header, and / or CTB levels, respectively.

[0114] To derive the filter for ALF, the activity and / or directivity of the current block (or ALF target block) are derived, and the filter can be derived based on the activity and / or directivity. For example, ALF processing can be applied in 4×4 block units (based on the luma component). The current block or ALF target block can be, for example, a CU, or a 4×4 block within a CU. Specifically, for example, the filter for ALF can be derived based on a first filter derived from information included in the ALF data and a predefined second filter, and the encoding device can select one of the filters based on the activity and / or directivity. The encoding device can use the filter coefficients included in the selected filter for ALF.

[0115] The encoding device performs filtering based on filters (S620). Modified reconstructed samples can be derived based on the filtering. For example, filter coefficients in the filters can be arranged or assigned according to the filter shape, and filtering can be performed on the reconstructed samples in the current block. Here, the reconstructed samples in the current block can be the reconstructed samples after deblocking filtering and SAO processing. For example, a single filter shape can be used, or a filter shape can be selected from multiple predetermined filter shapes. For example, the filter shape applied to the luma component and the filter shape applied to the chroma component can be different. For example, a 7×7 diamond filter shape can be used for the luma component, and a 5×5 diamond filter shape can be used for the chroma component.

[0116] Figure 7 Examples of filter shapes for ALF are shown. C0 to C11 in (a) and C0 to C5 in (b) can be filter coefficients that depend on their position within each filter shape.

[0117] Figure 7 (a) shows the shape of the 7×7 rhombus filter, and Figure 8 (b) shows the shape of a 5×5 rhombus filter. Figure 8In this document, Cn in the filter shape represents the filter coefficients. When n in Cn is the same, it indicates that the same filter coefficients can be assigned. In this document, the position and / or unit where filter coefficients are assigned according to the filter shape of the ALF can be called a filter tap. In this case, one filter coefficient can be assigned to each filter tap, and the arrangement of the filter taps can correspond to the filter shape. The filter tap located at the center of the filter shape can be called the center filter tap. The same filter coefficients can be assigned to two filter taps with the same n value that exist at positions corresponding to each other relative to the center filter tap. For example, in the case of a 7×7 rhombus filter shape, which includes 25 filter taps, and since filter coefficients C0 to C11 are assigned in a centrally symmetric manner, only 13 filter coefficients can be used to assign filter coefficients to the 25 filter taps. Alternatively, for example, in the case of a 5×5 rhombus filter shape, which includes 13 filter taps, and since filter coefficients C0 to C5 are assigned in a centrally symmetric manner, only 7 filter coefficients can be used to assign filter coefficients to the 13 filter taps. For example, to reduce the amount of data related to the filter coefficients that need to be signaled, 12 of the 13 filter coefficients in a 7×7 diamond filter shape can be signaled explicitly, and one filter coefficient can be derived implicitly. Alternatively, for example, 6 of the 7 filter coefficients in a 5×5 diamond filter shape can be signaled explicitly, and one filter coefficient can be derived implicitly.

[0118] According to the implementation described in this document, the ALF parameters used for ALF processing can be signaled via an adaptive parameter set (APS). The ALF parameters can be derived from filter information used for ALF or ALF data.

[0119] ALF is a loop filtering technique that can be applied to video / image coding as described above. ALF can be performed using Wiener-based adaptive filters. This can be done to minimize the mean square error (MSE) between the original sample and the decoded sample (or reconstructed sample). Advanced designs for ALF tools can include syntax elements accessible from the SPS and / or slice header (or tile group header).

[0120] Furthermore, the image header includes syntax elements applied to it, and these syntax elements can be applied to all slices of the image associated with the image header. When a specific syntax element is applied only to a specific slice, the specific syntax element should be signaled from the slice header, not the image header.

[0121] In existing technologies, signaling for enabling or disabling multiple tools used for image encoding or decoding resides in the image header and is overridden in the slice header. This method provides the flexibility to allow tool control at both the image and slice levels. However, when using this method, the need to verify the slice header after verifying the image header can impose a burden on the decoder.

[0122] Therefore, embodiments of this disclosure propose indication information that indicates whether at least one tool is applied at the image level or the slice level. In this case, the indication information can be included in either a sequence parameter set (SPS) or a picture parameter set (PPS). That is, when a specific tool is activated (or enabled) within the CLV, an indication or flag for indicating whether the specific tool is applied at the image level or the slice level can be signaled from a parameter set such as SPS or PPS. Although this indication or flag may correspond to a single tool, this disclosure is not limited to this. For example, an indication or flag for indicating whether all tools, not just a specific tool, are applied at the image level or the slice level can be signaled from a parameter set such as SPS or PPS.

[0123] Although the control flags and parameters used to enable or disable tools can be signaled at either the image or slice level, the signaling is not performed at either level. For example, if an instruction is received indicating whether to apply a specific tool at the image level, the control flags and parameters for enabling or disabling that tool can be signaled only at the image level. Similarly, if an instruction is received indicating whether to apply a specific tool at the slice level, the control flags and parameters for enabling or disabling that tool can be signaled only at the slice level.

[0124] Additionally, for example, a tool that is specified from a particular set of parameters for application at the image level can be specified from another set of parameters of the same type for application at the slice level.

[0125] For example, the PPS syntax that includes instruction information can be shown in the table below.

[0126] [Table 2]

[0127]

[0128] The semantics of the grammatical elements included in the grammar in Table 2 can be indicated, for example, as shown in Table 3.

[0129] [Table 3]

[0130]

[0131]

[0132] Referring to the table above, the indication information may include flags indicating whether the signaling of the reference image list is applied at the image level or the slice level. For example, the indication information may specify whether information related to the signaling of the reference image list appears (or exists) in the image header or in the slice header. For example, the flag could be indicated as `rpl_present_in_ph_flag`. If the value of the corresponding flag is 1, the information related to the signaling of the reference image list exists in the image header. And if the value of the corresponding flag is 0, the information related to the signaling of the reference image list exists in the slice header.

[0133] Additionally, the indication information may include flags indicating whether the Sample Adaptive Offset (SAO) process is applied at the image level or the slice level. For example, the indication information may specify whether information related to the SAO process appears (or exists) in the image header or in the slice header. For instance, the flag could be indicated as `sao_present_in_ph_flag`. If the value of the corresponding flag is 1, the information related to the SAO process exists in the image header. And if the value of the corresponding flag is 0, the information related to the SAO process exists in the slice header.

[0134] Additionally, the indication information may include flags indicating whether the Adaptive Loop Filter (ALF) procedure is applied at the image level or the slice level. For example, the indication information may specify whether ALF-related information appears (or exists) in the image header or in the slice header. For instance, the flag could be indicated as `alf_present_in_ph_flag`. If the corresponding flag value is 1, the ALF-related information exists in the image header. And if the corresponding flag value is 0, the ALF-related information exists in the slice header.

[0135] Additionally, the indication information may include at least one flag indicating whether the deblocking process is applied at the image level or the slice level. For example, based on at least one flag, information related to the deblocking process may appear (or may exist) in either the image header or the slice header. For example, at least one flag may be indicated as `deblocking_filter_ph_override_enabled_flag` or `deblocking_filter_sh_override_enabled_flag`. For example, if the value of at least one flag is equal to 1, a flag indicating whether parameters related to the deblocking process exist in the image header may exist in the image header. And, if the value of at least one flag is equal to 0, a flag indicating whether parameters related to the deblocking process exist in the image header may not exist in the image header.

[0136] Alternatively, based on the case where at least one flag is equal to 1, a flag indicating whether parameters related to the deblocking process exist in the slice header may be present in the slice header. Furthermore, based on the case where at least one flag is equal to 0, a flag indicating whether parameters related to the deblocking process exist in the slice header may not be present in the slice header. In this case, the values ​​of `deblocking_filter_ph_override_enabled_flag` and `deblocking_filter_sh_override_enabled_flag` may not both be equal to 1.

[0137] In addition, the image header syntax can be as shown in the table below.

[0138] [Table 4]

[0139]

[0140] The semantics of the grammatical elements included in the grammar in Table 4 can be indicated, for example, as shown in Table 5 below.

[0141] [Table 5]

[0142]

[0143] Referring to the table above, when the value of `deblocking_filter_ph_override_enabled_flag` corresponding to the flag indicating whether deblocking is applied at the image level is equal to 1, `pic_deblocking_filter_override_present_flag` can be signaled. When the value of `pic_deblocking_filter_override_present_flag` is equal to 1, `pic_deblocking_filter_override_flag` corresponding to the flag indicating whether parameters related to deblocking are present in the image header can be present in the image header. Alternatively, when the value of `pic_deblocking_filter_override_present_flag` is equal to 0, `pic_deblocking_filter_override_flag` corresponding to the flag indicating whether parameters related to deblocking are present in the image header can be absent from the image header.

[0144] Additionally, when the value of `pic_deblocking_filter_override_flag` corresponding to the flag indicating whether parameters related to the deblocking process exist in the image header is equal to 1, the deblocking parameters may appear in (or may exist in) the image header. Furthermore, when the value of `pic_deblocking_filter_override_flag` corresponding to the flag indicating whether parameters related to the deblocking process exist in the image header is equal to 0, the deblocking parameters may not appear in (or may not exist in) the image header.

[0145] Additionally, when the value of `pic_deblocking_filter_disabled_flag` is equal to 1, deblocking filtering is not applied to slices related to the image header. Conversely, when the value of `pic_deblocking_filter_disabled_flag` is equal to 0, deblocking filtering is applied to slices related to the image header.

[0146] Additionally, `pic_beta_offset_div2` and `pic_tc_offset_div2` can specify deblocking parameter offsets of β and tC (values ​​divided by 2) for the slices associated with the image header, respectively. The values ​​of `pic_beta_offset_div2` and `pic_tc_offset_div2` can both be in the range of -6 to 6.

[0147] In addition, the slice header syntax can be shown in the table below.

[0148] [Table 6]

[0149]

[0150]

[0151] The semantics of the grammatical elements included in the grammar in Table 6 can be indicated, for example, as shown in Table 7.

[0152] [Table 7]

[0153]

[0154]

[0155] Referring to the table above, when the value of `deblocking_filter_sh_override_enabled_flag` corresponding to the flag indicating whether deblocking is applied at the slice level is equal to 1, `slice_deblocking_filter_override_present_flag` can be signaled. When the value of `slice_deblocking_filter_override_present_flag` is equal to 1, `slice_deblocking_filter_override_flag` corresponding to the flag indicating whether parameters related to deblocking are present in the slice header can exist in the slice header. Alternatively, when the value of `slice_deblocking_filter_override_present_flag` is equal to 0, `slice_deblocking_filter_override_flag` corresponding to the flag indicating whether parameters related to deblocking are present in the slice header can be absent from the slice header.

[0156] Additionally, when the value of `slice_deblocking_filter_override_flag` corresponding to the flag indicating whether parameters related to the deblocking process exist in the slice header is equal to 1, the deblocking parameters may appear in (or may exist in) the slice header. Furthermore, when the value of `slice_deblocking_filter_override_flag` corresponding to the flag indicating whether parameters related to the deblocking process exist in the slice header is equal to 0, the deblocking parameters may not appear in (or may not exist in) the slice header.

[0157] Additionally, when the value of slice_deblocking_filter_disabled_flag is equal to 1, the deblocking filter is not applied to slices associated with the slice header. Conversely, when the value of slice_deblocking_filter_disabled_flag is equal to 0, the deblocking filter is applied to slices associated with the slice header.

[0158] Additionally, slice_beta_offset_div2 and slice_tc_offset_div2 can specify deblocking parameter offsets of β and tC (values ​​divided by 2) for the slice, respectively. The values ​​of slice_beta_offset_div2 and slice_tc_offset_div2 can both be in the range of -6 to 6.

[0159] Figure 8This is a flowchart illustrating the operation of an image encoding device according to an embodiment, and Figure 9 This is a block diagram illustrating the configuration of an image encoding device according to an embodiment.

[0160] Figure 8 The method disclosed in the article can be derived from Figure 2 or Figure 9 The encoding device disclosed in the document is used for execution. Figure 8 The S810 and S820 in the text can be generated by Figure 2 The image predictor 220, residual processor 230, or filter 260 shown in the figure are executed, and Figure 8 The S830 in the middle can be made by Figure 2 The entropy encoder 240 shown in the diagram is executed. Furthermore, the operations according to S810 to S830 are based on the above... Figures 1 to 7 This is part of the description given. Therefore, for simplicity, it will be omitted or briefly given in relation to... Figures 1 to 7 The description overlaps with the detailed description.

[0161] refer to Figure 8 The encoding device according to the embodiment can generate indication information indicating whether at least one tool to be applied to the current block is applied at the picture level or the slice level (S810).

[0162] For example, the image predictor 220 of the encoding device can generate indication information including a flag indicating whether the signaling of the reference image list is applied at the image level or the slice level. For example, if the flag value is equal to 1, the information related to the signaling of the reference image list can be stored in the image header. And if the flag value is equal to 0, the information related to the signaling of the reference image list can be stored in the slice header.

[0163] For example, the filter 260 of the encoding device can generate indication information including a flag indicating whether the Sample Adaptive Offset (SAO) process is applied at the picture level or the slice level. For example, if the flag value is equal to 1, the information related to the SAO process can be stored in the picture header. And if the flag value is equal to 0, the information related to the SAO process can be stored in the slice header.

[0164] For example, the filter 260 of the encoding device can generate indication information including a flag indicating whether the Adaptive Loop Filter (ALF) process is applied at the picture level or the slice level. For example, if the flag value is equal to 1, the information related to the ALF process can be stored in the picture header. And if the flag value is equal to 0, the information related to the ALF process can be stored in the slice header.

[0165] Alternatively, for example, the filter 260 of the encoding device can generate indication information including at least one flag indicating whether the deblocking process is applied at the picture level or the slice level. Based on at least one flag, information related to the deblocking process can exist in either the picture header or the slice header. For example, if the value of at least one flag is equal to 1, a flag indicating whether parameters related to the deblocking process exist in the picture header can exist in the picture header. And if the value of at least one flag is equal to 0, a flag indicating whether parameters related to the deblocking process exist in the picture header can not exist in the picture header.

[0166] The encoding apparatus according to the embodiment can generate information related to at least one tool (S820). For example, the image predictor 220 of the encoding apparatus can generate information related to signaling of a list of reference images. Alternatively, for example, the filter 260 of the encoding apparatus can generate at least one of information related to the SAO process, information related to the ALF process, and information related to the deblocking process.

[0167] The encoding apparatus according to the embodiment can encode image information including indication information and information related to at least one tool (S830). Additionally, the image information may include prediction information for the current block. The prediction information may include information about the inter-frame prediction mode or intra-frame prediction mode performed on the current block. Furthermore, the image information may include residual information generated from the original samples by the residual processor 230 of the encoding apparatus.

[0168] Furthermore, the bitstream containing image information can be transmitted to the decoding device via a network or (digital) storage medium. In this document, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.

[0169] Figure 10 This is a flowchart illustrating the operation of an image decoding device according to an embodiment, and Figure 11 This is a block diagram illustrating the configuration of an image decoding device according to an embodiment.

[0170] Figure 10 The method disclosed in the article can be derived from Figure 3 or Figure 11 The decoding is performed by the publicly disclosed decoding device. More specifically, it can be performed by... Figure 3 The entropy decoder 310 shown executes S1010 to S1030. Additionally, S1040 can be performed by... Figure 3 The predictor 330, residual processor 320, filter 350, or adder 340 shown are executed. Furthermore, the operations according to S1010 to S1040 are based on the above. Figures 1 to 7This is part of the description given. Therefore, for simplicity, it will be omitted or briefly given in relation to... Figures 1 to 7 The description overlaps with the detailed description.

[0171] The decoding apparatus according to the embodiment can obtain indication information (S1010) indicating whether at least one tool used for the current block is applied at the picture level or the slice level. For example, the indication information may include a flag indicating whether signaling for a list of reference pictures for the current block is applied at the picture level or the slice level. For example, the indication information may include a flag indicating whether a Sample Adaptive Offset (SAO) process is applied at the picture level or the slice level. For example, the indication information may include a flag indicating whether an Adaptive Loop Filter (ALF) process is applied at the picture level or the slice level. Alternatively, for example, the indication information may include at least one flag indicating whether a deblocking process is applied at the picture level or the slice level.

[0172] According to the embodiment, the decoding device can determine whether the information related to at least one tool exists in the image header or the slice header based on the indication information (S1020).

[0173] For example, if the flag indicating whether the signaling for the reference image list is applied at the image level or the slice level is valued at 1, it can be determined that information related to the signaling for the reference image list exists in the image header. Conversely, if the flag is valued at 0, it can be determined that information related to the signaling for the reference image list exists in the slice header.

[0174] For example, if the flag indicating whether the SAO process is applied at the image level or the slice level is equal to 1, it can be determined that information related to the SAO process exists in the image header. Conversely, if the flag is equal to 0, it can be determined that information related to the SAO process exists in the slice header.

[0175] For example, if the flag indicating whether the ALF procedure is applied at the image level or the slice level is equal to 1, it can be determined that information related to the ALF procedure exists in the image header. Conversely, if the flag is equal to 0, it can be determined that information related to the ALF procedure exists in the slice header.

[0176] Alternatively, for example, based on at least one flag indicating whether the deblocking process is applied at the image level or the slice level, information related to the deblocking process may exist in either the image header or the slice header. For example, if the value of at least one flag is equal to 1, it can be determined that the flag indicating whether parameters related to the deblocking process exist in the image header is present in the image header. And, if the value of at least one flag is equal to 0, it can be determined that the flag indicating whether parameters related to the deblocking process exist in the image header is not present in the image header. Alternatively, for example, if the value of at least one flag is equal to 1, it can be determined that the flag indicating whether parameters related to the deblocking process exist in the slice header is present in the slice header. And, if the value of at least one flag is equal to 0, it can be determined that the flag indicating whether parameters related to the deblocking process exist in the slice header is not present in the slice header.

[0177] According to the implementation, the decoding device can parse information related to at least one tool from the image header or slice header based on this determination (S1030).

[0178] The decoding device according to the embodiment can decode the current block based on information related to at least one tool (S1040). For example, based on signaling related to a list of reference images parsed by receiving one of the image header and slice header, the predictor 330 of the decoding device can perform a prediction on the current block. For example, based on information related to the SAO process parsed by receiving one of the image header and slice header, the filter 350 of the decoding device can perform the SAO process on the reconstructed sample. For example, based on information related to the ALF process parsed by receiving one of the image header and slice header, the filter 350 of the decoding device can perform the ALF process on the reconstructed sample. Alternatively, for example, based on information related to the deblocking process parsed by receiving one of the image header and slice header, the filter 350 of the decoding device can perform a deblocking process on the reconstructed sample.

[0179] Although the method has been described based on a flowchart listing the steps or blocks in the above embodiments, the steps in this document are not limited to a specific order, and specific steps may be performed in different steps or in different orders or simultaneously relative to the steps described above. Furthermore, those skilled in the art will understand that the steps in the flowchart are not exclusive, and one or more steps may be included or removed from the flowchart without affecting the scope of this disclosure.

[0180] The methods mentioned above according to this disclosure can be implemented in software form, and the encoding and / or decoding devices according to this disclosure can be included, for example, in an apparatus for performing image processing (e.g., TV, computer, smartphone, set-top box, display device, etc.).

[0181] When the embodiments of this disclosure are implemented in software, the methods described above can be implemented using modules (processes or functions) that perform the functions mentioned above. Modules can be stored in memory and executed by a processor. Memory can be installed internally or externally to the processor and can be connected to the processor via various known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, embodiments of this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information about the implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.

[0182] Furthermore, the decoding and encoding devices using the embodiments described in this document can be included in multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, portable cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, vehicle-mounted terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, or ship terminals), and medical video devices; and can be used to process image signals or data. For example, OTT video devices can include game consoles, Blu-ray players, networked TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).

[0183] Furthermore, the processing methods applying the embodiments of this document can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data with data structures according to the embodiments of this document can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Computer-readable recording media also include media implemented in the form of carrier waves (e.g., transmission over the Internet). Additionally, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks.

[0184] Furthermore, the embodiments described in this document can be implemented as a computer program product based on program code, and the program code can be executed on a computer according to the embodiments described in this document. The program code can be stored on a computer-readable medium.

[0185] Figure 12 Examples of content streaming systems to which the embodiments disclosed in this specification can be applied are illustrated.

[0186] refer to Figure 12 The content streaming system using the implementation methods described in this document can mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0187] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generate a bitstream, and then transmit it to a streaming server. As another example, if the multimedia input device, such as a smartphone, camera, or camcorder, directly generates the bitstream, the encoding server can be omitted.

[0188] The encoding method or bitstream generation method applied to the embodiments described in this document can be used to generate bitstreams. Furthermore, the streaming server can temporarily store the bitstream during the sending or receiving process.

[0189] A streaming server transmits multimedia data to a user's device via a web server based on a user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server forwards the request to the streaming server, which then delivers the multimedia data to the user. In this respect, the content streaming system may include a separate control server, which in this case controls the commands / responses between the various devices within the content streaming system.

[0190] A streaming server can receive content from media storage devices and / or encoding servers. For example, if content is received from an encoding server, it can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to provide a smooth streaming service.

[0191] For example, user equipment may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smartwatches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.

[0192] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.

[0193] The claims in this specification can be combined in various ways. For example, the technical features in the method claims can be combined to be implemented or performed in a device, and the technical features in the device claims can be combined to be implemented or performed in a method. Furthermore, the technical features in the method claims and the device claims can be combined to be implemented or performed in a device.

Claims

1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Obtain indication information related to whether the deblocking filter information exists in the image header or the slice header; Based on the indication information, determine whether the deblocking filter information exists in the image header or the slice header; Based on the determination, the deblocking filter information is parsed from the image header or the slice header; as well as The current block is decoded based on the deblocking filter information. The indication information is included in the image parameter set PPS. The image header includes information commonly applied to all slices in the image. The PPS includes information publicly applied to one or more images, and The syntax level in which the indication information is included is different from the syntax level in which the deblocking filter information is included.

2. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Generate deblocking filter information; Determine whether the deblocking filter information exists in the image header or the slice header; Generate indication information related to whether the deblocking filter information exists in the image header or the slice header; as well as The image information, including the indication information and the deblocking filter information, is encoded. The indication information is configured to indicate whether the deblocking filter information exists in the image header or the slice header. The indication information is included in the image parameter set PPS. The image header includes information commonly applied to all slices in the image. The PPS includes information publicly applied to one or more images, and The syntax level in which the indication information is included is different from the syntax level in which the deblocking filter information is included.

3. A method for transmitting image data, the method comprising the following steps: A bitstream for the image is generated by performing an image encoding method, the image encoding method comprising the following steps: Generate deblocking filter information, and determine whether the deblocking filter information exists in the image header or the slice header. Generate indication information related to whether the deblocking filter information exists in the image header or the slice header, and Encoding image information including the indication information and the deblocking filter information; and Send the data including the bit stream. The indication information is configured to indicate whether the deblocking filter information exists in the image header or the slice header. The indication information is included in the image parameter set PPS. The image header includes information commonly applied to all slices in the image. The PPS includes information publicly applied to one or more images, and The syntax level in which the indication information is included is different from the syntax level in which the deblocking filter information is included.