Method and device for signalling image information applied on picture level or slice level
By signaling image information at the picture or slice level, the method improves image coding efficiency and decoding performance, addressing the need for cost-effective high-resolution image transmission and storage.
Patent Information
- Application Number
- JP2025081985
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-12-12
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-12-10
AI Technical Summary
The increasing demand for high-resolution and high-quality images has led to a need for more efficient image compression technologies to reduce transmission and storage costs.
A method and apparatus for signaling image information at a picture or slice level, allowing tools to be applied at either the picture or slice level, enabling improved decoding efficiency by parsing information from the appropriate header based on indication signals.
Enhances image coding efficiency and decoding performance by optimizing tool application at the picture or slice level, reducing overall data transmission and storage costs.
Smart Images

Figure 2025113313000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image coding technology, and more particularly to a method and apparatus for signaling image information applied at a picture level or a slice level in an image coding system.
Background Art
[0002] In recent years, the demand for high-resolution and high-quality images such as HD (High Definition) images and UHD (Ultra High Definition) images has been increasing in various fields. As the image data becomes higher in resolution and quality, the amount of information or bits to be transmitted relatively increases compared to the existing image data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line, or storing image data using an existing storage medium, the transmission cost and the storage cost increase.
[0003] Accordingly, there is a need for a highly efficient image compression technology to effectively transmit, store, and reproduce high-resolution and high-quality image information.
Summary of the Invention
Problems to be Solved by the Invention
[0004] A technical problem of the present disclosure is to provide a method and apparatus for increasing image coding efficiency.
[0005] Another technical problem of the present disclosure is to provide a method and apparatus for signaling image information applied at a picture level or a slice level.
[0006] Still another technical problem of the present disclosure is to provide a method and apparatus for performing decoding on a current block based on image information applied at a picture level or a slice level.
Means for Solving the Problems
[0007] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method includes obtaining indication information indicating whether at least one tool for a current block is applied at a picture level or a slice level; determining, based on the indication information, in which of a picture header and a slice header the information related to the at least one tool exists; parsing, based on the determination result, the information related to the at least one tool from the picture header or the slice header; and decoding the current block based on the information related to the at least one tool.
[0008] According to another embodiment of the present disclosure, an image encoding method performed by an encoding device is provided. The method includes generating indication information indicating whether at least one tool applied to a current block is applied at a picture level or a slice level; generating information related to the at least one tool; and encoding image information including the indication information and the information related to the at least one tool, where the indication information indicates in which of a picture header and a slice header the information related to the at least one tool exists.
[0009] According to still another embodiment of the present disclosure, there is provided a computer-readable digital storage medium storing encoded image information that causes a decoding device to perform an image decoding method. The decoding method according to the one embodiment includes: obtaining instruction information indicating whether at least one tool applied to a current block is applied at a picture level or a slice level; determining, based on the instruction information, in which of a picture header and a slice header information related to the at least one tool exists; parsing, based on the determination result, information related to the at least one tool from the picture header or the slice header; and performing decoding on the current block based on the information related to the at least one tool.
Advantages of the Invention
[0010] According to this specification, general image / video compression efficiency can be improved.
[0011] According to this specification, the efficiency of image decoding can be improved based on instruction information indicating whether at least one tool for a current block is applied at a picture level or a slice level.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0013] This document can be modified in various ways, can have various embodiments, and specific embodiments are illustrated in the drawings and will be described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are merely used to describe specific embodiments and are not used with the intention of limiting the technical idea of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof, etc., is not precluded in advance.
[0014] On the one hand, each component in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each component is realized by separate hardware or separate software. For example, among the components, two or more components can be combined to form one component, and one component can also be divided into multiple components. As long as the implementation forms in which each component is integrated and / or separated do not deviate from the essence of this document, they are included in the scope of rights of this document.
[0015] In this specification, "A or B" can mean "only A", "only B", or "both A and B". In other words, in this specification, "A or B" can be interpreted as "A and / or B". For example, in this specification, "A, B, or C" can mean "only A", "only B", "only C", or "any combination of A, B, and C".
[0016] The slashes ( / ) and commas used in this specification can mean "and / or". For example, "A / B" can mean "A and / or B". Thus, "A / B" can mean "only A", "only B", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0017] As used herein, "at least one of A and B" can mean "only A," "only B," or "both A and B." Furthermore, as used herein, the expressions "at least one of A or B" and "at least one of A and / or B" can be interpreted as "at least one of A and B."
[0018] Furthermore, in this specification, "at least one of A, B, and C" can mean "only A," "only B," "only C," or "any combination of A, B, and C." Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C."
[0019] Furthermore, parentheses used herein may mean "for example." Specifically, when "prediction (intra prediction)" is used, "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra prediction," and "intra prediction" is proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is used, "intra prediction" is proposed as an example of "prediction."
[0020] In this specification, technical features individually described in one drawing may be realized individually or simultaneously.
[0021] Hereinafter, with reference to the accompanying drawings, preferred embodiments of the present disclosure will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and duplicate descriptions of the same components can be omitted.
[0022] FIG. 1 schematically shows an example of a video / image coding system to which the present disclosure can be applied.
[0023] As shown in FIG. 1, the video / image coding system can include a first device (source device) and a second device (receiver device). The source device can transmit encoded video / image information or data to the receiver device in a file or streaming form via a digital storage medium or a network.
[0024] The source device can include a video source, an encoding device, and a transmission unit. The receiver device can include a reception unit, a decoding device, and a renderer. The encoding device can be called a video / image encoding device, and the decoding device can be called a video / image decoding device. A transmitter can be included in the encoding device. A receiver can be included in the decoding device. The renderer can also include a display unit, and the display unit can also be composed of a separate device or an external component.
[0025] The video source can obtain video / images through processes such as video / image capture, synthesis, or generation. The video source can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate a video / image. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced by a process in which related data is generated.
[0026] The encoding device can encode the input video / image. The encoding device can execute a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0027] The transmitting unit can transmit the encoded video / image information or data output in the form of a bitstream to the receiving unit of the receiving device via a digital storage medium or a network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file through a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.
[0028] The decoding device can decode the video / image by executing a series of procedures such as inverse quantization, inverse transformation, prediction, etc. corresponding to the operation of the encoding device.
[0029] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.
[0030] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to the methods disclosed in the VVC (versatile video coding) standard, EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVS2 (2nd generation of audio video coding standard), or next-generation video / image coding standards (e.g., H.267 or H.268, etc.).
[0031] This document presents various embodiments related to video / image coding. Unless otherwise mentioned, the above embodiments, etc. can also be combined with each other.
[0032] In this document, video can mean a collection of a series of images, etc. over time. A picture generally means a unit representing one image at a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile can include one or more CTUs (coding tree units). One picture can be composed of one or more slices / tiles.
[0033] A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set. The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the width of the picture.A tile scan may represent a specific sequential ordering of CTUs partitioning a picture, where the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice may include multiple complete tiles or multiple consecutive CTU rows in one tile of a picture, which may be included in one NAL unit. In this document, the terms tile group and slice may be used interchangeably. For example, in this document, tile group / tile group header may be referred to as slice / slice header.
[0034] On the other hand, a picture can be divided into two or more sub-pictures, each of which can be a rectangular region of one or more slices within a picture.
[0035] A pixel or a pel can refer to the smallest unit that constitutes one picture (or image). A term corresponding to a pixel can also be used: "sample." A sample can generally represent a pixel or a pixel value, and can represent only a pixel / pixel value of a luma component, or can represent only a pixel / pixel value of a chroma component.
[0036] A unit can represent the basic unit of image processing. A unit can include at least one of a specific region of a picture and information related to the region. One unit can include one luma block and two chroma (e.g., cb, cr) blocks. A unit can, in some cases, be used interchangeably with terms such as block or area. In a general case, an M×N block can include a sample (or, sample array) consisting of M columns and N rows, or a set (or, array) of transform coefficients.
[0037] Figure 2 is a diagram schematically illustrating the configuration of a video / image encoding device to which this document can be applied. Hereinafter, the video encoding device can include an image encoding device.
[0038] As shown in FIG. 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor (231). The adder 250 can be called a reconstructor or a reconstructed block generator. The above-described image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.
[0039] The image segmentation unit 210 can divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the image characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that a coding unit of an optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can each be divided or partitioned from the above-described final coding unit.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.
[0040] The unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel for one picture (or image).
[0041] The encoding device 200 can subtract the prediction signal (predicted block, predicted sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, the unit that subtracts the prediction signal (predicted block, predicted sample array) from the input image signal (original block, original sample array) within the encoding device 200 can be called the subtraction unit 231. The prediction unit can perform a prediction on the processing target block (hereinafter referred to as the current block) and generate a predicted block including the predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 240 as described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bitstream.
[0042] The intra prediction unit 222 can predict the current block by referring to the samples within the current picture. The samples to be referred to can be located adjacent to the current block depending on the prediction mode, or can also be located far away. In intra prediction, the prediction mode can include a plurality of non - directional modes and a plurality of directional modes. The non - directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.
[0043] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (col CUs), and the reference picture including the temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0044] The prediction unit 220 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can apply not only intra prediction or inter prediction for the prediction of one block, but also can apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the intra block copy (IBC) prediction mode for the prediction of a block, or can be based on the palette mode. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values within the picture can be signaled based on the information regarding the palette table and the palette index.
[0045] The prediction signal generated via the prediction unit (including the inter-prediction unit 221 and / or the intra-prediction unit 222) can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include at least one of DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), KLT (Karhunen-Loeve Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when representing the relationship information between pixels as a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process can be applied to a pixel block having the same size of a square and can also be applied to a block of variable size that is not square.
[0046] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be referred to as residual information. The quantization unit 233 can reorder the block-form quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. In addition to the quantized transform coefficients, the entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) can be transmitted or stored in units of NAL (network abstraction layer) units in bitstream form. The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The information and / or syntax elements transmitted / signaled from the encoding device to the decoding device in this document can be included in the video / image information. The video / image information can be encoded through the encoding procedure described above and included in the bitstream.The bitstream can be transmitted via a network or stored in a digital storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured as an internal / external element of the encoding device 200 by a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage, or the transmission unit can also be included in the entropy encoding unit 240.
[0047] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transformation unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The addition unit 250 can be called a restoration unit or a reconstructed block generation unit. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and as will be described later, can also be used for inter prediction of the next picture after passing through filtering.
[0048] On the other hand, LMCS (luma mapping with chrom ascaling) can also be applied during the picture encoding and / or restoration process.
[0049] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, etc. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 290, as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 290 and output in the form of a bitstream.
[0050] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 280. When inter prediction is applied through this, the encoding device can avoid prediction mismatches between the encoding device 200 and the decoding device, and can also improve the encoding efficiency.
[0051] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks in which the motion information within the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks within the current picture and transmit them to the intra prediction unit 222.
[0052] FIG. 3 is a diagram illustrating the configuration of a video / image decoding device to which this document can be applied.
[0053] As shown in FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured as a single hardware component (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.
[0054] If a bitstream including video / image information is input, the decoding device 300 can restore an image corresponding to the process in which the video / image information is processed by the encoding device in FIG. 3. For example, the decoding device 300 can derive units / blocks based on the block splitting related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be split according to a quad tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored image signal decoded and output via the decoding device 300 can be played back via a playback device.
[0055] The decoding device 300 can receive the signal output from the encoding device of FIG. 3 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / image information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements necessary for image restoration, quantized values of transform coefficients regarding residuals, etc. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the surrounding and decoded information of the block to be decoded, or the information of symbols / bins decoded in the previous step, predicts the occurrence probability of a bin according to the determined context model, performs arithmetic decoding of the bin, and generates a symbol corresponding to the value of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin.Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit (inter prediction unit 332 and intra prediction unit 331), and the residual value obtained by performing entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficient and related parameter information, can be input to the residual processing unit 320. The residual processing unit 320 can derive a residual signal (residual block, residual sample, residual sample array). Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, inverse transform unit 322, addition unit 340, filtering unit 350, memory 360, inter prediction unit 332, and intra prediction unit 331.
[0056] In the inverse quantization unit 321, the quantized transform coefficient can be inverse quantized to output a transform coefficient. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed in the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficient using a quantization parameter (for example, quantization step size information) to obtain a transform coefficient.
[0057] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0058] The prediction unit can perform prediction on the current block and generate a predicted block including predicted samples for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.
[0059] The prediction unit 330 can generate a prediction signal based on various prediction methods described later. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can be based on the intra block copy (IBC) prediction mode or the palette mode for the prediction of a block. The IBC prediction mode or the palette mode can be used for content image / video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, information regarding the palette table and the palette index can be included in and signaled in the video / image information.
[0060] The intra prediction unit 332 can predict the current block by referring to samples within the current picture. The samples to be referred can be located adjacent to the current block or at a distance therefrom depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 332 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent block.
[0061] The inter prediction unit 331 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between the adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. For example, the inter prediction unit 331 can configure a motion information candidate list based on the adjacent blocks and derive the motion vector and / or the reference picture index of the current block based on the received candidate selection information. Inter prediction can be executed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of the inter prediction for the current block.
[0062] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the predictor 330. When there is no residual for the block to be processed as in the case where the skip mode is applied, the predicted block can be used as the restored block.
[0063] The adder 340 can be referred to as a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next block to be processed within the current picture, and as will be described later, it can also be output after passing through filtering, or can be used for inter prediction of the next picture.
[0064] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.
[0065] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be sent to the memory 60, specifically, the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0066] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 331. The memory 360 can store the motion information of the block for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the block in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 331 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 332.
[0067] In this specification, the embodiments described in the filtering unit 260, the inter prediction unit 221, and the intra prediction unit 222 of the encoding device 100 can be applied to the filtering unit 350, the inter prediction unit 332, and the intra prediction unit 331 of the decoding device 300 in the same or corresponding manner, respectively.
[0068] On the other hand, as described above, prediction is performed to improve the compression efficiency when performing video coding. Thereby, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is derived in the same manner in the encoding device and the decoding device, and the encoding device can improve the image coding efficiency by signaling information (residual information) regarding the residual between the original block, which is not the original sample value of the original block itself, and the predicted block to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, add the residual block and the predicted block to generate a restored block including restored samples, and generate a restored picture including the restored block.
[0069] The residual information can be generated through conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, and thus signal the relevant residual information (via a bitstream) to a decoding device. Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in inter prediction of subsequent pictures to derive a residual block and generate a restored picture based on this.
[0070] FIG. 4 exemplarily shows a hierarchical structure for coded data.
[0071] As shown in FIG. 4, the coded data can be divided into a VCL (video coding layer) that handles video / image coding processing and itself, and a NAL (Network abstraction layer) that is between the lower system that stores and transmits the coded video / image data.
[0072] VCL can generate parameter sets (Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc.) corresponding to headers such as sequences and pictures, and Supplemental Enhancement Information (SEI) messages that are additionally required in the video / image coding process. The SEI message is separated from the information (slice data) regarding the video / image. VCL containing information regarding the video / image consists of slice data and slice headers. On the other hand, the slice header can be referred to as a tile group header, and the slice data can be referred to as tile group data.
[0073] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the Raw Byte Sequence Payload (RBSP) generated by VCL. At this time, RBSP refers to slice data, parameter sets, SEI messages, etc. generated by VCL. The NAL unit header can contain NAL unit type information specified by the RBSP data included in the NAL unit.
[0074] The NAL unit, which is the basic unit of NAL, serves to map the coded image to a bit sequence of a lower-level system such as a file format according to a predetermined standard, Real-time Transport Protocol (RTP), Transport Stream (TS), etc.
[0075] As shown in the figure, NAL units can be classified into VCL NAL units and Non-VCL NAL units by the RBSP generated in the VCL. A VCL NAL unit can mean a NAL unit containing information (slice data) related to an image, and a Non-VCL NAL unit can mean a NAL unit containing information (parameter set or SEI message) necessary for decoding an image.
[0076] As described above, the above-mentioned VCL NAL units and Non-VCL NAL units can be transmitted via a network with header information attached according to the data standard of the lower system. For example, NAL units can be transformed into data forms of predetermined standards such as the H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and transmitted via various networks.
[0077] As described above, the NAL unit type can be specified by the RBSP data structure (structure) included in the NAL unit, and information regarding such a NAL unit type can be stored in the NAL unit header and signaled.
[0078] For example, depending on whether the NAL unit contains information (slice data) related to an image, it can be roughly classified into a VCL NAL unit type and a Non-VCL NAL unit type. The VCL NAL unit type can be classified according to the nature and type of the picture included in the VCL NAL unit, and the Non-VCL NAL unit type can be classified according to the type of the parameter set, etc.
[0079] The following is an example of the NAL unit type specified by the type of parameter set included in the Non-VCL NAL unit type, etc. The NAL unit type can be specified by the type of parameter set, etc. For example, the NAL unit type is the APS (Adaptation Parameter Set) NAL unit which is the type for the NAL unit including APS, the DPS (Decoding Parameter Set) NAL unit which is the type for the NAL unit including DPS, the VPS (Video Parameter Set) NAL unit which is the type for the NAL unit including VPS, the SPS (Sequence Parameter Set) NAL unit which is the type for the NAL unit including SPS, and the PPS (Picture Parameter Set) NAL unit which is the type for the NAL unit including PPS, and can be specified as any one of them.
[0080] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.
[0081] On one hand, as described above, one picture can include a plurality of slices, and one slice can include a slice header and slice data. In this case, one picture header can be further added for a plurality of slices (slice header and slice data set) within one picture. The picture header (picture header syntax) can include information / parameters that can be commonly applied to the picture. The slice header (slice header syntax) can include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) can include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) can include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) can include information / parameters that can be commonly applied to the entire video. The DPS can include information / parameters related to the concatenation of coded video sequences (CVS). In this document, the high level syntax (HLS) can include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, a picture header syntax, and slice header syntax.
[0082] In this document, the image / video information encoded by the encoding device and signaled in the form of a bitstream to the decoding device can include not only information related to partitioning within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but also information included in the slice header, information included in the Picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. Further, the image / video information can further include information of the NAL unit header.
[0083] On the other hand, as described above, the encoding device / decoding device can perform an in-loop filtering procedure on the restored picture in order to improve subjective / objective image quality. A restored picture modified through the in-loop filtering procedure can be generated, and the modified restored picture can be output as a decoded picture by the decoding device and can also be stored in the decoded picture buffer or memory of the encoding device / decoding device. Further, the restored picture modified later can be used as a reference picture in the inter prediction procedure when encoding / decoding. The in-loop filtering procedure can include, as described above, a deblocking filtering procedure, a sample adaptive offset (SAO) procedure, and / or an adaptive loop filter (ALF) procedure. In this case, one or some of the deblocking filtering procedure, the sample adaptive offset (SAO) procedure, and the adaptive loop filter (ALF) procedure can be sequentially applied, or all of them can be sequentially applied. For example, after the deblocking filtering procedure is applied to the restored picture, the SAO procedure can be performed. Or, for example, after the deblocking filtering procedure is applied to the restored picture, the ALF procedure can be performed. This can be similarly performed in the encoding device.
[0084] The deblocking filtering procedure is a procedure for removing distortion generated at the boundary between blocks in the restored picture. The deblocking filtering procedure can, for example, derive a target boundary in the restored picture, determine a boundary strength (bS) for the target boundary, and perform deblocking filtering on the target boundary based on the bS. The bS can be determined based on, for example, the prediction mode of two blocks adjacent to the target boundary, the motion vector difference, whether the reference pictures are the same, and whether there are non-zero valid coefficients.
[0085] SAO is a method for compensating for the offset difference between the restored picture and the original picture on a sample-by-sample basis. For example, SAO can be applied according to types such as band offset and edge offset. According to SAO, samples can be classified into different categories according to the SAO type, and an offset value can be added to each sample according to the category. The filtering information for SAO can include whether SAO can be applied, SAO type information, SAO offset value information, etc. For example, SAO can be applied to the restored picture after deblocking filtering is applied.
[0086] The ALF (Adaptive Loop Filter) procedure is a procedure for filtering the restored picture on a sample-by-sample basis based on filter coefficients according to the filter shape. The encoding device can compare the restored picture with the original picture to determine whether to apply ALF, the ALF shape and / or ALF filtering coefficients, etc., and can signal them to the decoding device. That is, the filtering information for ALF can include whether ALF can be applied, ALF filter shape information, ALF filtering coefficient information, etc. The ALF procedure can be applied to the restored picture after deblocking filtering is applied.
[0087] FIG. 5 is a flowchart showing an embodiment of a method for performing deblocking filtering.
[0088] As described above, the encoding device / decoding device can restore a picture in units of blocks. When such block-based image restoration is performed, block distortion may occur at the boundaries between blocks in the restored picture. Therefore, the encoding device and the decoding device can use a deblocking filter to remove the block distortion occurring at the boundaries (boundaries) between blocks in the restored picture.
[0089] Therefore, the encoding device / decoding device can derive the boundaries between blocks in the restored picture where deblocking filtering is performed. On the other hand, the boundary where deblocking filtering is performed can be called an edge. Also, the boundary where the deblocking filtering is performed can include two types, and the two types can be a vertical boundary and a horizontal boundary. The vertical boundary can be called a vertical edge, and the horizontal boundary can be called a horizontal edge. The encoding device / decoding device can perform deblocking filtering on the vertical edge and can perform deblocking filtering on the horizontal edge.
[0090] For example, the encoding device / decoding device can derive a target boundary to be filtered in the restored picture (S510).
[0091] Also, the encoding device / decoding device can determine the boundary strength (bS) for the boundary where deblocking filtering is performed (S520). bS can also be expressed as the boundary filtering strength. For example, it can be assumed that the case of obtaining the bS value for the boundary (block edge) between block P and block Q is considered. In this case, the encoding device / decoding device can obtain the bS value for the boundary (block edge) between block P and block Q based on block P and block Q. For example, bS can be determined by the following table.
[0092]
Table 1-1
[0093]
Table 1-2
[0094] Here, p can represent the samples of block P adjacent to the boundary to be deblocked filtered, and q can represent the samples of block Q adjacent to the boundary to be deblocked filtered.
[0095] Also, for example, the p0 can represent samples of a block adjacent to the left or upper side of the deblocking filtering target boundary, and the q0 can represent samples of a block adjacent to the right or lower side of the deblocking filtering target boundary. As an example, when the direction of the target boundary is the vertical direction (i.e., when the target boundary is a vertical boundary), the p0 can represent samples of a block adjacent to the left side of the deblocking filtering target boundary, and the q0 can represent samples of a block adjacent to the right side of the deblocking filtering target boundary. Or, as an example, when the direction of the target boundary is the horizontal direction (i.e., when the target boundary is a horizontal boundary), the p0 can represent samples of a block adjacent to the upper side of the deblocking filtering target boundary, and the q0 can represent samples of a block adjacent to the lower side of the deblocking filtering target boundary.
[0096] Furthermore, referring to FIG. 5, the encoding device / decoding device can perform deblocking filtering based on the bS (S530). For example, if the bS value is 0, filtering is not applied to the target boundary. On the other hand, based on the determined bS value, a filter applied to the boundary between blocks can be determined. The filter can be divided into a strong filter and a weak filter. The encoding device / decoding device can increase the coding efficiency by performing filtering with different filters for the boundary at a position where the probability of block distortion occurring in the reconstructed picture is high and the boundary at a position where the probability of block distortion occurring is low.
[0097] FIG. 6 is a flowchart schematically showing an example of the ALF procedure. The ALF procedure disclosed in FIG. 6 can be performed by an encoding device and a decoding device. In this document, the coding device can include the encoding device and / or the decoding device.
[0098] As shown in FIG. 6, the coding device derives a filter for ALF (S610). The filter can include filter coefficients. The coding device can determine whether ALF can be applied, and when it is determined to apply the ALF, it can derive a filter including the filter coefficients for the ALF. The filter (coefficients) for ALF or the information for deriving the filter (coefficients) for ALF can be called ALF parameters. Information regarding whether ALF can be applied (e.g., ALF available flag) and the ALF data for deriving the filter can be signaled from the encoding device to the decoding device. The ALF data can include information for deriving the filter for the ALF. Also, as an example, for hierarchical control of ALF, the ALF available flag can be signaled at the SPS, picture header, slice header, and / or CTB level, respectively.
[0099] To derive the filter for the ALF, the activity and / or directivity of the current block (or ALF target block) are derived, and the filter can be derived based on the activity and / or the directivity. For example, the ALF procedure can be applied in units of 4×4 blocks (based on luma components). The current block or ALF target block can be, for example, a CU, or a 4×4 block within a CU. Specifically, for example, a filter for ALF can be derived based on a first filter derived from the information included in the ALF data and a predefined second filter, and the coding device can select one of the filters based on the activity and / or the directivity. The coding device can use the filter coefficients included in the selected filter for the ALF.
[0100] The coding device performs filtering based on the filter (S620). Restored samples modified based on the filtering can be derived. For example, the filter coefficients in the filter can be arranged or assigned according to the filter shape, and the filtering can be performed on the restored samples in the current block. Here, the restored samples in the current block can be the restored samples after the deblocking filter procedure and the SAO procedure are completed. As an example, one filter shape can be used, or one filter shape can be selected and used from a predetermined plurality of filter shapes. For example, the filter shape applied to the luma component can be different from the filter shape applied to the chroma component. For example, a 7×7 diamond filter shape can be used for the luma component, and a 5×5 diamond filter shape can be used for the chroma component.
[0101] FIG. 7 shows an example of a filter shape for the ALF. C0 to C11 in (a) and C0 to C5 in (b) can be filter coefficients depending on the positions within each filter shape.
[0102] Fig. 7(a) shows a 7×7 diamond filter shape, and (b) shows a 5×5 diamond filter shape. In Fig. 7, Cn within the filter shape indicates the filter coefficient. When n is the same for the Cn, this represents that the same filter coefficient can be assigned. In this document, the position and / or unit where the filter coefficient is assigned according to the filter shape of the ALF can be called a filter tab. At this time, one filter coefficient can be assigned to each filter tab, and the form in which the filter tabs are arranged can correspond to the filter shape. The filter tab located at the center of the filter shape can be called the center filter tab. The same filter coefficient can be assigned to two filter tabs with the same n value existing at positions corresponding to each other with the center filter tab as a reference. For example, in the case of a 7×7 diamond filter shape, it includes 25 filter tabs, and since the filter coefficients of C0 to C11 are assigned in a centrally symmetric form, the filter coefficients of the 25 filter tabs can be assigned with only 13 filter coefficients. Also, for example, in the case of a 5×5 diamond filter shape, it includes 13 filter tabs, and since the filter coefficients of C0 to C5 are assigned in a centrally symmetric form, the filter coefficients of the 13 filter tabs can be assigned with only 7 filter coefficients. For example, in order to reduce the data amount of information regarding the signaled filter coefficient, among the 13 filter coefficients for the 7×7 diamond filter shape, 12 filter coefficients can be (explicitly) signaled, and 1 filter coefficient can be (implicitly) derived. Also, for example, among the 7 filter coefficients for the 5×5 diamond filter shape, 6 filter coefficients can be (explicitly) signaled, and 1 filter coefficient can be (implicitly) derived.
[0103] According to an embodiment of this document, the ALF parameters used for the ALF procedure can be signaled via an APS (adaptation parameter set). The ALF parameters can be derived from the filter information or ALF data for the ALF.
[0104] As described above, ALF is a type of in-loop filtering technique that can be applied in video / image coding. ALF can be performed using a Wiener-based adaptive filter. This can be for minimizing the mean square error (MSE) between the original sample and the decoded sample (or, the restored sample). The high level design for the ALF tool can incorporate syntax elements accessible in the SPS and / or slice header (or, tile group header).
[0105] On the other hand, the picture header contains syntax elements applied in the picture header, and the syntax elements can be applied to all slices of the picture related to the picture header. If a specific syntax element is applied only to a specific slice, the specific syntax element must be signaled in a slice header that is not a picture header.
[0106] Conventionally, the signaling of control flags and parameters for enabling or disabling various tools for picture encoding or decoding could be present in the picture header and could be overridden in the slice header. Such a scheme provides the flexibility that tool control can be performed both at the picture level and at the slice level. However, such a scheme can impose a burden on the decoder because the slice header has to be checked after checking the picture header.
[0107] Accordingly, one embodiment of the present invention proposes indication information indicating whether at least one tool is applied at the picture level or the slice level. At this time, the indication information can be included in any one of the SPS (Sequence Parameter Set) and PPS (Picture Parameter Set). That is, when a specific tool is activated within the CLVS, an indication or flag indicating whether the specific tool is applied at the picture level or the slice level can be signaled in a parameter set such as the SPS or PPS. The indication or flag can be for one tool, but is not limited thereto. For example, an indication or flag indicating whether all tools other than the specific tool are applied at the picture level or the slice level can be signaled in a parameter set such as the SPS or PPS.
[0108] Control flags and parameters for tool activation or deactivation can be signaled at the picture level or the slice level, but cannot be signaled at both the picture level and the slice level. For example, when obtaining indication information indicating whether a specific tool is applied at the picture level, the control flags and parameters for tool activation or deactivation can be signaled only in the picture header. Similarly, when obtaining indication information indicating whether a specific tool is applied at the slice level, the control flags and parameters for tool activation or deactivation can be signaled only in the slice header.
[0109] Also, for example, a tool designated to be applied at the picture level in a specific parameter set can be designated to be applied at the slice level in another parameter set of the same type.
[0110] For example, the PPS syntax including the indication information can be as shown in the following table.
[0111] [Table 2]
[0112] The semantics of the syntax elements included in the syntax of Table 2 above can be represented, for example, as shown in Table 3 below.
[0113] [Table 3]
[0114] Referring to the above table, the indication information can include a flag indicating whether the signaling of the reference picture list is applied at the picture level or the slice level. For example, the indication information can specify whether the information related to the signaling of the reference picture list exists in the picture header or the slice header. For example, the flag can be referred to as rpl_present_in_ph_flag. Based on the case where the value of the flag is the same as 1, the information related to the signaling of the reference picture list exists in the picture header, and based on the case where the value of the flag is the same as 0, the information related to the signaling of the reference picture list can exist in the slice header.
[0115] Also, the indication information can include a flag indicating whether the SAO (Sample Adaptive Offset) procedure is applied at the picture level or the slice level. For example, the indication information can specify whether the information related to the SAO procedure exists in the picture header or the slice header. For example, the flag can be referred to as sao_present_in_ph_flag. Based on the case where the value of the flag is the same as 1, the information related to the SAO procedure exists in the picture header, and based on the case where the value of the flag is the same as 0, the information related to the SAO procedure can exist in the slice header.
[0116] Also, the indication information can include a flag indicating whether the ALF (Adaptive Loop Filter) procedure is applied at the picture level or the slice level. For example, the indication information can specify whether information related to the ALF procedure is present in the picture header or the slice header. For example, the flag can be referred to as alf_present_in_ph_flag. Based on the case where the value of the flag is the same as 1, information related to the ALF procedure can be present in the picture header, and based on the case where the value of the flag is the same as 0, information related to the ALF procedure can be present in the slice header.
[0117] Also, the indication information can include at least one flag indicating whether the deblocking procedure is applied at the picture level or the slice level. For example, based on the at least one flag, information related to the deblocking procedure can be present in either one of the picture header and the slice header. For example, the at least one flag can be referred to as deblocking_filter_ph_override_enabled_flag or deblocking_filter_sh_override_enabled_flag. For example, based on the case where the value of the at least one flag is the same as 1, a flag indicating whether parameters related to the deblocking procedure are present in the picture header can be present in the picture header, and based on the case where the value of the at least one flag is the same as 0, a flag indicating whether parameters related to the deblocking procedure are present in the picture header can be absent from the picture header.
[0118] Alternatively, based on whether the value of the at least one flag is the same as 1, a flag indicating whether parameters related to the deblocking procedure are present in the slice header can be present in the slice header, and based on whether the value of the at least one flag is the same as 0, a flag indicating whether parameters related to the deblocking procedure are present in the slice header can be absent from the slice header. At this time, the values of deblocking_filter_ph_override_enabled_flag and deblocking_filter_sh_override_enabled_flag can both not be the same as 1.
[0119] On the other hand, the picture header syntax can be as follows in the following table.
[0120]
Table 4-1
[0121]
Table 4-2
[0122] The semantics of the syntax elements included in the syntax of Table 4 above can be represented, for example, as in Table 5 below.
[0123]
Table 5
[0124] Referring to the above table, when the value of deblocking_filter_ph_override_enabled_flag, which corresponds to the flag indicating whether the deblocking procedure is applied at the picture level, is 1, pic_deblocking_filter_override_present_flag can be signaled. When the value of pic_deblocking_filter_override_present_flag is 1, pic_deblocking_filter_override_flag, which corresponds to the flag indicating whether the parameters related to the deblocking procedure exist in the picture header, can exist in the picture header. Or, when the value of pic_deblocking_filter_override_present_flag is 0, pic_deblocking_filter_override_flag, which corresponds to the flag indicating whether the parameters related to the deblocking procedure exist in the picture header, can not exist in the picture header.
[0125] Also, when the value of pic_deblocking_filter_override_flag, which corresponds to the flag indicating whether the parameters related to the deblocking procedure exist in the picture header, is the same as 1, deblocking parameters can exist in the picture header. When the value of pic_deblocking_filter_override_flag is the same as 0, deblocking parameters can not exist in the picture header.
[0126] Also, when the value of pic_deblocking_filter_disabled_flag is the same as 1, the deblocking filter can not be applied to the slice associated with the picture header. When the value of pic_deblocking_filter_disabled_flag is the same as 0, the deblocking filter can be applied to the slice associated with the picture header.
[0127] Also, pic_beta_offset_div2 and pic_tc_offset_div2 can specify the deblocking parameter offsets for β and tC (divided by 2) for slices associated with the picture header. The values of pic_beta_offset_div2 and pic_tc_offset_div2 are both in the range from -6 to 6.
[0128] On the other hand, the slice header syntax can be as follows.
[0129]
Table 6-1
[0130]
Table 6-2
[0131] The semantics of the syntax elements included in the syntax of Table 6 above can be represented, for example, as in Table 7 below.
[0132]
Table 7
[0133] Referring to the above table, when the value of deblocking_filter_sh_override_enabled_flag, which corresponds to a flag indicating whether the deblocking procedure is applied at the slice level, is 1, slice_deblocking_filter_override_present_flag can be signaled. When the value of slice_deblocking_filter_override_present_flag is 1, slice_deblocking_filter_override_flag, which corresponds to a flag indicating whether parameters related to the deblocking procedure exist in the slice header, can exist in the slice header. Or, when the value of slice_deblocking_filter_override_present_flag is 0, slice_deblocking_filter_override_flag, which corresponds to a flag indicating whether parameters related to the deblocking procedure exist in the slice header, can not exist in the picture header.
[0134] Also, when the value of slice_deblocking_filter_override_flag, which corresponds to a flag indicating whether parameters related to the deblocking procedure exist in the slice header, is the same as 1, deblocking parameters can exist in the slice header. When the value of slice_deblocking_filter_override_flag is the same as 0, deblocking parameters can not exist in the slice header.
[0135] Also, when the value of slice_deblocking_filter_disabled_flag is the same as 1, the deblocking filter can not be applied to the slice associated with the slice header. When the value of slice_deblocking_filter_disabled_flag is the same as 0, the deblocking filter can be applied to the slice associated with the slice header.
[0136] In addition, slice_beta_offset_div2 and slice_tc_offset_div2 can specify deblocking parameter offsets for β and tC (divided by 2) for the slice. The values of slice_beta_offset_div2 and slice_tc_offset_div2 are both within the range of -6 to 6.
[0137] FIG. 8 is a flowchart showing the operation of the encoding device according to one embodiment, and FIG. 9 is a block diagram showing the configuration of the encoding device according to one embodiment.
[0138] The method disclosed in FIG. 8 may be performed by the encoding device disclosed in FIG. 2 or FIG. 9. S810 and S820 in FIG. 8 may be performed by the image prediction unit 220, the residual processing unit 230, or the filtering unit 260 disclosed in FIG. 2, and S830 may be performed by the entropy encoding unit 240 disclosed in FIG. 2. Furthermore, the operations of S810 to S830 are based on some of the content described above in FIGS. 1 to 7. Therefore, the description of specific content that overlaps with the content described above in FIGS. 1 to 7 will be omitted or simplified.
[0139] As shown in FIG. 8, an encoding apparatus according to an embodiment may generate indication information indicating whether at least one tool applied to a current block is applied at a picture level or a slice level (S810).
[0140] For example, the image prediction unit 220 of the encoding device can generate instruction information including a flag indicating whether the signaling of the reference picture list is applied at the picture level or the slice level. For example, based on the case where the value of the flag is the same as 1, the information related to the signaling of the reference picture list exists in the picture header, and based on the case where the value of the flag is the same as 0, the information related to the signaling of the reference picture list can exist in the slice header.
[0141] For example, the filtering unit 260 of the encoding device can generate instruction information including a flag indicating whether the SAO (Sample Adaptive Offset) procedure is applied at the picture level or the slice level. For example, based on the case where the value of the flag is the same as 1, the information related to the SAO procedure exists in the picture header, and based on the case where the value of the flag is the same as 0, the information related to the SAO procedure can exist in the slice header.
[0142] For example, the filtering unit 260 of the encoding device can generate instruction information including a flag indicating whether the ALF (Adaptive Loop Filter) procedure is applied at the picture level or the slice level. For example, based on the case where the value of the flag is the same as 1, the information related to the ALF procedure exists in the picture header, and based on the case where the value of the flag is the same as 0, the information related to the ALF procedure can exist in the slice header.
[0143] Alternatively, for example, the filtering unit 260 of the encoding device can generate instruction information including at least one flag indicating whether the deblocking procedure is applied at the picture level or the slice level. Based on the at least one flag, information related to the deblocking procedure can exist in either the picture header or the slice header. For example, based on the case where the value of the at least one flag is the same as 1, a flag indicating whether parameters related to the deblocking procedure exist in the picture header exists in the picture header, and based on the case where the value of the at least one flag is the same as 0, a flag indicating whether parameters related to the deblocking procedure exist in the picture header may not exist in the picture header.
[0144] An encoding device according to an embodiment can generate information related to at least one tool (S820). For example, the image prediction unit 220 of the encoding device can generate information related to the signaling of the reference picture list. Alternatively, for example, the filtering unit 260 of the encoding device can generate at least one of information related to the SAO procedure, information related to the ALF procedure, and information related to the deblocking procedure.
[0145] An encoding device according to an embodiment can encode image information including instruction information and information related to at least one tool (S830). Also, the image information can include prediction information for the current block. The prediction information can include information related to the inter prediction mode or the intra prediction mode performed on the current block. Also, the image information can include residual information generated from the original samples by the residual processing unit 230 of the encoding device.
[0146] On the one hand, the bitstream encoded with the image information can be transmitted to the decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0147] FIG. 10 is a flowchart showing the operation of the decoding device according to an embodiment, and FIG. 11 is a block diagram showing the configuration of the decoding device according to an embodiment.
[0148] The method disclosed in FIG. 10 can be performed by the decoding device disclosed in FIG. 3 or FIG. 11. Specifically, S1010 to S1030 can be performed by the entropy decoding unit 310 disclosed in FIG. 3. Also, S1040 can be performed by the prediction unit 330, the residual processing unit 320, the filtering unit 350, or the addition unit 340 disclosed in FIG. 3. Furthermore, the operations by S1010 and S1040 are based on a part of the content described above in FIGS. 1 to 7. Therefore, the specific content overlapping with the content described above in FIGS. 1 to 7 will be omitted or simplified in the description.
[0149] The decoding device according to one embodiment can obtain instruction information indicating whether at least one tool for the current block is applied at the picture level or the slice level (S1010). For example, the instruction information may include a flag indicating whether the signaling of the reference picture list is applied at the picture level or the slice level. For example, the instruction information may include a flag indicating whether the SAO (Sample Adaptive Offset) procedure is applied at the picture level or the slice level. For example, the instruction information may include a flag indicating whether the ALF (Adaptive Loop Filter) procedure is applied at the picture level or the slice level. Or, for example, the instruction information may include at least one flag indicating whether the deblocking procedure is applied at the picture level or the slice level.
[0150] The decoding device according to one embodiment can determine whether information related to at least one tool exists in either the picture header or the slice header based on the instruction information (S1020).
[0151] For example, based on the case where the value of the flag indicating whether the signaling of the reference picture list is applied at the picture level or the slice level is the same as 1, it can be determined that the information related to the signaling of the reference picture list exists in the picture header, and based on the case where the value of the flag is the same as 0, it can be determined that the information related to the signaling of the reference picture list exists in the slice header.
[0152] For example, based on the case where the value of the flag indicating whether the SAO procedure is applied at the picture level or the slice level is the same as 1, it can be determined that the information related to the SAO procedure exists in the picture header, and based on the case where the value of the flag is the same as 0, it can be determined that the information related to the SAO procedure exists in the slice header.
[0153] For example, based on the case where the value of a flag indicating whether the ALF procedure is applied at the picture level or the slice level is the same as 1, the information related to the ALF procedure exists in the picture header, and based on the case where the value of the flag is the same as 0, it can be determined that the information related to the ALF procedure exists in the slice header.
[0154] Or, for example, based on at least one flag indicating whether the deblocking procedure is applied at the picture level or the slice level, the information related to the deblocking procedure can exist in either one of the picture header and the slice header. For example, based on the case where the value of the at least one flag is the same as 1, a flag indicating whether the parameters related to the deblocking procedure exist in the picture header exists in the picture header, and based on the case where the value of the at least one flag is the same as 0, it can be determined that the flag indicating whether the parameters related to the deblocking procedure exist in the picture header does not exist in the picture header. Or, for example, based on the case where the value of the at least one flag is the same as 1, a flag indicating whether the parameters related to the deblocking procedure exist in the slice header exists in the slice header, and based on the case where the value of the at least one flag is the same as 0, it can be determined that the flag indicating whether the parameters related to the deblocking procedure exist in the slice header does not exist in the slice header.
[0155] The decoding apparatus according to an embodiment can parse information related to at least one tool from a picture header or a slice header based on the determination result (S1030).
[0156] The decoding device according to an embodiment can perform decoding on the current block based on information related to at least one tool (S1040). For example, the prediction unit 330 of the decoding device can perform prediction on the current block based on information related to the signaling of the reference picture list obtained by receiving and parsing either one of the picture header and the slice header. For example, the filtering unit 350 of the decoding device can perform the SAO procedure on the restored samples based on information related to the SAO procedure obtained by receiving and parsing either one of the picture header and the slice header. For example, the filtering unit 350 of the decoding device can perform the ALF procedure on the restored samples based on information related to the ALF procedure obtained by receiving and parsing either one of the picture header and the slice header. Or, for example, the filtering unit 350 of the decoding device can perform the deblocking procedure on the restored samples based on information related to the deblocking procedure obtained by receiving and parsing either one of the picture header and the slice header.
[0157] In the above-described embodiment, the method is described based on a flowchart in a series of steps or blocks, but the corresponding embodiment is not limited to the order of the steps, and a certain step can occur in a different order or simultaneously with steps different from the above. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, other steps are included, or one or more steps of the flowchart can be deleted without affecting the scope of the embodiments of this document.
[0158] The method according to the above-described embodiments of this document can be realized in software form, and the encoding device and / or decoding device according to this document can be included in devices that perform image processing such as TVs, computers, smartphones, set-top boxes, display devices, etc.
[0159] In this document, when an embodiment is realized by software, the foregoing method can be realized by modules (processes, functions, etc.) that perform the foregoing functions. The modules can be stored in a memory and executed by a processor. The memory can be inside or outside the processor and can be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), a RAM (random access memory), a flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be realized and executed on a processor, a microprocessor, a controller, or a chip. For example, the functional units shown in each drawing can be realized and executed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, the information for realization (for example, information on instructions) or algorithms can be stored in a digital storage medium.
[0160] In addition, the decoding device and encoding device to which the embodiments of this document are applied can be included in multimedia broadcast transmission / reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, pay-per-view (VoD) service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, picture phone video devices, transportation means terminals (e.g., vehicle terminals including autonomous driving vehicles, airplane terminals, ship terminals, etc.), and medical video devices, etc., and can be used to process video signals or data signals. For example, as an over-the-top (OTT) video device, it can include game consoles, Blu-ray players, Internet-connected TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0161] In addition, the processing method to which the embodiments of this document are applied can be produced in the form of a program executed by a computer and can be stored in a computer-readable recording medium. Also, multimedia data having a data structure according to the embodiments of this document can be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which data that can be read by a computer is stored. The computer-readable recording medium can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Further, the computer-readable recording medium includes a medium realized in the form of a carrier wave (for example, transmission via the Internet). Also, a bit stream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0162] In addition, the embodiments of this document can be realized by a computer program product with program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.
[0163] FIG. 12 shows an example of a content streaming system to which the disclosure of this document can be applied.
[0164] As shown in FIG. 12, the content streaming system to which the present disclosure is applied can generally include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0165] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and plays the role of transmitting this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server can be omitted.
[0166] The bitstream can be generated by an encoding method or a bitstream generation method applied to the embodiments of this document, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0167] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server plays the role of a medium to inform the user of what services are available. When the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include a separate control server. In this case, the control server plays the role of controlling commands / responses between each device in the content streaming system.
[0168] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0169] Examples of the user device include a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (e.g., a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, and a digital signage.
[0170] Each server in the content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.
[0171] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented in a device, and the technical features of the device claims in this specification can be combined and implemented in a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented in a method.
Claims
1. In an image decoding method performed by a decoding apparatus, obtaining instruction information indicating whether at least one tool for a current block is applied at a picture level or whether the at least one tool for the current block is applied at a slice level; judging, based on the instruction information, in which of a picture header and a slice header information related to the at least one tool exists; parsing the information related to the at least one tool from the picture header or the slice header based on the judgment; decoding the current block based on the information related to the at least one tool, comprising: the instruction information is included in a PPS (Picture Parameter Set); the picture header includes information commonly applied to all slices in a picture; the PPS includes information commonly applied to one or more pictures.
2. In an image encoding method performed by an encoding apparatus, generating instruction information indicating whether at least one tool applied to a current block is applied at a picture level or whether the at least one tool for the current block is applied at a slice level; generating information related to the at least one tool; encoding image information including the instruction information and the information related to the at least one tool, comprising: the instruction information indicates in which of a picture header and a slice header the information related to the at least one tool exists; the instruction information is included in a PPS (Picture Parameter Set); the picture header includes information commonly applied to all slices in a picture; the PPS includes information commonly applied to one or more pictures.
3. A method for transmitting data for an image, a method for generating a bitstream for the image, the bitstream comprising: generating instruction information indicating whether at least one tool applied to the current block is applied at the picture level or whether the at least one tool for the current block is applied at the slice level; generating information related to the at least one tool; encoding image information including the instruction information and the information related to the at least one tool; and sending the data including the bitstream, wherein the instruction information indicates in which of a picture header and a slice header the information related to the at least one tool exists; wherein the instruction information is included in a PPS (Picture Parameter Set); wherein the picture header includes information commonly applied to all slices in a picture; wherein the PPS includes information commonly applied to one or more pictures.
Citation Information
Patent Citations
Signaling of non-picture-level syntax elements at the picture level
WO2021022264A1