Image coding apparatus and method based on filtering

Efficient filtering methods for high-resolution images/videos, including deblocking and ALF, address the high data volume challenge by enhancing compression efficiency and visual quality through independent subpicture coding and virtual boundary-based in-loop filtering.

JP7856829B2Active Publication Date: 2026-05-11LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-07-17
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality images/videos, including VR and AR content, has led to higher transmission and storage costs due to increased data volume, necessitating a highly efficient image/video compression technology.

Method used

Implementing an efficient filtering method that includes deblocking, SAO, and ALF, with in-loop filtering based on virtual boundaries, and coding subpictures independently to enhance compression efficiency and subjective/objective visual quality.

Benefits of technology

Improves overall image/video compression efficiency and subjective/objective visual quality by efficiently performing virtual boundary-based in-loop filtering and signaling subpicture-related information, reducing hardware resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007856829000017
    Figure 0007856829000017
  • Figure 0007856829000018
    Figure 0007856829000018
  • Figure 0007856829000019
    Figure 0007856829000019
Patent Text Reader

Abstract

To provide a method for increasing image / video coding efficiency.SOLUTION: According to the embodiments of the present document, etc., sub-pictures and / or virtual boundaries can be used for coding an image. For example, sub-pictures in the current picture can be used for predicting, reconstructing, and / or filtering the current picture. Virtual boundaries can be used for filtering reconstructed samples of the current picture. Through image coding based on the sub-pictures and / or virtual boundaries according to the embodiments of the present document, etc., the subjective / objective quality of an image can be improved, and the consumption of hardware resources necessary for the coding can be reduced.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0006] , ,

[0001] This document relates to an image coding apparatus and method based on filtering.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality images / videos such as 4K or UHD (Ultra High Definition) images / videos of 8K or higher has been increasing in various fields. As the image / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to the existing image / video data. Therefore, when transmitting image data using a medium such as an existing wired or wireless broadband line or storing image / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of images / videos having image characteristics different from real images, such as game images, has been increasing.

[0004] Thus, there is a need for a highly efficient image / video compression technology to effectively compress, transmit, store, and reproduce the information of high-resolution and high-quality images / videos having various characteristics as described above.

[0005] Specifically, in order to improve subjective / objective visual quality, an in-loop filtering procedure is performed, and there is a discussion on a scheme for improving the signaling efficiency of information for performing in-loop filtering based on virtual boundaries. In addition, the application of sub-pictures for improving the prediction and restoration performance in image coding is being studied.

Summary of the Invention

Means for Solving the Problems

[0006] According to one embodiment of this document, a method and apparatus for improving the efficiency of image / video coding are provided.

[0007] According to one embodiment of this document, an efficient filtering method and apparatus are provided.

[0008] According to one embodiment of this document, a method and apparatus for efficiently applying deblocking, SAO (sample adaptive loop), and ALF (adaptive loop filtering) are provided.

[0009] According to one embodiment of this document, in-loop filtering is performed based on a virtual boundary.

[0010] According to one embodiment of this document, whether or not the SPS (sequence parameter set) includes additional virtual boundary-related information (e.g., information regarding the number and location of virtual boundaries) is determined based on whether or not resampling is available for the reference picture.

[0011] According to one embodiment of this document, image coding is performed based on a subpicture.

[0012] According to one embodiment of this document, subpictures used for image coding are coded independently.

[0013] According to one embodiment of this document, a picture contains only one subpicture, and the subpicture is coded independently.

[0014] According to one embodiment of this document, a picture is generated based on a sub-picture merging procedure, and the sub-picture may be independently coded sub-pictures.

[0015] According to one embodiment of this document, subpictures used for image coding are each treated as pictures.

[0016] According to one embodiment of this document, an encoding device for video / image encoding is provided.

[0017] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded video / image information generated by a video / image encoding method disclosed in at least one embodiment of this document.

[0018] According to one embodiment of this document, a computer-readable digital storage medium is provided which stores encoded information or encoded video / image information that causes a decoding device to perform a video / image decoding method disclosed in at least one embodiment of this document. [Effects of the Invention]

[0019] According to one embodiment of this document, the overall image / video compression efficiency can be improved.

[0020] According to one embodiment of this document, subjective / objective visual quality can be enhanced through efficient filtering.

[0021] According to one embodiment of this document, a virtual boundary-based in-loop filtering procedure can be performed efficiently, and filtering performance can be improved.

[0022] According to one embodiment of this document, information for virtual boundary-based in-loop filtering can be efficiently signaled.

[0023] According to one embodiment of this document, subpicture-related information is efficiently signaled, thereby improving the subjective / objective quality of the image and reducing the consumption of hardware resources required for coding. [Brief explanation of the drawing]

[0024] [Figure 1] An example of a video / image coding system that can be applied to embodiments of this document is schematically shown. [Figure 2] It is a drawing schematically explaining the configuration of a video / image encoding device that can be applied to embodiments of this document. [Figure 3] It is a drawing schematically explaining the configuration of a video / image decoding device that can be applied to embodiments of this document. [Figure 4] A hierarchical structure for a coded image / video is exemplarily shown. [Figure 5] It is a sequence diagram for explaining an encoding method based on filtering in an encoding device. [Figure 6] It is a sequence diagram for explaining a decoding method based on filtering in a decoding device. [Figure 7] An example of a video / image encoding method and related components according to embodiments (etc.) of this document is schematically shown. [Figure 8] An example of a video / image encoding method and related components according to embodiments (etc.) of this document is schematically shown. [Figure 9] An example of an image / video decoding method and related components according to embodiments (etc.) of this document is schematically shown. [Figure 10] An example of an image / video decoding method and related components according to embodiments (etc.) of this document is schematically shown. [Figure 11] An example of a content streaming system to which embodiments etc. disclosed in this document can be applied is shown.

Mode for Carrying Out the Invention

[0025] This document may be modified in various ways and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail. However, this is not intended to limit this document to any particular embodiment. Terms used herein are used solely to describe specific embodiments and are not intended to limit the technical ideas of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as “includes” or “has” herein are intended to specify the existence of features, figures, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood not to preemptively exclude the existence or possibility of adding one or more other features, figures, steps, actions, components, parts, or combinations thereof.

[0026] On the other hand, each configuration shown in the diagrams described in this document is illustrated independently for the purpose of explaining its distinct characteristic functions, and does not mean that each configuration is implemented with separate hardware or separate software. For example, two or more of the configurations may be combined to form a single configuration, and one configuration may be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are also included in the scope of the rights of this document, as long as they do not deviate from the essence of this document.

[0027] Preferred embodiments of this document will be described in more detail below with reference to the attached drawings. Hereafter, the same reference numerals will be used for identical components in the drawings, and redundant descriptions of identical components will be omitted.

[0028] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document relate to the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), next-generation video / image coding standards after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), EVC (essential video coding) standard, AVS2 standard, etc.).

[0029] This document presents various embodiments relating to video / image coding, and unless otherwise noted, these embodiments may be implemented in combination with each other.

[0030] In this document, "video" can mean a collection of images over time. "Picture" generally refers to a single image representing a specific time period, while "slice" or "tile" is a unit that constitutes part of a picture in coding. A slice or tile may contain one or more CTUs (coding tree units). A single picture may consist of one or more slices or tiles. A single picture may consist of one or more tile groups. A tile group may contain one or more tiles.

[0031] A pixel or pel can refer to the smallest unit that makes up a picture (or image). Alternatively, the term "sample" may be used as a counterpart to pixel. A sample generally refers to a pixel or a pixel value, and may refer only to the pixel / pixel value of the luma component, or only to the pixel / pixel value of the chroma component. Alternatively, a sample can refer to a pixel value in the spatial domain, and if such a pixel value is converted to the frequency domain, it can also refer to the conversion coefficient in the frequency domain.

[0032] A unit represents a basic unit of image processing. A unit contains at least one of the following: a specific region of a picture and information about that region. One unit contains one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block contains a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.

[0033] In this document, " / " and "," are interpreted as "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Additionally, "A / B / C" means "at least one of A, B and / or C". Similarly, "A, B, C" also means "at least one of A, B and / or C".

[0034] Additionally, in this document, "or" is interpreted as "and / or". For example, "A or B" could mean 1) only "A", 2) only "B", or 3) "A and B". In other words, "or" in this document can mean "additionally or alternatively".

[0035] In this specification, "at least one of A and B" may mean "A only," "B only," or "both A and B." Furthermore, in this specification, the expressions "at least one of A or B" and "at least one of A and / or B" may be interpreted similarly to "at least one of A and B."

[0036] Furthermore, in this specification, "at least one of A, B and C" may mean "A only," "B only," "C only," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and C" may mean "at least one of A, B and C."

[0037] Furthermore, parentheses used in this specification may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be proposed as an example of "prediction." Also, when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be proposed as an example of "prediction."

[0038] Technical features described individually within a single drawing in this specification may be implemented individually or simultaneously.

[0039] Figure 1 schematically shows an example of a video / image coding system to which this document can be applied.

[0040] As shown in Figure 1, a video / image coding system may comprise a source device and a receiving device. The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or network.

[0041] The source device may comprise a video source, an encoding device, and a transmitter. The receiving device may comprise a receiver, a decoding device, and a renderer. The encoding device may be called a video / image encoding device, and the decoding device may be called a video / image decoding device. The transmitter may be provided in the encoding device. The receiver may be provided in the decoding device. The renderer may comprise a display unit, which may consist of a separate device or external component.

[0042] A video source can acquire video / images through processes such as video / image capture, synthesis, or generation. A video source may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, or a video / image archive containing previously captured video / images. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images may be generated via a computer, in which case the video / image capture process may be replaced by the process of generating the associated data.

[0043] An encoding device can encode input video / images. For compression and coding efficiency, the encoding device can perform a series of steps including prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output in bitstream format.

[0044] The transmitting unit can transmit encoded video / image information or data output in bitstream format to the receiving unit of a receiving device via a digital storage medium or network in file or streaming format. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. The transmitting unit may include elements for generating media files via a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.

[0045] A decoding device can decode video / images by performing a series of steps, such as inverse quantization, inverse transformation, and prediction, corresponding to the operation of an encoding device.

[0046] The renderer can render the decoded video / image. The rendered video / image can be displayed via the display unit.

[0047] Figure 2 is a schematic diagram illustrating the configuration of a video / image encoding device to which this document applies. Hereinafter, the term "video encoding device" may include an image encoding device.

[0048] As shown in Figure 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be called a reconstructor or a reconstructed block generator. The aforementioned image segmentation unit 210, prediction unit 220, residual processing unit 230, entropy encoding unit 240, addition unit 250, and filtering unit 260 can be configured by one or more hardware components (e.g., an encoder chipset or processor) depending on the embodiment. The memory 270 may also include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.

[0049] The image splitting unit 210 can split an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units can be recursively split from a coding tree unit (CTU) or the largest coding unit (LCU) using a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be split into multiple coding units of deeper depth based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, followed by the binary tree structure and / or the ternary structure. Alternatively, the binary tree structure may be applied first. The coding procedure according to this disclosure may be performed based on the final coding unit that is not further split. In this case, based on coding efficiency due to image characteristics, the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into lower-depth coding units so that the optimally sized coding unit is used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further comprise a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit can each be separated or partitioned from the final coding unit described above.The prediction unit may be a unit of sample prediction, and the conversion unit may be a unit for deriving conversion coefficients and / or a unit for deriving a residual signal from conversion coefficients.

[0050] The term "unit" can sometimes be used interchangeably with terms such as "block" or "area." Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and may represent only the luminance (luma) component pixel / pixel value, or only the chroma component pixel / pixel value. A sample can be used to refer to a single picture (or image) as a pixel or pel.

[0051] The subtraction unit 231 can generate a residual signal (residual block, residual sample, or residual sample array) by subtracting the predicted signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input image signal (original block, original sample, or original sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can make predictions for the block to be processed (hereinafter referred to as the current block) and generate a predicted block that includes a predicted sample for the current block. The prediction unit 220 can determine whether intra-prediction or inter-prediction is applied on a current block or CU basis. As will be described later in the explanation of each prediction mode, the prediction unit can generate various information related to prediction, such as prediction mode information, and transmit it to the entropy encoding unit 240. The information related to prediction can be encoded by the entropy encoding unit 240 and output in bitstream form.

[0052] The intra-prediction unit 222 can predict the current block by referring to a sample in the current picture. The referenced sample may be located adjacent to the current block or at a distance, depending on the prediction mode. The prediction mode in intra-prediction may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is illustrative, and more or fewer directional prediction modes may be used depending on the settings. The intra-prediction unit 222 may also determine the prediction mode to apply to the current block using the prediction modes applied to adjacent blocks.

[0053] The interprediction unit 221 can derive a predicted block relative to the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, col CU, etc., and the reference picture containing the temporal neighboring block may be called a collocated picture (colPic). For example, the inter-prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter-prediction can be performed based on various prediction modes; for example, in skip mode and merge mode, the inter-prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted.In motion vector prediction (MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0054] The prediction unit 220 can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra-prediction (CIIP). The prediction unit can also perform intra-block copy (IBC) for predictions on blocks. The intra-block copy can be used for content image / video coding in games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can use at least one of the inter-prediction techniques described in this document.

[0055] The prediction signals generated via the interpretation unit 221 and / or intrapretation unit 222 can be used to generate a reconstructed signal or a residual signal. The transformation unit 232 can apply transformation techniques to the residual signal to generate transformation coefficients. For example, transformation techniques may include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT refers to a transformation obtained from a graph when relational information between pixels is represented by this graph. CNT refers to a transformation obtained by generating a prediction signal using all previously reconstructed pixels and obtaining a transformation based on it. The transformation process may also be applied to pixel blocks of the same size and square shape, or to non-square blocks of variable size.

[0056] The quantization unit 233 quantizes the conversion coefficients and transmits them to the entropy encoding unit 240, which can encode the quantized signal (information about the quantized conversion coefficients) and output it as a bitstream. The information about the quantized conversion coefficients can be called residual information. The quantization unit 233 can rearrange the block-form quantized conversion coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized conversion coefficients based on the one-dimensional vector form of the quantized conversion coefficients. The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, CAVLC (context-adaptive variable length coding), and CABAC (context-adaptive binary arithmetic coding). In addition to the quantized conversion coefficients, the entropy encoding unit 240 can also encode information necessary for video / image restoration (e.g., the values ​​of syntax elements) together with or separately from the quantized conversion coefficients. Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form in network abstraction layer (NAL) units. The video / image information may further include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). The video / image information may also further include general constraint information. In this document, the signaling / transmitted information and / or syntax elements described later may be encoded via the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium.Here, the network may include broadcasting networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be transmitted by a transmitting unit (not shown) and / or stored by a storage unit (not shown) which are configured as internal / external elements of the encoding device 200, or the transmitting unit may be included in the entropy encoding unit 240.

[0057] The quantized conversion coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transformation to the quantized conversion coefficients via the inverse quantization unit 234 and the inverse transformation unit 235. The adder 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the prediction unit 220. If there is no residual for the block to be processed, as in the case of skip mode, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra-prediction of the next block to be processed in the current picture, and can also be used for inter-prediction of the next picture after filtering, as described later.

[0058] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.

[0059] The filtering unit 260 can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and the modified restored picture can be stored in the memory 270, specifically in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, and bilateral filter. The filtering unit 260 can generate various filtering-related information and transmit it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The filtering-related information can be encoded by the entropy encoding unit 240 and output in bitstream form.

[0060] The corrected restored picture sent to memory 270 can be used as a reference picture in the interpretation unit 221. When interpretation is applied via this, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve encoding efficiency.

[0061] The DPB in memory 270 can store the corrected restored picture for use as a reference picture in the inter-prediction unit 221. Memory 270 can store motion information of blocks from which motion information in the current picture has been derived (or encoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 221 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 270 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 222.

[0062] Figure 3 is a schematic diagram illustrating the configuration of a video / image decoding device to which this document can be applied.

[0063] As shown in Figure 3, the decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. The aforementioned entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 can be configured by a single hardware component (e.g., a decoder chipset or processor) depending on the embodiment. The memory 360 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The aforementioned hardware component may also further include memory 360 as an internal / external component.

[0064] When a bitstream containing video / image information is input, the decoding device 300 can reconstruct the image corresponding to the process by which the video / image information was processed in the encoding device shown in Figure 3. For example, the decoding device 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, a coding unit, which can be divided from a coding tree unit or a maximum coding unit according to a quad-tree structure, a binary tree structure, and / or a terminally tree structure. One or more conversion units can be derived from the coding unit. The reconstructed image signal decoded and output via the decoding device 300 can then be reproduced via a playback device.

[0065] The decoding device 300 can receive the signal output from the encoding device shown in Figure 3 in bitstream form, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information necessary for image restoration (or picture restoration) (e.g., video / image information). The video / image information may further include information about various parameter sets, such as the adaptation parameter set (APS), picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS). The video / image information may also further include general constraint information. The decoding device can further decode the picture based on the parameter set information and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values ​​of syntax elements necessary for image reconstruction and the quantized values ​​of conversion coefficients related to the residual. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded and the decoded information of adjacent and decoded blocks or symbol / bin information decoded in a previous step, predicts the probability of bin occurrence based on the determined context model, performs arithmetic decoding of the bins, and generates symbols corresponding to the values ​​of each syntax element. At this time, after determining the context model, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin.Of the information decoded by the entropy decoding unit 310, information related to prediction is provided to the prediction unit 330, and residual information that has been entropy decoded by the entropy decoding unit 310, i.e., quantized conversion coefficients and related parameter information, can be input to the inverse quantization unit 321. In addition, of the information decoded by the entropy decoding unit 310, information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoding unit 310. On the other hand, the decoding device relating to this document can be called a video / image / picture decoding device, and the decoding device can also be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoding unit 310, and the sample decoder may include at least one of the inverse quantization unit 321, inverse transformation unit 322, prediction unit 330, addition unit 340, filtering unit 350, and memory 360.

[0066] The inverse quantization unit 321 can inverse quantize the quantized transformation coefficients and output the transformation coefficients. The inverse quantization unit 321 can rearrange the quantized transformation coefficients in a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain the transformation coefficients.

[0067] In the inverse conversion unit 322, the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).

[0068] The prediction unit can make predictions for the current block and generate a predicted block containing prediction samples for the current block. Based on the prediction information output from the entropy decoding unit 310, the prediction unit can determine whether intra-prediction or inter-prediction is applied to the current block and can determine a specific intra / inter-prediction mode.

[0069] The prediction unit can generate prediction signals based on various prediction methods described later. For example, the prediction unit can apply intra-prediction or inter-prediction for a single block, and can also apply intra-prediction and inter-prediction simultaneously. This can be called combined inter and intra prediction (CIIP). The prediction unit can also perform intra-block copying (IBC) for predictions on blocks. This intra-block copying can be used for content image / video coding in games, for example, as in SCC (screen content coding). IBC basically performs prediction within the current picture, but can be done similarly to inter-prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter-prediction techniques described in this document. Palette mode can be considered an example of intra-coding or intra-prediction.

[0070] The intra-prediction unit 331 can predict the current block by referring to a sample in the current picture. The referenced sample can be located adjacent to or far from the current block depending on the prediction mode. In intra-prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra-prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction modes applied to adjacent blocks.

[0071] The interprediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) identified by motion vectors on the reference picture. In this case, in order to reduce the amount of motion information transmitted in interprediction mode, motion information can be predicted in blocks, subblocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information may include motion vectors and reference picture indices. The motion information may further include interprediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of interprediction, adjacent blocks may include spatially adjacent blocks that exist in the current picture and temporally adjacent blocks that exist in the reference picture. For example, the interprediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Interprediction can be performed based on various prediction modes, and the prediction information may include information indicating the mode of interprediction for the current block.

[0072] The summing unit 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the prediction unit. If there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as the restored block.

[0073] The addition unit 340 may be called the restoration unit or restoration block generation unit. The generated restoration signal can be used for intra-prediction of the next block to be processed in the current picture, and can be output after filtering as described later, or it can be used for intra-prediction of the next picture.

[0074] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture decoding process.

[0075] The filtering unit 350 can apply filtering to the restored signal to improve subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can transmit the modified restored picture to the memory 360, specifically to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, and bilateral filter.

[0076] The (modified) restored picture stored in the DPB of memory 360 can be used as a reference picture by the inter-prediction unit 332. Memory 360 can store motion information of blocks from which motion information in the current picture has been derived (or decoded) and / or motion information of blocks in the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit 332 for use as motion information of spatially adjacent blocks or motion information of temporally adjacent blocks. Memory 360 can store restored samples of restored blocks in the current picture and transmit them to the intra-prediction unit 331.

[0077] In this specification, embodiments described in relation to the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, and filtering unit 350 of the decoding device 300 can be applied identically to or corresponding to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, and filtering unit 260 of the encoding device 200, respectively.

[0078] As mentioned above, prediction is performed to improve compression efficiency when performing video coding. Through this, a predicted block containing predicted samples for the current block, which is the block to be coded, can be generated. Here, the predicted block contains predicted samples in the spatial domain (or pixel domain). The predicted block is derived in both the encoding and decoding devices, and the encoding device can improve image coding efficiency by signaling the decoding device with information (residual information) about the residual between the original block and the predicted block, which is not the original sample value of the original block itself. The decoding device can derive a residual block containing residual samples based on the residual information, and can generate a restored block containing restored samples by combining the residual block and the predicted block, and can generate a restored picture containing the restored block.

[0079] The residual information can be generated through transformation and quantization procedures. For example, an encoding device can signal the relevant residual information (via a bitstream) to a decoding device by deriving a residual block between the original block and the predicted block, performing a transformation procedure on the residual samples (residual sample array) contained in the residual block to derive transformation coefficients, and performing a quantization procedure on the transformation coefficients to derive quantized transformation coefficients. Here, the residual information may include information such as the value information, position information, transformation technique, transformation kernel, and quantization parameters of the quantized transformation coefficients. The decoding device can derive a residual sample (or residual block) by performing an inverse quantization / inverse transformation procedure based on the residual information. The decoding device can generate a reconstructed picture based on the predicted block and the residual block. The encoding device can also derive a residual block by inverse quantization / inverse transformation of the quantized transformation coefficients for reference for subsequent interpretation of the picture, and generate a reconstructed picture based on this.

[0080] In this document, at least one of quantization / inverse quantization and / or transformation / inverse transformation may be omitted. If quantization / inverse quantization is omitted, the quantized transformation coefficients may be called transformation coefficients. If transformation / inverse transformation is omitted, the transformation coefficients may also be called coefficients or residual coefficients, or for consistency of expression, they may still be called transformation coefficients.

[0081] In this document, quantized transformation coefficients and transformation coefficients may be referred to as transformation coefficients and scaled transformation coefficients, respectively. In this case, residual information may include information about the transformation coefficients, which may be signaled via residual coding syntax. Transformation coefficients may be derived based on the residual information (or information about the transformation coefficients), and scaled transformation coefficients may be derived via an inverse transformation (scaling) of the transformation coefficients. Residual samples may be derived based on an inverse transformation (transformation) of the scaled transformation coefficients. This may be applied / expressed similarly in other parts of this document.

[0082] The prediction unit of the encoding / decoding device can perform interpretation on a block-by-block basis to derive predicted samples. Interpretation can indicate predictions derived in a manner dependent on data elements (e.g., sample values ​​or motion information) of pictures other than the current picture. When interpretation is applied to the current block, a predicted block (predicted sample array) for the current block can be derived based on the reference block (reference sample array) identified by the motion vector on the reference picture pointed to by the index of the reference picture. In this case, in order to reduce the amount of motion information transmitted in interpretation mode, the motion information of the current block can be predicted on a block, subblock, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information may include motion vectors and the index of the reference picture. The motion information may further include information on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). When interpretation is applied, neighboring blocks may include spatial neighboring blocks that exist in the current picture and temporal neighboring blocks that exist in the reference picture. The reference picture containing the aforementioned reference block and the reference picture containing the aforementioned temporally adjacent block may be the same or different. The temporally adjacent block may be referred to by names such as collocated reference block, colCU, etc., and the reference picture containing the temporally adjacent block may be referred to as collocated picture (colPic). For example, a candidate list of motion information can be constructed based on the adjacent blocks of the current block, and flags or index information can be signaled to indicate which candidate is selected (used) in order to derive the motion vector and / or index of the reference picture of the current block. Interpretation is performed based on various prediction modes; for example, in skip mode and merge mode, the motion information of the current block may be the same as the motion information of the selected adjacent block.In skip mode, unlike merge mode, the residual signal may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the selected adjacent block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0083] The motion information may include L0 motion information and / or L1 motion information depending on the interpretation type (L0 prediction, L1 prediction, Bi prediction, etc.). A motion vector in the L0 direction may be called an L0 motion vector or MVL0, and a motion vector in the L1 direction may be called an L1 motion vector or MVL1. A prediction based on an L0 motion vector may be called an L0 prediction, a prediction based on an L1 motion vector may be called an L1 prediction, and a prediction based on both the L0 motion vector and the L1 motion vector may be called a bi (Bi) prediction. Here, an L0 motion vector may represent a motion vector associated with a reference picture list L0 (L0), and an L1 motion vector may represent a motion vector associated with a reference picture list L1 (L1). The reference picture list L0 may include pictures earlier in the output order than the current picture, and the reference picture list L1 may include pictures later in the output order than the current picture. The aforementioned earlier picture may be called a forward (reference) picture, and the aforementioned later picture may be called a reverse (reference) picture. The reference picture list L0 may include further reference pictures that are later in the output order than the current picture. In this case, the earlier picture may be indexed first in the reference picture list L0, and the later picture may be indexed afterward. The reference picture list L1 may include further reference pictures that are earlier in the output order than the current picture. In this case, the later picture may be indexed first in the reference picture list L1, and the earlier picture may be indexed afterward. Here, the output order may correspond to the POC (picture order count) order.

[0084] Figure 4 illustrates the hierarchical structure for coded images / videos.

[0085] As shown in Figure 4, coded images / videos are divided into the VCL (video coding layer), which handles the decoding process and the images / videos themselves; a lower-level system that transmits and stores the coded information; and the NAL (network abstraction layer), which exists between the VCL and the lower-level system and is responsible for network adaptation functions.

[0086] VCL can generate VCL data containing compressed image data (slice data), or generate parameter sets containing information such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS), or SEI (Supplemental Enhancement Information) messages that are additionally necessary during the image decoding process.

[0087] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the RBSP (Raw Byte Sequence Payload) generated by VCL. In this case, the RBSP refers to the slice data, parameter set, SEI message, etc., generated by VCL. The NAL unit header can include NAL unit type information, which is identified by the RBSP data contained in the NAL unit.

[0088] As shown in the above diagram, NAL units can be divided into VCL NAL units and Non-VCL NAL units by the RBSP generated in VCL. VCL NAL units can mean NAL units that contain information about the image (slice data), and Non-VCL NAL units can mean NAL units that contain information necessary for decoding the image (parameter set or SEI message).

[0089] The aforementioned VCL NAL units and Non-VCL NAL units can be transmitted over a network with header information added according to the data standards of the lower-level system. For example, NAL units can be transformed into data formats of predetermined standards such as H.266 / VVC file format, RTP (Real-time Transport Protocol), and TS (Transport Stream) and transmitted over various networks.

[0090] As mentioned above, the NAL unit type can be identified by the RBSP data structure contained within the NAL unit, and information about such NAL unit types can be stored in the NAL unit header and signaled.

[0091] For example, NAL units can be broadly classified into VCL NAL unit types and Non-VCL NAL unit types depending on whether or not they contain information (slice data) about the image. VCL NAL unit types can be further classified by the nature and type of picture they contain, while Non-VCL NAL unit types can be further classified by the type of parameter set.

[0092] The following is an example of a NAL unit type identified by the type of parameter set included in the Non-VCL NAL unit type.

[0093] -APS (Adaptation Parameter Set) NAL unit: Type for NAL units that include APS

[0094] -DPS (Decoding Parameter Set) NAL unit: Type for NAL units including DPS

[0095] -VPS (Video Parameter Set) NAL unit: Type for NAL unit including VPS

[0096] -SPS (Sequence Parameter Set) NAL unit: Type for NAL units that include SPS

[0097] -PPS (Picture Parameter Set) NAL unit: Type for NAL units that include PPS

[0098] -PH (Picture header) NAL unit: Type for NAL units that include PH

[0099] The aforementioned NAL unit type has syntax information for the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be identified by the nal_unit_type value.

[0100] On the other hand, as mentioned above, a single picture can contain multiple slices, and a single slice can contain a slice header and slice data. In this case, a picture header may be added to each of the multiple slices (slice headers and slice data sets) within a single picture. The picture header (picture header syntax) may contain information / parameters that can be commonly applied to the picture. In this document, slices may be mixed with or replaced by tile groups. Also in this document, slice headers may be mixed with or replaced by type group headers.

[0101] The slice header (slice header syntax, slice header information) may include information / parameters that can be commonly applied to the slice. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. The VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. The DPS (DPS syntax) may include information / parameters that can be commonly applied to video in general. The DPS may include information / parameters related to the concatenation of CVS (coded video sequence). In this document, High-level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.

[0102] In this document, the image / video information encoded from the encoding device to the decoding device and signaled in bitstream form may include not only partitioning-related information, intra / inter prediction information, residual information, and in-loop filtering information within the picture, but also information contained in the slice header, the picture header, the APS, the PPS, the SPS, the VPS, and / or the DPS. Furthermore, the image / video information may further include information from the NAL unit header.

[0103] On the other hand, to compensate for differences between the original image and the reconstructed image due to errors that occur during the compression encoding process, such as quantization, an in-loop filtering procedure can be performed on the reconstructed sample or reconstructed picture, as described above. As described above, in-loop filtering can be performed in the filter section of the encoding device and the filter section of the decoding device, and a deblocking filter, SAO, and / or adaptive loop filter (ALF) can be applied. For example, the ALF procedure can be performed after the deblocking filtering procedure and / or SAO procedure are completed. However, even in this case, the deblocking filtering procedure and / or SAO procedure may be omitted.

[0104] The following provides a detailed explanation of picture restoration and filtering. In image / video coding, a restored block can be generated based on intra-prediction / inter-prediction for each block, and a restored picture containing the restored block can be generated. If the current picture / slice is an I-picture / slice, the blocks contained in the current picture / slice can be restored based solely on intra-prediction. On the other hand, if the current picture / slice is a P or B-picture / slice, the blocks contained in the current picture / slice can be restored based on either intra-prediction or inter-prediction. In this case, intra-prediction may be applied to some blocks within the current picture / slice, while inter-prediction may be applied to the remaining blocks.

[0105] Intra prediction can represent a prediction that generates prediction samples for the current block based on reference samples within the picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, adjacent reference samples to be used for intra prediction of the current block can be derived. The adjacent reference samples of the current block may include samples adjacent to the left boundary of the current block of size nW × nH and a total of 2 × nH samples adjacent to the bottom left, samples adjacent to the top boundary of the current block and a total of 2 × nW samples adjacent to the top right, and one sample adjacent to the top left of the current block. Alternatively, the adjacent reference samples of the current block may include multiple columns of upper adjacent samples and multiple rows of left adjacent samples. Furthermore, the adjacent reference samples of the current block may also include a total of nH samples adjacent to the right boundary of the current block, which is of size nW × nH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right side of the current block.

[0106] However, some of the adjacent reference samples in the current block may not yet be decoded or available. In this case, the decoder can construct adjacent reference samples to be used for prediction by substituting the unavailable samples as available samples, or by constructing adjacent reference samples to be used for prediction through interpolation of available samples.

[0107] If neighboring reference samples are derived, (i) predicted samples can be derived based on the average or interpolation of neighboring reference samples in the current block, or (ii) predicted samples can be derived based on reference samples in the current block that are located in a specific (predicted) direction relative to the predicted sample. Case (i) is called the non-directional mode or non-angular mode, and case (ii) is called the directional mode or angular mode. Alternatively, the predicted sample can be generated by interpolation between the first neighboring sample and a second neighboring sample located in the opposite direction to the prediction direction of the current block's intra-prediction mode, based on the predicted sample of the current block. In this case, it can be called linear interpolation intra-prediction (LIP). Alternatively, chroma predicted samples can be generated based on chroma samples using a linear model. In this case, it can be called the LM mode. Alternatively, a temporary predicted sample for the current block can be derived based on filtered adjacent reference samples, and the predicted sample for the current block can be derived by performing a weighted sum of the temporary predicted sample and at least one reference sample derived by the intra-prediction mode from the existing adjacent reference samples, i.e., unfiltered adjacent reference samples. In the above case, it can be called PDPC (Position dependent intra prediction). In addition, intra-predictive coding can be performed by selecting the reference sample line with the highest prediction accuracy from the adjacent multiple reference sample lines of the current block, deriving the predicted sample using the reference sample located in the prediction direction on that line, and instructing (signaling) the decoding device with the reference sample line used at this time.In the aforementioned cases, this can be called multi-reference line (MRL) intra prediction or MRL-based intra prediction. Furthermore, the current block can be divided into vertical or horizontal subpartitions, and intra prediction can be performed based on the same intra prediction mode, with adjacent reference samples derived and available for use on a subpartition basis. That is, in this case, the intra prediction mode for the current block is also applied to the subpartition, and by deriving and using adjacent reference samples on a subpartition basis, intra prediction performance can be improved in some cases. Such prediction methods can be called intra subpartitions (ISP) or ISP-based intra prediction. The aforementioned intra prediction methods can be distinguished from the intra prediction modes in Table of Contents 1 and 2 and referred to as intra prediction types. These intra prediction types can be referred to by various terms, such as intra prediction techniques or additional intra prediction modes. For example, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the aforementioned LIP, PDPC, MRL, and ISP. A general intra-prediction method that excludes specific intra-prediction types such as LIP, PDPC, MRL, and ISP can be called a normal intra-prediction type. The normal intra-prediction type can be generally applied when the aforementioned specific intra-prediction types are not applicable, and predictions can be performed based on the intra-prediction modes described above. Meanwhile, post-processing filtering can be performed on the derived prediction samples as needed.

[0108] Specifically, the intra-prediction procedure may include an intra-prediction mode / type determination step, an adjacent reference sample derivation step, and an intra-prediction mode / type-based predictive sample derivation step. Additionally, a post-filtering step may be performed on the derived predictive samples as needed.

[0109] A modified restored picture is generated by the in-loop filtering procedure, and the decoder outputs the modified restored picture as a decoded picture. This modified restored picture is also stored in the decoded picture buffer or memory of the encoding / decoding device and can be used as a reference picture in the interpretation procedure during subsequent encoding / decoding of the picture. The in-loop filtering procedure includes, as described above, a deblocking filtering procedure, an SAO (sample adaptive offset) procedure, and / or an ALF (adaptive loop filter) procedure. In this case, one or part of the deblocking filtering procedure, the SAO (sample adaptive offset) procedure, the ALF (adaptive loop filter) procedure, and the bi-lateral filter procedure may be applied sequentially, or all of them may be applied sequentially. For example, the SAO procedure may be performed after the deblocking filtering procedure is applied to the restored picture. Or, for example, the ALF procedure may be performed after the deblocking filtering procedure is applied to the restored picture. This is also done in the encoding device.

[0110] Deblocking filtering is a filtering technique that removes distortion occurring at the boundaries between blocks in a restored picture. The deblocking filtering procedure can, for example, involve deriving a target boundary in the restored picture, determining a boundary strength (bS) for the target boundary, and performing deblocking filtering on the target boundary based on the bS. The bS can be determined based on the prediction modes of two adjacent blocks, the difference in motion vectors, whether the reference picture is identical, and whether a non-zero effectiveness coefficient exists.

[0111] SAO is a method for compensating for the offset difference between a restored picture and the original picture on a sample-by-sample basis, and can be applied based on types such as Band Offset and Edge Offset. According to SAO, each SAO type can classify samples into different categories, and an offset value can be added to each sample based on the category. Filtering information for SAO can include information on whether SAO is applicable, SAO type information, SAO offset value information, etc. SAO can also be applied to the restored picture after the deblocking filtering has been applied.

[0112] ALF (Adaptive Loop Filter) is a technique that filters a restored picture on a sample-by-sample basis based on filter coefficients determined by the filter shape. The encoding device can determine whether ALF is applicable, the ALF shape, and / or ALF filtering coefficients by comparing the restored picture with the original picture, and can signal this to the decoding device. That is, filtering information for ALF can include information on whether ALF is applicable, ALF filter shape information, ALF filtering coefficient information, etc. ALF can also be applied to the restored picture after the deblocking filtering has been applied.

[0113] Figure 5 is a sequence diagram illustrating a filtering-based encoding method in an encoding device. The method in Figure 5 may include steps S500 to S530.

[0114] In step S500, the encoding device can generate a restored picture. Step S500 can be performed based on the restored picture (or restored sample) generation procedure described above.

[0115] In step S510, the encoding device can determine whether or not in-loop filtering is applied (across the virtual boundary) based on in-loop filtering-related information. Here, in-loop filtering may include at least one of the deblocking filtering, SAO, or ALF described above.

[0116] In step S520, the encoding device can generate a modified restored picture (modified restored sample) based on the decision in step S510. Here, the modified restored picture (modified restored sample) may be a filtered restored picture (filtered restored sample).

[0117] In step S530, the encoding device can encode image / video information including in-loop filtering-related information based on the in-loop filtering procedure.

[0118] Figure 6 is a sequence diagram illustrating a filtering-based decoding method in a decoding device. The method in Figure 6 may include steps S600 to S630.

[0119] In step S600, the decoding device can obtain image / video information, including in-loop filtering-related information, from the bitstream. Here, the bitstream can be based on encoded image / video information transmitted from the encoding device.

[0120] In step S610, the decoding device can generate a restored picture. Step S610 can be performed based on the restored picture (or restored sample) generation procedure described above.

[0121] In step S620, the decoding device can determine whether or not in-loop filtering is applied (across the virtual boundary) based on in-loop filtering-related information. Here, in-loop filtering may include at least one of the deblocking filtering, SAO, or ALF described above.

[0122] In step S630, the decoding device can generate a modified restored picture (modified restored sample) based on the decision in step S620. Here, the modified restored picture (modified restored sample) may be a filtered restored picture (filtered restored sample).

[0123] As described above, an in-loop filtering procedure can be applied to the restored picture. In this case, a virtual boundary can be defined to further enhance the subjective / objective visual quality of the restored picture, and the in-loop filtering procedure can be applied across the virtual boundary. The virtual boundary includes, for example, discontinuous edges such as 360-degree images, VR images, or PIPs (picture in picture). For example, the virtual boundary may exist at a predetermined, promised location, and its presence and / or location may be signaled. As an example, the virtual boundary may be located above the fourth sample line from the top of the CTU row (specifically, for example, above the fourth sample line from the top of the CTU row). As another example, information regarding the presence and / or location of the virtual boundary may also be signaled via an HLS. The HLS includes, as described above, SPS, PPS, picture headers, slice headers, etc.

[0124] The following describes the high-level syntax signaling and semantics related to the embodiments described in this document.

[0125] One embodiment of this document includes a method for controlling a loop filter. This method for controlling a loop filter can be applied to a restored picture. The in-loop filter (loop filter) can be used for decoding an encoded bitstream. The loop filter includes the deblocking, SAO, and ALF mentioned above. The SPS includes flags associated with each of the deblocking, SAO, and ALF. The flags indicate whether each tool is available for coding a CLVS (coded layer video sequence) or CVS (coded video sequence) that references the SPS.

[0126] In one example, if a loop filter is available for coding pictures within a CVS, the application of the loop filter is controlled so as not to cross certain boundaries. For example, the loop filter may be controlled so as not to cross subpicture boundaries, tile boundaries, slice boundaries, and / or virtual boundaries.

[0127] In-loop filtering-related information includes information, syntax, syntax elements, and / or semantics described in this document (or embodiments contained herein). In-loop filtering-related information includes information regarding whether the in-loop filtering procedure (in whole or in part) is applicable across a specific boundary (e.g., a virtual boundary, subpicture boundary, slice boundary, and / or tile boundary). Image information contained in a bitstream includes high-level syntax (HLS), which includes the in-loop filtering-related information. Based on the decision regarding whether the in-loop filtering procedure is applied across a specific boundary, modified (or filtered) reconstructed samples (reconstructed pictures) are generated. In one example, if the in-loop filtering procedure is disabled for all blocks / boundaries, the modified reconstructed samples may be identical to the reconstructed samples. In another example, the modified reconstructed samples include modified reconstructed samples derived based on in-loop filtering, however, in this case, some of the reconstructed samples (e.g., reconstructed samples across a virtual boundary) may not be in-loop filtered based on the aforementioned decision. For example, a recovery sample crossing a specific boundary (including at least one of a virtual boundary, subpicture boundary, slice boundary, and / or tile boundary for which in-loop filtering is enabled) may be subjected to in-loop filtering, while a recovery sample crossing another boundary (including at least one of a virtual boundary, subpicture boundary, and / or tile boundary for which in-loop filtering is disabled) may not be subjected to in-loop filtering.

[0128] In one example, in relation to whether or not an in-loop filtering procedure is performed across a virtual boundary, in-loop filtering-related information includes an SPS virtual boundary existence flag, a picture header virtual boundary existence flag, information about the number of virtual boundaries, and information about the location of the virtual boundaries.

[0129] In the embodiments described herein, the information regarding the position of a virtual boundary includes information regarding the x-coordinate of a vertical virtual boundary and / or information regarding the y-coordinate of a horizontal virtual boundary. Specifically, the information regarding the position of a virtual boundary includes information regarding the x-coordinate of a vertical virtual boundary and / or information regarding the y-coordinate of a horizontal virtual boundary in units of luma samples. Furthermore, the information regarding the position of a virtual boundary includes information regarding the number of x-coordinates of vertical virtual boundaries (syntax elements) present in the SPS. Furthermore, the information regarding the position of a virtual boundary includes information regarding the number of y-coordinates of horizontal virtual boundaries (syntax elements) present in the SPS. Alternatively, the information regarding the position of a virtual boundary includes information regarding the number of x-coordinates of vertical virtual boundaries (syntax elements) present in the picture header. Furthermore, the information regarding the position of a virtual boundary includes information regarding the number of y-coordinates of horizontal virtual boundaries (syntax elements) present in the picture header.

[0130] The following table shows exemplary syntax and semantics of the SPS (sequence parameter set) according to this embodiment.

[0131] [Table 1]

[0132] [Table 2]

[0133] The following table shows exemplary syntax and semantics of the PPS (picture parameter set) according to this embodiment.

[0134] [Table 3]

[0135] [Table 4]

[0136] The following table shows the exemplary syntax and semantics of the picture header according to this embodiment.

[0137] [Table 5-1]

[0138] [Table 5-2]

[0139] [Table 6-1]

[0140] [Table 6-2]

[0141] The following table shows the exemplary syntax and semantics of a slice header according to this embodiment.

[0142] [Table 7]

[0143] [Table 8]

[0144] The following sections describe information related to subpictures, information about virtual boundaries that can be used in in-loop filtering, and their signaling.

[0145] If a picture contains multiple subpictures but none of the subpictures have a boundary that is treated like a picture boundary, the advantages of using subpictures cannot be realized. In one embodiment of this document, the image / video information for image coding includes information for treating subpictures like pictures, which is called a picture-treated flag (e.g., subpic_treated_as_pic_flag[i]).

[0146] To signal the layout of subpictures, a flag related to whether or not subpictures exist (e.g., subpic_present_flag) is signaled. This may also be called the subpicture existence flag. If the value of subpic_present_flag is 1, information about the number of subpictures that divide the picture (e.g., sps_num_subpics_minus1) is signaled. In one example, the number of subpictures that divide the picture may be the same as sps_num_subpics_minus1 + 1 (sps_num_subpics_minus1 plus 1). Possible values ​​for sps_num_subpics_minus1 include 0, which means that there is only one subpicture in the picture. If a picture contains only one subpicture, the signaling of subpicture-related information is considered a redundant step because the subpicture itself is a picture.

[0147] In existing embodiments, if a picture contains only one subpicture and subpicture signaling exists, the value of the picture handling flag (e.g., subpic_treated_as_pic_flag[i]) and / or the flag related to whether loop filtering is performed across the subpicture (e.g., loop_filter_across_enabled_flag) can be 0 or 1. Here, if the value of subpic_treated_as_pic_flag[i] is 0, a problem arises that contradicts the case where the subpicture boundary is the picture boundary. This requires an additional redundant step to make the decoder confirm that the picture boundary is the subpicture boundary.

[0148] When a picture is generated based on a procedure for merging two or more subpictures, all subpictures used in the merging procedure must be independently coded subpictures (subpictures with a picture handling flag (subpic_treated_as_pic_flag[i]) value of 1). This is because merging a subpicture that is not independently coded (hereinafter referred to as the "first subpicture") can cause problems after merging because blocks within the first subpicture may be coded by referencing reference blocks that exist outside the first subpicture.

[0149] Furthermore, if a picture is partitioned into subpictures, subpicture ID signaling may or may not exist. If subpicture ID signaling exists, it is present (included) in the SPS, PPS and / or picture header (PH). If subpicture ID signaling is not present in the SPS, this includes cases where a bitstream is generated as a result of the subpicture merging procedure. Therefore, when subpicture ID signaling is not included in the SPS, it is preferable that all subpictures be coded independently.

[0150] In image coding procedures that utilize virtual boundaries, information regarding the location of the virtual boundary can be signaled in the SPS or picture header. Signaling information regarding the location of the virtual boundary in the SPS means that there is no change in that location within the CLVS. However, if reference picture resampling (RPR) is enabled for the CLVS, the pictures within the CLVS will have different sizes. Here, reference picture resampling (also called adaptive resolution change, ARC) is performed for the normal coding operation of pictures with different resolutions (spatial resolutions). For example, reference picture resampling includes upsampling and downsampling. Reference picture resampling achieves high coding efficiency for the adaptation of bitrate and spatial resolution. Taking reference picture resampling into account, it is necessary to ensure that the location of the virtual boundary is always within a single picture.

[0151] In existing ALF procedures, k-order exponential Golomb coding with k=3 is used to signal the absolute values ​​of the Luma and Chroma ALF coefficients. However, k-order exponential Golomb coding is problematic because it induces considerable computational overhead and complexity.

[0152] The embodiments described in the following paragraphs offer solutions to the aforementioned problems. The embodiments may be applied independently, or at least two or more embodiments may be applied in combination.

[0153] In one embodiment of this document, when subpicture signaling exists and a picture has only one subpicture, the single subpicture is an independently coded subpicture. For example, when a picture has only one subpicture, the single subpicture is an independently coded subpicture, and the value of the picture handling flag (e.g., subpic_treated_as_pic_flag[i]) for the single subpicture is 1. This allows for the omission of redundant steps associated with the subpicture.

[0154] In one embodiment of this document, if subpicture signaling exists, the number of subpictures may be greater than one. In one example, if subpicture signaling exists (for example, the value of subpics_present_flag is 1), then the information regarding the number of subpictures (for example, sps_num_subpics_minus1) is greater than 0, and the number of subpictures may be sps_num_subpics_minus1+1 (sps_num_subpics_minus1 plus 1). In another example, the information regarding the number of subpictures is sps_num_subpics_minus2, and the number of subpictures may be sps_num_subpics_minus2+2 (sps_num_subpics_minus2 plus 2). In another example, the subpics_present_flag can be replaced by sps_num_subpics_minus1, which is information about the number of subpics, and therefore subpics signaling can exist if sps_num_subpics_minus1 is greater than 0.

[0155] In one embodiment of this document, when a picture is divided into subpictures, at least one of the subpictures may be an independently coded subpicture. Here, the value of the picture handling flag (e.g., subpic_treated_as_pic_flag[i]) for the independently coded subpicture is 1.

[0156] In one embodiment of this document, the subpictures of a picture based on a procedure for merging two or more subpictures may be independently coded subpictures.

[0157] In one embodiment of this document, if the subpicture ID (identification) signaling is located in a location other than the SPS (other syntax, other high-level syntax information), then all subpictures are independently coded subpictures, and the value of the picture handling flag (e.g., subpic_treated_as_pic_flag) for all subpictures may be 1. In one example, the subpicture ID signaling is located in the PPS, in which case all subpictures may be independently coded subpictures. In another example, the subpicture ID signaling is located in the picture header, in which case all subpictures may be independently coded subpictures.

[0158] In one embodiment of this document, if virtual boundary signaling for CLVS is present in the SPS and reference picture resampling is possible, then all horizontal virtual boundary positions may be within the minimum picture height of the picture referencing the SPS, and all vertical virtual boundary positions may be within the minimum picture width of the picture referencing the SPS.

[0159] In one embodiment of this document, when reference picture resampling (RPR) is enabled, virtual boundary signaling is included in the picture header. That is, when reference picture resampling is enabled, virtual boundary signaling may not be included in the SPS.

[0160] In one embodiment of this document, fixed-length coding (FLC) is used to signal ALF data, with a corresponding number of bits (or bit length). In one example, information about the ALF data includes information about the bit length of the absolute value of the ALF luma coefficient (e.g., alf_luma_coeff_abs_len_minus1) and / or information about the bit length of the absolute value of the ALF chroma coefficient (e.g., alf_chroma_coeff_abs_len_minus1). For example, information about the bit length of the absolute value of the ALF luma coefficient and / or information about the bit length of the absolute value of the ALF chroma coefficient can be ue(v) coded.

[0161] The following table shows an exemplary syntax of SPS according to this embodiment.

[0162] [Table 9]

[0163] The following table shows exemplary semantics for the syntax elements included in the aforementioned syntax.

[0164] [Table 10]

[0165] The following table shows an exemplary syntax of SPS according to this embodiment.

[0166] [Table 11]

[0167] The following table shows exemplary semantics for the syntax elements included in the aforementioned syntax.

[0168] [Table 12]

[0169] The following table shows an exemplary syntax of ALF data according to this embodiment.

[0170] [Table 13]

[0171] The following table shows exemplary semantics for the syntax elements included in the aforementioned syntax.

[0172] [Table 14]

[0173] According to the embodiments of this document described in conjunction with the aforementioned table, image coding based on subpictures and / or virtual boundaries improves the subjective and objective quality of images and reduces the consumption of hardware resources required for coding.

[0174] Figures 7 and 8 schematically show an example of a video / image encoding method and related components according to the embodiments (etc.) described in this document.

[0175] The method disclosed in Figure 7 can be performed by the encoding device disclosed in Figure 2 or Figure 8. Specifically, for example, steps S700 and S730 in Figure 7 can be performed by the prediction unit 220 of the encoding device in Figure 8, steps S710 and S720 in Figure 7 can be performed by the residual processing unit 230 of the encoding device in Figure 8, step S740 in Figure 7 can be performed by the filtering unit 260 of the encoding device in Figure 8, and step S750 in Figure 7 can be performed by the entropy encoding unit 240 of the encoding device in Figure 8. Although not shown in Figure 7, the prediction unit 220 of the encoding device in Figure 7 can derive prediction samples or prediction-related information, and the entropy encoding unit 240 of the encoding device can generate a bitstream from the residual information or prediction-related information. The method disclosed in Figure 7 may include the embodiments described above in this document.

[0176] As shown in Figure 7, the encoding device can derive at least one reference picture (S700). The encoding device can perform a prediction procedure based on the at least one reference picture. Specifically, the encoding device can generate prediction samples for the current block based on the prediction mode. In this case, various prediction methods disclosed in this document, such as interpretation or intrapretation, may be applied. The encoding device can generate prediction samples for the current block in the current picture based on the prediction procedure. For example, the encoding device can perform an interpretation procedure based on the at least one reference picture and then generate prediction samples based on the interpretation procedure.

[0177] The encoding device can generate / derive a residual sample (S710). The encoding device can derive a residual sample for the current block, which can be derived based on the original sample and the predicted sample of the current block. In one example, the encoding device can generate a residual sample based on the at least one reference picture in step S700. For example, the encoding device can generate a predicted sample for the current block based on the at least one reference picture, and then generate a residual sample based on the predicted sample.

[0178] The encoding device can derive conversion coefficients. The encoding device can derive conversion coefficients based on the conversion procedure for the residual sample. For example, the conversion procedure may include at least one of DCT, DST, GBT, or CNT.

[0179] The encoding device can derive quantized transformation coefficients. The encoding device can derive quantized transformation coefficients based on a quantization procedure for the transformation coefficients. The quantized transformation coefficients may have a one-dimensional vector form based on the coefficient scan order.

[0180] The encoding device can generate residual information (S720). The encoding device can generate residual information based on the residual sample for the current block. The encoding device can generate residual information representing the quantized conversion coefficients. The residual information can be generated via various encoding methods such as exponential golomb, CAVLC, CABAC, etc.

[0181] The encoding device can generate a reconstructed sample. The encoding device can generate a reconstructed sample based on the residual information. The reconstructed sample can be generated by adding the residual sample based on the residual information to the predicted sample. Specifically, the encoding device can perform a prediction (intra or inter prediction) for the current block and generate a reconstructed sample based on the original sample and the predicted sample generated from the prediction.

[0182] A restored sample may include a restored luma sample and a restored chroma sample. Specifically, a residual sample may include a residual luma sample and a residual chroma sample. A residual luma sample can be generated based on an original luma sample and a predicted luma sample. A residual chroma sample can be generated based on an original chroma sample and a predicted chroma sample. The encoding device can derive conversion coefficients (luma conversion coefficients) for the residual luma sample and / or conversion coefficients (chroma conversion coefficients) for the residual chroma sample. Quantized conversion coefficients may include quantized luma conversion coefficients and / or quantized chroma conversion coefficients.

[0183] The encoding device can generate reference picture-related information (S730). The encoding device can generate reference picture-related information based on at least one reference picture. The reference picture-related information can be used for interpretation by the decoding device.

[0184] The encoding device can now generate in-loop filtering-related information for the restored picture sample (S740). The encoding device can perform an in-loop filtering procedure on the restored sample and generate in-loop filtering-related information based on the in-loop filtering procedure. For example, the in-loop filtering-related information may include the virtual boundary information described in this document (such as the SPS virtual boundary availability flag, picture header virtual boundary availability flag, SPS virtual boundary existence flag, picture header virtual boundary existence flag, and information regarding the location of the virtual boundary).

[0185] The encoding device can encode video / image information (S750). The image information may include residual information, prediction-related information, reference picture-related information, virtual boundary-related information (and / or additional virtual boundary-related information), and / or in-loop filtering-related information. The encoded video / image information can be output in bitstream form. The bitstream can be transmitted to a decoding device via a network or storage medium.

[0186] The aforementioned image / video information may include various types of information according to the embodiments of this document. For example, the image / video information may include information disclosed in at least one of the tables 1 to 14 described above.

[0187] In one embodiment, the image information may include an SPS (sequence parameter set). For example, whether the SPS includes additional virtual boundary-related information may be determined based on whether resampling for at least one reference picture is available. Here, resampling for at least one reference picture can be performed by the RPR (reference picture resampling) described above. The additional virtual boundary-related information may also be simply referred to as virtual boundary-related information. The term "additional" is used to distinguish it from virtual boundary-related information such as the SPS virtual boundary presence flag and / or the PH virtual boundary presence flag.

[0188] In one embodiment, the additional virtual boundary-related information may include the number of virtual boundaries and the locations of the virtual boundaries.

[0189] In one embodiment, the additional virtual boundary-related information may include information regarding the number of vertical virtual boundaries, information regarding the positions of vertical virtual boundaries, information regarding the number of horizontal virtual boundaries, and information regarding the positions of horizontal virtual boundaries.

[0190] In one embodiment, the image information may include a reference picture resampling availability flag. For example, it may be determined whether resampling is available for the at least one reference picture based on the reference picture resampling availability flag.

[0191] In one embodiment, the SPS may include an SPS virtual boundary presence flag related to whether or not the SPS includes the additional virtual boundary-related information. Based on the fact that resampling is available for the at least one reference picture, the value of the SPS virtual boundary presence flag may be determined to be 0.

[0192] In one embodiment, the additional virtual boundary-related information may not be included in the SPS based on the fact that resampling is available for the at least one reference picture. The image information includes picture header information, and the picture header information may include the additional virtual boundary-related information.

[0193] In one embodiment, the current picture may include a subpicture as just one subpicture. The subpicture may be independently coded. The restored sample may be generated based on the subpicture, subpicture-related information may be generated based on the subpicture, and the image information may include the subpicture-related information.

[0194] In one embodiment, the image information may not contain a picture-treated flag for the subpicture. Therefore, the value of the picture-treated flag for the subpicture can be set by the decoding device through inference (guessing or predicting). In one example, the value of the picture-treated flag for the subpicture may be set to 1.

[0195] In one embodiment, the current picture may include subpictures. In one example, the subpictures may be derived based on a merging procedure of two or more independently-coded subpictures. The restored sample may be generated based on the subpictures, subpicture-related information may be generated based on the subpictures, and the image information may include the subpicture-related information.

[0196] Figures 9 and 10 schematically show an example of a video / image decoding method and related components according to the embodiments (etc.) of this document.

[0197] The method disclosed in Figure 9 can be performed by the decoding device disclosed in Figure 3 or Figure 10. Specifically, for example, S900 in Figure 9 can be performed by the entropy decoding unit 310 of the decoding device, S910 in Figure 9 can be performed by the prediction unit 310 of the decoding device, S920 can be performed by the residual processing unit 320 and / or addition unit 340 of the decoding device, and S930 can be performed by the filtering unit 350 of the decoding device. The method disclosed in Figure 9 may include the embodiments described above in this document.

[0198] As shown in Figure 9, the decoding device can receive / acquire video / image information (S900). The video / image information may include residual information, prediction-related information, reference picture-related information, virtual boundary-related information (and / or additional virtual boundary-related information), and / or in-loop filtering-related information. The decoding device can receive / acquire the image / video information via a bitstream.

[0199] The aforementioned image / video information may include various types of information according to the embodiments of this document. For example, the image / video information may include information disclosed in at least one of the tables 1 to 14 described above.

[0200] The decoding device can derive quantized transformation coefficients. The decoding device can derive quantized transformation coefficients based on the residual information. The quantized transformation coefficients may have a one-dimensional vector form based on the coefficient scan order. The quantized transformation coefficients may include quantized luma transformation coefficients and / or quantized chroma transformation coefficients.

[0201] The decoding device can derive conversion coefficients. The decoding device can derive conversion coefficients based on an inverse quantization procedure for the quantized conversion coefficients. The decoding device can derive Luma conversion coefficients via inverse quantization based on quantized Luma conversion coefficients. The decoding device can derive Chroma conversion coefficients via inverse quantization based on quantized Chroma conversion coefficients.

[0202] The decoding device can generate / derive a resistive sample. The decoding device can derive a resistive sample based on an inverse transformation procedure for the transformation coefficients. The decoding device can derive a resistive luma sample via an inverse transformation procedure based on luma transformation coefficients. The decoding device can derive a resistive chroma sample via an inverse transformation procedure based on chroma transformation coefficients.

[0203] The decoding device can derive at least one reference picture based on reference picture-related information (S910). The decoding device can perform a prediction procedure based on the at least one reference picture. Specifically, the decoding device can generate prediction samples for the current block based on a prediction mode. In this case, various prediction methods disclosed in this document, such as interpretation or intrapretation, may be applied. The decoding device can generate prediction samples for the current block in the current picture based on the prediction procedure. For example, the decoding device can perform an interpretation procedure based on the at least one reference picture and then generate prediction samples based on the interpretation procedure.

[0204] The decoding device can generate / derive a reconstructed sample (S920). The decoding device can generate a reconstructed sample based on a predicted sample and a residual sample. The decoding device can generate a reconstructed sample based on the sum of the predicted sample and the original sample. For example, the decoding device can generate / derive a reconstructed lumina sample and / or a reconstructed chroma sample. The decoding device can generate a reconstructed lumina sample and / or a reconstructed chroma sample based on the residual information. The decoding device can generate a reconstructed sample based on the residual information. The reconstructed sample may include a reconstructed lumina sample and / or a reconstructed chroma sample. The lumina component of the reconstructed sample may correspond to the reconstructed lumina sample, and the chroma component of the reconstructed sample may correspond to the reconstructed chroma sample. The decoding device can generate a predicted lumina sample and / or a predicted chroma sample via a prediction procedure. The decoding device can generate a reconstructed lumina sample based on a predicted lumina sample and a residual lumina sample. The decoding device can generate a reconstructed chroma sample based on a predicted chroma sample and a residual chroma sample.

[0205] The decoding device can generate a corrected (filtered) reconstructed sample (S930). The decoding device can generate a corrected reconstructed sample based on an in-loop filtering procedure applied to the reconstructed sample. The decoding device can generate a corrected reconstructed sample based on in-loop filtering related information. The decoding device can utilize a deblocking procedure, an SAO procedure, and / or an ALF procedure to generate a corrected reconstructed sample.

[0206] In one embodiment, the image information may include SPS. For example, whether the SPS includes additional virtual boundary-related information may be determined based on whether resampling is available for the at least one reference picture. Here, resampling for the at least one reference picture can be performed by the RPR described above. The additional virtual boundary-related information may also be simply referred to as virtual boundary-related information. The term "additional" is used to distinguish it from virtual boundary-related information such as the SPS virtual boundary presence flag and / or the PH virtual boundary presence flag.

[0207] In one embodiment, the additional virtual boundary-related information may include the number of virtual boundaries and the locations of the virtual boundaries.

[0208] In one embodiment, the additional virtual boundary-related information may include information regarding the number of vertical virtual boundaries, information regarding the positions of vertical virtual boundaries, information regarding the number of horizontal virtual boundaries, and information regarding the positions of horizontal virtual boundaries.

[0209] In one embodiment, the image information may include a reference picture resampling availability flag. For example, it may be determined whether resampling is available for the at least one reference picture based on the reference picture resampling availability flag.

[0210] In one embodiment, the SPS may include an SPS virtual boundary presence flag related to whether or not the SPS includes the additional virtual boundary-related information. Based on the fact that resampling is available for the at least one reference picture, the value of the SPS virtual boundary presence flag may be determined to be 0.

[0211] In one embodiment, the additional virtual boundary-related information may not be included in the SPS based on the fact that resampling is available for the at least one reference picture. The image information includes picture header information, and the picture header information may include the additional virtual boundary-related information.

[0212] In one embodiment, the current picture may include a subpicture as just one subpicture. The subpicture may be coded independently. The restored sample may be generated based on the subpicture, subpicture-related information may be generated based on the subpicture, and the image information may include the subpicture-related information.

[0213] In one embodiment, the image information may not contain a picture-treated flag for the subpicture. Therefore, the value of the picture-treated flag for the subpicture can be set by the decoding device through inference (guessing or predicting). In one example, the value of the picture-treated flag for the subpicture may be set to 1.

[0214] In one embodiment, the current picture includes a subpicture. In one example, the subpicture is derived based on a merging procedure of two or more independently coded subpictures. The restored sample is generated based on the subpicture, subpicture-related information is generated based on the subpicture, and the image information includes the subpicture-related information.

[0215] The decoding device can receive information about the residual for the current block if a residual sample exists for the current block. The information about the residual may include transformation coefficients for the residual sample. Based on the residual information, the decoding device can derive a residual sample (or a residual sample array) for the current block. Specifically, the decoding device can derive quantized transformation coefficients based on the residual information. The quantized transformation coefficients may have a one-dimensional vector form based on the coefficient scan order. The decoding device can derive transformation coefficients based on an inverse quantization procedure for the quantized transformation coefficients. Based on the transformation coefficients, the decoding device can derive a residual sample.

[0216] The decoding device can generate a reconstructed sample based on an (intra) predicted sample and a residual sample, and derive a reconstructed block or reconstructed picture based on the reconstructed sample. Specifically, the decoding device can generate a reconstructed sample based on the sum of an (intra) predicted sample and a residual sample. Thereafter, as described above, the decoding device may apply deblocking filtering and / or in-loop filtering procedures such as the SAO procedure to the reconstructed picture to improve subjective / objective image quality as needed.

[0217] For example, a decoding device can decode a bitstream or encoded information to obtain image information that includes all or part of the aforementioned information (or syntax elements). Furthermore, the bitstream or encoded information can be stored on a computer-readable storage medium, which can trigger the aforementioned decoding method.

[0218] In the embodiments described above, the method is explained based on a flowchart as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps than those described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and that different steps may be included, or one or more steps in the flowchart may be omitted without affecting the scope of the embodiments described herein.

[0219] The methods according to the embodiments of this document described above can be implemented in software form, and the encoding and / or decoding devices relating to this document may be included in, for example, image processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.

[0220] In this document, when embodiments are implemented in software, the methods described above can be implemented by modules (processes, functions, etc.) that perform the functions described above. These modules are stored in memory and can be executed by a processor. The memory may be internal or external to the processor and may be connected to the processor by various well-known means. The processor may include an ASIC (application-specific integrated circuit), other chipsets, logic circuits, and / or data processing devices. The memory may include ROM (read-only memory), RAM (random access memory), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing may be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.

[0221] Furthermore, the decoding and encoding devices to which the embodiments of this document apply may include multimedia broadcasting transceivers, mobile communication terminals, home cinema video equipment, digital cinema video equipment, surveillance cameras, video interaction devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, customized video (VoD) service providers, OTT video (Over the top video) devices, internet streaming service providers, 3D video devices, VR (virtual reality) devices, AR (argumente reality) devices, image phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video equipment, and may be used to process video signals or data signals. For example, OTT video (Over the top video) devices may include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, DVRs (Digital Video Recorders), etc.

[0222] Furthermore, the processing methods to which the embodiments of this document apply can be produced in the form of programs executed on a computer and stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of this document can also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices that store data to be read by a computer. The computer-readable recording medium may include, for example, Blu-ray discs (BDs), general-purpose serial buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media implemented in the form of carrier waves (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method can be stored on a computer-readable recording medium or transmitted over a wired wireless network.

[0223] Furthermore, the embodiments described herein can be implemented as computer program products using program code, and the program code can be executed on a computer according to the embodiments described herein. The program code can be stored on a computer-readable carrier.

[0224] Figure 11 shows an example of a content streaming system to which the embodiments disclosed in this document may be applied.

[0225] Referring to Figure 11, the content streaming system to which the embodiments described in this document apply can broadly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0226] The encoding server is responsible for compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then transmitting this bitstream to the streaming server. As an alternative, if a multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server may be omitted.

[0227] The bitstream can be generated by an encoding method or a bitstream generation method to which an embodiment of this document applies, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0228] The streaming server transmits multimedia data to user devices based on user requests via a web server, and the web server acts as an intermediary to inform users about available services. When a user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. In this case, the content streaming system may include a separate control server, in which case the control server controls the commands and responses between the devices within the content streaming system.

[0229] The streaming server can receive content from a media storage and / or encoding server. For example, if it starts receiving content from the encoding server, it can receive the content in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0230] Examples of user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (such as smartwatches, smart glasses, and HMDs), digital TVs, desktop computers, and digital signage.

[0231] Each server within the aforementioned content streaming system can be operated as a distributed server, in which case the data received by each server can be processed in a distributed manner.

[0232] The claims described herein can be combined in various ways. For example, the technical features of the method claims herein can be combined to realize an apparatus, and the technical features of the apparatus claims herein can be combined to realize a method. Furthermore, the technical features of the method claims and the technical features of the apparatus claims herein can be combined to realize an apparatus, and the technical features of the method claims and the technical features of the apparatus claims herein can be combined to realize a method.

Claims

1. In an image decoding method performed by a decoding device, The steps include obtaining image information via a bitstream, including residual information, reference picture-related information, and in-loop filtering-related information. A step of deriving at least one reference picture, wherein the at least one reference picture is pointed to by the reference picture-related information, A step of generating a restored sample of the current picture based on the residual information and the at least one reference picture, A step of determining whether an in-loop filtering procedure is performed on the restored sample across a virtual boundary, based on the in-loop filtering-related information, wherein the in-loop filtering-related information indicates whether the in-loop filtering procedure is performed across the virtual boundary. The step of generating a corrected restored sample based on the in-loop filtering procedure on the restored sample includes, The aforementioned image information includes SPS (sequence parameter set) and picture header information. Whether the SPS includes virtual boundary-related information is determined based on whether reference picture resampling is possible for at least one of the reference pictures. A method wherein, based on the fact that the reference picture resampling is possible for the at least one reference picture, the picture header information includes the virtual boundary-related information, and the SPS does not include the virtual boundary-related information.

2. In an image encoding method performed by an encoding device, The current step is to generate a residual sample for the block, The steps include generating residual information based on the residual sample for the current block, Currently, the steps involve deriving at least one reference picture for the picture restoration sample, A step of generating reference picture-related information based on the at least one reference picture, A step of determining whether an in-loop filtering procedure is performed on the restored sample across a virtual boundary, A step of generating in-loop filtering related information for the restoration sample of the current picture, wherein the in-loop filtering related information indicates whether the in-loop filtering procedure is performed across the virtual boundary. The process includes the step of encoding image information including the residual information, the reference picture-related information, and the in-loop filtering-related information, The aforementioned image information includes SPS (sequence parameter set) and picture header information. Whether the SPS includes virtual boundary-related information is determined based on whether reference picture resampling is possible for at least one of the reference pictures. A method wherein, based on the fact that the reference picture resampling is possible for the at least one reference picture, the picture header information includes the virtual boundary-related information, and the SPS does not include the virtual boundary-related information.

3. Regarding methods for transmitting image-related data, A step of obtaining a bitstream relating to the image, wherein the bitstream is The current step is to generate a residual sample for the block, The steps include generating residual information based on the residual sample for the current block, Currently, the steps involve deriving at least one reference picture for the picture restoration sample, A step of generating reference picture-related information based on the at least one reference picture, A step of determining whether an in-loop filtering procedure is performed on the restored sample across a virtual boundary, A step of generating in-loop filtering related information for the restoration sample of the current picture, wherein the in-loop filtering related information indicates whether the in-loop filtering procedure is performed across the virtual boundary. A step of encoding image information including the residual information, the reference picture-related information, and the in-loop filtering-related information, which is generated based on the steps of: The step of transmitting the data, which includes the bitstream, The aforementioned image information includes SPS (sequence parameter set) and picture header information. Whether the SPS includes virtual boundary-related information is determined based on whether reference picture resampling is possible for at least one of the reference pictures. A method wherein, based on the fact that the reference picture resampling is possible for the at least one reference picture, the picture header information includes the virtual boundary-related information, and the SPS does not include the virtual boundary-related information.