Device and method for coding pictures to control loop filtering

Efficient filtering techniques for high-resolution video, including deblocking and ALF, applied across virtual boundaries and with sub-pictures as independent units, address the need for improved compression and quality in high-quality video formats like VR and AR, enhancing efficiency and reducing resource use.

JP2025113314APending Publication Date: 2025-08-01LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025081999
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-12
Filing Date
2025-05-15
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The increasing demand for high-resolution and high-quality video, including immersive media such as VR and AR content, has led to a need for more efficient video compression technologies to manage the higher data volumes and transmission costs associated with these formats.

Method used

The implementation of efficient filtering methods, including deblocking, SAO, and ALF, is performed based on virtual boundaries, with sub-pictures being independently coded and treated as pictures, and information signaling optimized for improved video coding efficiency.

Benefits of technology

This approach enhances overall video compression efficiency, improves subjective and objective visual quality, and reduces hardware resource consumption by efficiently managing filtering across virtual boundaries and signaling sub-picture information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025113314000001_ABST
    Figure 2025113314000001_ABST
Patent Text Reader

Abstract

To provide a device and a method for coding pictures to control in-loop filtering.SOLUTION: According to an embodiment of the present document, picture coding based on a sub-picture and / or a virtual boundary improves the subjective and objective quality and hardware resources necessary for coding is less consumed.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to a video coding apparatus and method for controlling loop filtering.

Background Art

[0002] In recent years, the demand for high-resolution and high-quality video / video such as 4K or UHD (Ultra High Definition) video / video of 8K or higher has been increasing in various fields. As the video / video data becomes higher in resolution and quality, the amount of information or bits transmitted relatively increases compared to existing video / video data. Therefore, when transmitting video data using a medium such as an existing wired or wireless broadband line, or storing video / video data using an existing storage medium, the transmission cost and storage cost increase.

[0003] In addition, in recent years, the interest and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms have been increasing, and the broadcast of video / video having video characteristics different from those of real-world video, such as game video, has been increasing.

[0004] Accordingly, there is a need for a highly efficient video / video compression technology to effectively compress, transmit, store, and reproduce information of high-resolution and high-quality video / video having various characteristics as described above.

[0005] Specifically, there is a discussion on a scheme for efficiently controlling loop filtering performed across virtual boundaries and a strategy for efficiently signaling information related to sub-pictures.

Summary of the Invention

Means for Solving the Problems

[0006] According to one embodiment of this document, a method and apparatus for enhancing the efficiency of video / video coding are provided.

[0007] According to one embodiment of this document, an efficient filtering application method and apparatus are provided.

[0008] According to one embodiment of this document, a method and apparatus for efficiently applying deblocking, SAO (sample adaptive loop), and ALF (adaptive loop filtering) are provided.

[0009] According to one embodiment of this document, in-loop filtering is performed based on a virtual boundary.

[0010] According to one embodiment of this document, based on whether resampling for a reference picture is available, it is determined whether an SPS (sequence parameter set) includes additional virtual boundary-related information (for example, information regarding the number and position of virtual boundaries).

[0011] According to one embodiment of this document, video coding is performed based on sub-pictures.

[0012] According to one embodiment of this document, the sub-pictures used for video coding are independently coded.

[0013] According to one embodiment of this document, a picture includes only one sub-picture, and the sub-picture is independently coded.

[0014] According to one embodiment of this document, a picture is generated based on a sub-picture merging procedure, and the sub-picture can be an independently coded sub-picture.

[0015] According to one embodiment of this document, the sub-pictures used for video coding are each treated as a picture.

[0016] According to one embodiment of this document, an encoding apparatus for performing video / video encoding is provided.

[0017] According to one embodiment of this document, there is provided a computer-readable digital storage medium storing encoded video / video information generated by a video / video encoding method disclosed in at least one of the embodiments of this document.

[0018] According to one embodiment of this document, there is provided a computer-readable digital storage medium storing encoded information or encoded video / video information that causes a decoding device to perform a video / video decoding method disclosed in at least one of the embodiments of this document.

Advantages of the Invention

[0019] According to one embodiment of this document, the overall video / video compression efficiency can be increased.

[0020] According to one embodiment of this document, subjective / objective visual quality can be improved through efficient filtering.

[0021] According to one embodiment of this document, a virtual boundary-based in-loop filtering procedure can be efficiently performed, and the filtering performance can be improved.

[0022] According to one embodiment of this document, information for virtual boundary-based in-loop filtering can be efficiently signaled.

[0023] According to one embodiment of this document, sub-picture related information can be efficiently signaled, thus improving the subjective / objective quality of the video and reducing the consumption of hardware resources required for coding.

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Best Mode for Carrying Out the Invention

[0025] This document can be modified in various ways and can have various embodiments. Specific embodiments will be illustrated in the drawings and described in detail. However, this is not intended to limit this document to specific embodiments. The terms commonly used in this specification are merely used to describe specific embodiments and are not intended to limit the technical concept of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "including" or "having" in this specification are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it should be understood that the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof is not precluded in advance.

[0026] On the other hand, each configuration in the drawings described in this document is independently illustrated for the convenience of explaining different characteristic functions, and it does not mean that each configuration is implemented by separate hardware or separate software. For example, among the configurations, two or more configurations can be combined to form one configuration, and one configuration can also be divided into multiple configurations. Embodiments in which each configuration is integrated and / or separated are included in the scope of rights of this document as long as they do not deviate from the essence of this document.

[0027] Hereinafter, with reference to the accompanying drawings, preferred embodiments of this document will be described in more detail. Hereinafter, the same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components will be omitted.

[0028] This document relates to video / video coding. For example, the methods / embodiments disclosed in this document relate to the VVC (Versatile Video Coding) standard (ITU-T Rec.H.266), next-generation video / image coding standards after VVC, or other video coding-related standards (e.g., the HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), EVC (essential video coding) standard, AVS2 standard, etc.).

[0029] This document presents various embodiments related to video / video coding, and unless otherwise stated, the embodiments can also be implemented in combination with each other.

[0030] In this document, video can mean a collection of a series of images over time. A picture generally means a unit representing one image at a specific time period, and a slice / tile is a unit that constitutes a part of a picture in coding. A slice / tile may include one or more CTUs (coding tree units). One picture may be composed of one or more slices / tiles. One picture may be composed of one or more tile groups. One tile group may include one or more tiles.

[0031] A pixel or pel can mean the smallest unit that constitutes one picture (or video). Also, the term "sample" may be used as a term corresponding to a pixel. A sample generally indicates a pixel or a pixel value, and can indicate only the pixel / pixel value of the luma component, or only the pixel / pixel value of the chroma component. Or, a sample can mean a pixel value in the spatial domain, and when such a pixel value is converted to the frequency domain, it can mean a conversion coefficient in the frequency domain.

[0032] A "unit" indicates the basic unit of video processing. A unit includes at least one of a specific area of a picture and information regarding the area. One unit includes one luma block and two chroma (e.g., cb, cr) blocks. A unit may, in some cases, be used interchangeably with terms such as "block" or "area". In general, an M×N block includes a set (or, array) of samples (or, sample array) or transform coefficients consisting of M columns and N rows.

[0033] In this document, " / " and "," are interpreted as "and / or". For example, "A / B" is interpreted as "A and / or B", and "A, B" is interpreted as "A and / or B". Additionally, "A / B / C" means "at least one of A, B, and / or C". Also, "A, B, C" also means "at least one of A, B, and / or C".

[0034] Additionally, in this document, "or" is interpreted as "and / or". For example, "A or B" can mean 1) only "A", 2) only "B", or 3) "A and B". In other words, "or" in this document can mean "additionally or alternatively".

[0035] In this specification, "at least one of A and B" can mean "only A", "only B", or "both A and B". Also, in this specification, expressions such as "at least one of A or B" and "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0036] Also, in this specification, "at least one of A, B and C" may mean "only A", "only B", "only C" or "any combination of A, B and C". Also, "at least one of A, B or C" and "at least one of A, B and / or C" may mean "at least one of A, B and C".

[0037] Also, the parentheses used in this specification may mean "for example". Specifically, when it is displayed as "prediction (intra prediction)", "intra prediction" may be proposed as an example of "prediction". In other words, "prediction" in this specification is not limited to "intra prediction", and "intra prediction" may be proposed as an example of "prediction". Also, when it is displayed as "prediction (i.e., intra prediction)", "intra prediction" may be proposed as an example of "prediction".

[0038] The technical features separately described in one drawing in this specification may be realized separately or simultaneously.

[0039] FIG. 1 schematically shows an example of a video / image coding system to which this document can be applied.

[0040] As shown in FIG. 1, the video / image coding system can include a source device and a receiving device. The source device can transmit encoded video / image information or data in file or streaming form to the receiving device via a digital storage medium or a network.

[0041] The source device can include a video source, an encoding device, and a transmitting unit. The receiving device can include a receiving unit, a decoding device, and a renderer. The encoding device can be called a video / video encoding device, and the decoding device can be called a video / video decoding device. A transmitter can be provided in the encoding device. A receiver can be provided in the decoding device. The renderer can include a display unit, and the display unit can also be composed of a separate device or an external component.

[0042] The video source can obtain video / video through processes such as video / video capture, synthesis, or generation. The video source can include a video / video capture device and / or a video / video generation device. The video / video capture device can include, for example, one or more cameras, a video / video archive containing previously captured video / video, etc. The video / video generation device can include, for example, a computer, a tablet, and a smartphone, etc., and can (electronically) generate video / video. For example, virtual video / video can be generated through a computer or the like, and in this case, the video / video capture process can be replaced by the process of generating related data.

[0043] The encoding device can encode the input video / video. The encoding device can perform a series of procedures such as prediction, transformation, quantization, etc. for compression and coding efficiency. The encoded data (encoded video / video information) can be output in the form of a bitstream.

[0044] The transmitting unit can transmit the encoded video / video information or data output in bitstream form to the receiving unit of the receiving device via a digital storage medium or a network in file or streaming form. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating a media file via a predetermined file format and can include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and transmit it to the decoding device.

[0045] The decoding device can decode the video / video by performing a series of procedures such as inverse quantization, inverse transformation, prediction, etc., corresponding to the operation of the encoding device.

[0046] The renderer can render the decoded video / video. The rendered video / video can be displayed via the display unit.

[0047] Figure 2 is a drawing schematically explaining the configuration of a video / video encoding device to which this document can be applied. Hereinafter, the video encoding device can include the video encoding device.

[0048] As shown in FIG. 2, the encoding device 200 can be configured to include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 can include an inter-prediction unit 221 and an intra-prediction unit 222. The residual processor 230 can include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 can further include a subtractor (231). The adder 250 can be called a reconstructor or a reconstructed block generator. The aforementioned image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 can be configured by one or more hardware components (e.g., an encoder chipset or a processor) according to an embodiment. Also, the memory 270 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 270 as an internal / external component.

[0049] The video segmentation unit 210 can divide the input video (or picture, frame) input to the encoding device 200 into one or more processing units. As an example, the processing unit can be called a coding unit (CU). In this case, the coding unit can be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) by a QTBTTT (Quad-tree binary-tree ternary-tree) structure. For example, one coding unit can be divided into a plurality of coding units with a deeper depth based on a quad-tree structure, a binary-tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure can be applied first, and then the binary-tree structure and / or the ternary structure can be applied. Or, the binary-tree structure can also be applied first. The coding procedure according to the present disclosure can be performed based on the final coding unit that is no longer divided. In this case, based on the coding efficiency according to the video characteristics, etc., the largest coding unit can be used as the final coding unit, or, if necessary, the coding unit can be recursively divided into coding units with a deeper depth so that the coding unit with the optimal size can be used as the final coding unit. Here, the coding procedure can include procedures such as prediction, transformation, and restoration described later. As another example, the processing unit can further include a prediction unit (PU: Prediction Unit) or a transform unit (TU: Transform Unit). In this case, the prediction unit and the transform unit can be divided or partitioned from the final coding unit described above, respectively.The prediction unit can be a unit of sample prediction, and the conversion unit can be a unit for deriving a conversion coefficient and / or a unit for deriving a residual signal from the conversion coefficient.

[0050] The unit can, in some cases, be used interchangeably with terms such as block or area. In general, an M×N block can represent a set such as samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luma component, or can also represent only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a pixel or a pel in one picture (or video).

[0051] The subtraction unit 231 can subtract the prediction signal (predicted block, predicted sample, or predicted sample array) output from the prediction unit 220 from the input video signal (original block, original sample, or original sample array) to generate a residual signal (residual block, residual sample, or residual sample array), and the generated residual signal is transmitted to the conversion unit 232. The prediction unit 220 can perform prediction on the processing target block (hereinafter referred to as the current block) and generate a predicted block including predicted samples for the current block. The prediction unit 220 can determine whether intra prediction or inter prediction is applied in units of the current block or CU. The prediction unit can generate various pieces of information related to prediction, such as prediction mode information, and transmit them to the entropy encoding unit 240, as will be described later in the description of each prediction mode. The information related to prediction can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0052] The intra prediction unit 222 can predict the current block by referring to samples within the current picture. The samples to be referred can be located adjacent to the current block or can be located remotely, depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes can include, for example, the DC mode and the Planar mode. The directional modes can include, for example, 33 directional prediction modes or 65 directional prediction modes, depending on the degree of fineness of the prediction direction. However, this is an example, and more or fewer directional prediction modes can be used depending on the setting. The intra prediction unit 222 can also determine the prediction mode to be applied to the current block by using the prediction mode applied to the adjacent blocks.

[0053] The inter prediction unit 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between adjacent blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent blocks can include spatial neighboring blocks existing in the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block can be the same or different. The temporal neighboring blocks can be called by names such as collocated reference blocks and collocated CUs (col CUs), and the reference picture including the temporal neighboring blocks can also be called a collocated picture (colPic). For example, the inter prediction unit 221 can construct a motion information candidate list based on adjacent blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter prediction unit 221 can use the motion information of adjacent blocks as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of the motion information prediction (motion vector prediction, MVP) mode, the motion vector of an adjacent block is used as a motion vector predictor, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0054] The prediction unit 220 can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for the prediction of one block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). In addition, the prediction unit can execute intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / motion video coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document.

[0055] The prediction signal generated via the inter prediction unit 221 and / or the intra prediction unit 222 can be used to generate a restored signal or can be used to generate a residual signal. The conversion unit 232 can apply a conversion technique to the residual signal to generate transform coefficients. For example, the conversion technique can include DCT (Discrete Cosine Transform), DST (Discrete Sine Transform), GBT (Graph-Based Transform), or CNT (Conditionally Non-linear Transform). Here, GBT means the conversion obtained from this graph when representing the relationship information between pixels in a graph. CNT means the conversion obtained based on generating a prediction signal using all previously reconstructed pixels. Also, the conversion process may be applied to a pixel block having the same size of a square or may be applied to a block of variable size that is not square.

[0056] The quantization unit 233 quantizes the transform coefficients and transmits them to the entropy encoding unit 240. The entropy encoding unit 240 can encode the quantized signal (information regarding the quantized transform coefficients) and output it as a bitstream. The information regarding the quantized transform coefficients can be called residual information. The quantization unit 233 can reorder the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate the information regarding the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. The entropy encoding unit 240 can perform various encoding methods such as, for example, exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit 240 can also encode, together or separately, information necessary for video / image restoration (e.g., values of syntax elements, etc.) in addition to the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in units of NAL (network abstraction layer) units in bitstream form. The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS), etc. Also, the video / video information can further include general constraint information. In this document, the signaling / transmitted information and / or syntax elements described later can be encoded through the above-described encoding procedure and included in the bitstream. The bitstream can be transmitted via a network or stored in a digital storage medium.Here, the network can include a broadcast network and / or a communication network, etc., and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The signal output from the entropy encoding unit 240 can be configured such that a transmission unit (not shown) for transmission and / or a storage unit (not shown) for storage are internal / external elements of the encoding device 200, or the transmission unit can also be included in the entropy encoding unit 240.

[0057] The quantized transform coefficients output from the quantization unit 233 can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235, a residual signal (residual block or residual sample) can be restored. The addition unit 155 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample, or reconstructed sample array) by adding the restored residual signal to the prediction signal output from the prediction unit 220. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the reconstructed block. The generated reconstructed signal can be used for intra prediction of the next block to be processed within the current picture, and as will be described later, it can also be used for inter prediction of the next picture after passing through filtering.

[0058] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied during the picture encoding and / or restoration process.

[0059] The filtering unit 260 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 260 can apply various filtering methods to the restored picture to generate a modified restored picture, and can store the modified restored picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filter, bilateral filter, and the like. The filtering unit 260 can generate various information related to filtering and transmit it to the entropy encoding unit 240 as will be described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoding unit 240 and output in the form of a bit stream.

[0060] The modified restored picture transmitted to the memory 270 can be used as a reference picture in the inter prediction unit 221. Through this, when inter prediction is applied, the encoding device can avoid prediction mismatches between the encoding device 100 and the decoding device, and can also improve the encoding efficiency.

[0061] The DPB of the memory 270 can store the modified restored picture for use as a reference picture in the inter prediction unit 221. The memory 270 can store the motion information of the blocks in which the motion information within the current picture has been derived (or encoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 221 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 270 can store the restored samples of the restored blocks within the current picture and transmit them to the intra prediction unit 222.

[0062] FIG. 3 is a drawing schematically explaining the configuration of a video / video decoding apparatus to which this document can be applied.

[0063] As shown in FIG. 3, the decoding apparatus 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 can include an inter-prediction unit 331 and an intra-prediction unit 332. The residual processor 320 can include a dequantizer 321 and an inverse transformer 321. The above-described entropy decoder 310, residual processor 320, predictor 330, adder 340, and filtering unit 350 can be configured by one hardware component (for example, a decoder chipset or a processor) according to an embodiment. Further, the memory 360 can include a DPB (decoded picture buffer) and can also be configured by a digital storage medium. The hardware component can further include the memory 360 as an internal / external component.

[0064] If a bitstream including video / video information is input, the decoding device 300 can restore the video corresponding to the process in which the video / video information was processed by the encoding device in FIG. 3. For example, the decoding device 300 can derive units / blocks based on the block splitting related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing units applied in the encoding device. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be split according to a quad-tree structure, a binary tree structure, and / or a ternary tree structure from a coding tree unit or a maximum coding unit. One or more transform units can be derived from the coding unit. Then, the restored video signal decoded and output via the decoding device 300 can be played back via a playback device.

[0065] The decoding device 300 can receive the signal output from the encoding device in FIG. 3 in the form of a bitstream, and the received signal can be decoded via the entropy decoding unit 310. For example, the entropy decoding unit 310 can parse the bitstream to derive information (e.g., video / video information) necessary for video restoration (or, picture restoration). The video / video information can further include information regarding various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Also, the video / video information can further include general constraint information. The decoding device can further decode a picture based on the information regarding the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this document can be decoded via the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 can decode the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the value of the syntax element necessary for video restoration and the quantized value of the transform coefficient regarding the residual. More specifically, the CABAC entropy decoding method receives the bin corresponding to each syntax element in the bitstream, determines a context model using the syntax element information to be decoded, the information adjacent to the decoding target block, and the decoding information of the decoding target block or the information of the symbol / bin decoded in the previous step, predicts the occurrence probability of the bin based on the determined context model, performs arithmetic decoding of the bin, and can generate a symbol corresponding to the value of each syntax element. At this time, the CABAC entropy decoding method can update the context model using the information of the symbol / bin decoded for the context model of the next symbol / bin after determining the context model.Among the information decoded by the entropy decoding unit 310, the information related to prediction is provided to the prediction unit 330, and the residual information that has undergone entropy decoding in the entropy decoding unit 310, that is, the quantized transform coefficients and related parameter information, can be input to the inverse quantization unit 321. Also, among the information decoded by the entropy decoding unit 310, the information related to filtering can be provided to the filtering unit 350. On the other hand, a receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can also be a component of the entropy decoding unit 310. On the other hand, the decoding device according to this document can be called a video / video / picture decoding device, and the decoding device can also be classified into an information decoder (video / video / picture information decoder) and a sample decoder (video / video / picture sample decoder). The information decoder can include the entropy decoding unit 310, and the sample decoder can include at least one of the inverse quantization unit 321, the inverse transform unit 322, the prediction unit 330, the addition unit 340, the filtering unit 350, and the memory 360.

[0066] In the inverse quantization unit 321, the quantized transform coefficients can be inverse quantized to output transform coefficients. The inverse quantization unit 321 can reorder the quantized transform coefficients in a two-dimensional block form. In this case, the reordering can be performed based on the coefficient scan order performed in the encoding device. The inverse quantization unit 321 can perform inverse quantization on the quantized transform coefficients using a quantization parameter (for example, quantization step size information) to obtain transform coefficients.

[0067] In the inverse transform unit 322, the transform coefficients are inverse transformed to obtain a residual signal (residual block, residual sample array).

[0068] The prediction unit can perform a prediction on the current block and generate a predicted block that includes a prediction sample for the current block. The prediction unit can determine whether intra prediction or inter prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit 310, and can determine a specific intra / inter prediction mode.

[0069] The prediction unit can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can not only apply intra prediction or inter prediction for the prediction of one block, but also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP). Also, the prediction unit can execute intra block copy (IBC) for the prediction of a block. The intra block copy can be used for content video / moving picture coding such as games, for example, like SCC (screen content coding). IBC basically performs prediction within the current picture, but can be performed in the same way as inter prediction in terms of deriving a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction.

[0070] The intra prediction unit 331 can predict the current block by referring to samples within the current picture. The samples to be referred to can be located adjacent to the current block or at a distance therefrom depending on the prediction mode. In intra prediction, the prediction mode can include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 can also determine the prediction mode to be applied to the current block using the prediction mode applied to an adjacent block.

[0071] The inter prediction unit 332 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of the motion information between an adjacent block and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter prediction, the adjacent block can include a spatial neighboring block existing within the current picture and a temporal neighboring block existing in the reference picture. For example, the inter prediction unit 332 can construct a motion information candidate list based on adjacent blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and the information regarding the prediction can include information indicating the mode of inter prediction for the current block.

[0072] The adder 340 can generate a restored signal (restored picture, restored block, restored sample array) by adding the acquired residual signal to the predicted signal (predicted block, predicted sample array) output from the predictor. When there is no residual for the block to be processed, as in the case where the skip mode is applied, the predicted block can be used as the restored block.

[0073] The adder 340 can be called a restoration unit or a restored block generation unit. The generated restored signal can be used for intra prediction of the next block to be processed in the current picture, can be output after filtering as will be described later, or can also be used for inter prediction of the next picture.

[0074] On the other hand, LMCS (luma mapping with chroma scaling) can also be applied in the picture decoding process.

[0075] The filtering unit 350 can apply filtering to the restored signal to improve the subjective / objective image quality. For example, the filtering unit 350 can apply various filtering methods to the restored picture to generate a modified restored picture, and can send the modified restored picture to the memory 360, specifically, to the DPB of the memory 360. The various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.

[0076] The (corrected) restored picture stored in the DPB of the memory 360 can be used as a reference picture in the inter prediction unit 332. The memory 360 can store the motion information of the block for which the motion information in the current picture has been derived (or decoded) and / or the motion information of the blocks in the already restored picture. The stored motion information can be transmitted to the inter prediction unit 332 for utilization as the motion information of spatially adjacent blocks or temporally adjacent blocks. The memory 360 can store the restored samples of the restored blocks in the current picture and can transmit them to the intra prediction unit 331.

[0077] In this specification, the embodiments described in the prediction unit 330, inverse quantization unit 321, inverse transform unit 322, filtering unit 350, etc. of the decoding device 300 can be applied so as to be the same or corresponding to the prediction unit 220, inverse quantization unit 234, inverse transform unit 235, filtering unit 260, etc. of the encoding device 200, respectively.

[0078] As described above, when performing video coding, prediction is performed to improve the compression efficiency. Through this, a predicted block including prediction samples for the current block, which is the block to be coded, can be generated. Here, the predicted block includes prediction samples in the spatial domain (or pixel domain). The predicted block is also derived in the encoding device and the decoding device, and the encoding device can improve the video coding efficiency by signaling information (residual information) regarding the residual between the original block and the predicted block, which is not the original sample value of the original block, to the decoding device. The decoding device can derive a residual block including residual samples based on the residual information, and can generate a restored block including restored samples by combining the residual block and the predicted block, and can generate a restored picture including the restored block.

[0079] The residual information can be generated through conversion and quantization procedures. For example, an encoding device can derive a residual block between the original block and the predicted block, execute a conversion procedure on the residual samples (residual sample array) included in the residual block to derive conversion coefficients, and execute a quantization procedure on the conversion coefficients to derive quantized conversion coefficients, so as to signal (via a bitstream) the relevant residual information to a decoding device. Here, the residual information can include information such as the value information, position information, conversion technique, conversion kernel, quantization parameter, etc. of the quantized conversion coefficients. The decoding device can execute an inverse quantization / inverse conversion procedure based on the residual information to derive residual samples (or a residual block). The decoding device can generate a restored picture based on the predicted block and the residual block. Also, the encoding device can inverse quantize / inverse convert the quantized conversion coefficients for reference in the inter-prediction of subsequent pictures to derive a residual block, and generate a restored picture based on this.

[0080] In this document, at least one of quantization / inverse quantization and / or conversion / inverse conversion can be omitted. When the quantization / inverse quantization is omitted, the quantized conversion coefficients can be referred to as conversion coefficients. When the conversion / inverse conversion is omitted, the conversion coefficients can also be referred to as coefficients or residual coefficients, or, for the sake of uniformity of expression, can still be referred to as conversion coefficients.

[0081] In this document, the quantized transform coefficients and the transform coefficients can each be referred to as a transform coefficient and a scaled transform coefficient, respectively. In this case, the residual information can include information regarding the transform coefficient(s), and the information regarding the transform coefficient(s) can be signaled via a residual coding syntax. The transform coefficient can be derived based on the residual information (or the information regarding the transform coefficient(s)), and the scaled transform coefficient can be derived via an inverse transform (scaling) with respect to the transform coefficient. The residual sample can be derived based on an inverse transform (transformation) with respect to the scaled transform coefficient. This can be applied / expressed similarly in other parts of this document.

[0082] The prediction unit of the encoding device / decoding device can derive a prediction sample by performing inter prediction in units of blocks. Inter prediction can indicate a prediction derived in a method that depends on data elements (e.g., sample values or motion information) of pictures other than the current picture. When inter prediction is applied to the current block, a predicted block (predicted sample array) for the current block can be induced based on a reference block (reference sample array) specified by a motion vector on a reference picture pointed to by the index of the reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information of the current block can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between adjacent blocks and the current block. The motion information can include a motion vector and an index of a reference picture. The motion information can further include information on an inter prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). When inter prediction is applied, the adjacent blocks can include spatial neighboring blocks existing within the current picture and temporal neighboring blocks existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block can be called by names such as a collocated reference block and a collocated CU (colCU), and the reference picture including the temporal neighboring block may also be called a collocated picture (colPic). For example, a candidate list of motion information can be configured based on adjacent blocks of the current block, and flag or index information indicating which candidate is selected (used) can be signaled to derive the motion vector and / or index of the reference picture of the current block. Inter prediction is performed based on various prediction modes. For example, in the case of skip mode and merge mode, the motion information of the current block can be the same as the motion information of the selected adjacent block.In the case of the skip mode, unlike the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the selected adjacent block is used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived by using the sum of the motion vector predictor and the motion vector difference.

[0083] The motion information can include L0 motion information and / or L1 motion information according to the inter-prediction type (L0 prediction, L1 prediction, Bi prediction, etc.). The motion vector in the L0 direction may be referred to as the L0 motion vector or MVL0, and the motion vector in the L1 direction may be referred to as the L1 motion vector or MVL1. The prediction based on the L0 motion vector may be called L0 prediction, the prediction based on the L1 motion vector may be called L1 prediction, and the prediction based on both the L0 motion vector and the L1 motion vector may be called bi (Bi) prediction. Here, the L0 motion vector can indicate the motion vector related to the reference picture list L0 (L0), and the L1 motion vector can indicate the motion vector related to the reference picture list L1 (L1). The reference picture list L0 can include, as reference pictures, pictures that are earlier in output order than the current picture, and the reference picture list L1 can include pictures that are later in output order than the current picture. The earlier picture may be referred to as a forward (reference) picture, and the later picture may be referred to as a backward (reference) picture. The reference picture list L0 can further include, as reference pictures, pictures that are later in output order than the current picture. In this case, the earlier picture may be indexed first within the reference picture list L0, and the later picture may be indexed thereafter. The reference picture list L1 can further include, as reference pictures, pictures that are earlier in output order than the current picture. In this case, the later picture may be indexed first within the reference picture list 1, and the earlier picture may be indexed thereafter. Here, the output order may correspond to the POC (picture order count) order (order).

[0084] FIG. 4 exemplarily shows a hierarchical structure for the coded video / picture.

[0085] Referring to FIG. 4, the coded video is divided into a VCL (Video Coding Layer) that handles the decoding process of the video and itself, a lower system that transmits and stores the encoded information, and a NAL (Network Abstraction Layer) that exists between the VCL and the lower system and is responsible for the network adaptation function.

[0086] In the VCL, VCL data including compressed video data (slice data) can be generated, or parameter sets including information such as a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), and a Video Parameter Set (VPS), or a Supplemental Enhancement Information (SEI) message that is additionally required in the decoding process of the video can be generated.

[0087] In the NAL, a NAL unit can be generated by adding header information (NAL unit header) to the RBSP (Raw Byte Sequence Payload) generated in the VCL. At this time, the RBSP means slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header can include NAL unit type information specified by the RBSP data included in the corresponding NAL unit.

[0088] As shown in the above drawings, the NAL unit can be divided into a VCL NAL unit and a Non-VCL NAL unit according to the RBSP generated in the VCL. The VCL NAL unit can mean a NAL unit including information (slice data) for the video, and the Non-VCL NAL unit can mean a NAL unit including information (parameter set or SEI message) required for decoding the video.

[0089] The above-mentioned VCL NAL units and Non-VCL NAL units can be transmitted via a network with header information attached according to the data standard of the lower-level system. For example, the NAL unit can be transformed into a data form of a predetermined standard such as the H.266 / VVC file format, RTP (Real-time Transport Protocol), TS (Transport Stream), etc., and can be transmitted via various networks.

[0090] As mentioned above, the NAL unit type can be specified by the RBSP data structure included in the corresponding NAL unit, and the information regarding such NAL unit type can be stored in the NAL unit header and signaled.

[0091] For example, depending on whether the NAL unit contains information (slice data) for video, it can be broadly classified into a VCL NAL unit type and a Non-VCL NAL unit type. The VCL NAL unit type can be classified according to the nature and type of the picture included in the VCL NAL unit, and the Non-VCL NAL unit type can be classified according to the type of parameter set, etc.

[0092] The following is an example of the NAL unit type specified by the type of parameter set included in the Non-VCL NAL unit type, etc.

[0093] - APS (Adaptation Parameter Set) NAL unit: The type for the NAL unit containing APS

[0094] - DPS (Decoding Parameter Set) NAL unit: The type for the NAL unit containing DPS

[0095] -VPS (Video Parameter Set) NAL unit: Type for NAL unit containing VPS

[0096] -SPS (Sequence Parameter Set) NAL unit: Type for NAL unit containing SPS

[0097] -PPS (Picture Parameter Set) NAL unit: Type for NAL unit containing PPS

[0098] -PH (Picture header) NAL unit: Type for NAL unit containing PH

[0099] The above-mentioned NAL unit types have syntax information for the NAL unit type, and the syntax information can be stored in the NAL unit header and signaled. For example, the syntax information is nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.

[0100] On the other hand, as described above, one picture can include a plurality of slices, and one slice can include a slice header and slice data. In this case, one picture header can be further added for a plurality of slices (slice header and slice data set) within one picture. The picture header (picture header syntax) can include information / parameters that are commonly applicable to the picture. In this document, a slice can be mixed or replaced with a tile group. Also, in this document, a slice header can be mixed or replaced with a type group header.

[0101] The slice header (slice header syntax, slice header information) can include information / parameters that are commonly applicable to the slice. The APS (APS syntax) or PPS (PPS syntax) can include information / parameters that are commonly applicable to one or more slices or pictures. The SPS (SPS syntax) can include information / parameters that are commonly applicable to one or more sequences. The VPS (VPS syntax) can include information / parameters that are commonly applicable to multiple layers. The DPS (DPS syntax) can include information / parameters that are commonly applicable to the entire video. The DPS can include information / parameters related to the concatenation of CVS (coded video sequence). In this document, the high level syntax (HLS) can include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.

[0102] In this document, the video / video information encoded from an encoding device and signaled in bitstream form to a decoding device can include not only information related to partitioning within a picture, intra / inter prediction information, residual information, in-loop filtering information, etc., but also information included in the slice header, information included in the picture header, information included in the APS, information included in the PPS, information included in the SPS, information included in the VPS, and / or information included in the DPS. Further, the video / video information can further include information of the NAL unit header.

[0103] On the one hand, in order to compensate for the difference between the original video and the restored video due to errors occurring in the compression encoding process such as quantization, as described above, an in-loop filtering procedure can be executed on the restored sample or the restored picture. As described above, in-loop filtering can be executed in the filter section of the encoding device and the filter section of the decoding device, and a deblocking filter, SAO, and / or an adaptive loop filter (ALF) can be applied. For example, the ALF procedure can be executed after the deblocking filtering procedure and / or the SAO procedure is completed. However, also in this case, the deblocking filtering procedure and / or the SAO procedure can be omitted.

[0104] Specific descriptions for picture restoration and filtering are described below. In video / video coding, restored blocks can be generated based on intra prediction / inter prediction in each block unit, and a restored picture including the restored blocks can be generated. When the current picture / slice is an I picture / slice, the blocks included in the current picture / slice can be restored based only on intra prediction. On the other hand, when the current picture / slice is a P or B picture / slice, the blocks included in the current picture / slice can be restored based on intra prediction or inter prediction. In this case, intra prediction can be applied to some blocks within the current picture / slice, and inter prediction can also be applied to the remaining blocks.

[0105] Intra prediction can indicate a prediction that generates a prediction sample for a current block based on reference samples within a picture to which the current block belongs (hereinafter referred to as the current picture). When intra prediction is applied to the current block, adjacent reference samples to be used for intra prediction of the current block can be derived. The adjacent reference samples of the current block can include a total of 2×nH samples adjacent to the left boundary of the current block of size nW×nH and adjacent to the bottom - left, samples adjacent to the top boundary of the current block and a total of 2×nW samples adjacent to the top - right, and 1 sample adjacent to the top - left of the current block. Or, the adjacent reference samples of the current block can also include a plurality of columns of upper - adjacent samples and a plurality of rows of left - adjacent samples. Also, the adjacent reference samples of the current block can include a total of nH samples adjacent to the right boundary of the current block of size nW×nH, a total of nW samples adjacent to the bottom boundary of the current block, and 1 sample adjacent to the bottom - right of the current block.

[0106] However, some of the adjacent reference samples of the current block may not yet be decoded or may not be available. In this case, the decoder can substitute samples that are not available with samples that are available to form adjacent reference samples to be used for prediction. Or, adjacent reference samples to be used for prediction can be formed through interpolation of available samples.

[0107] When an adjacent reference sample is derived, (i) a predicted sample can be derived based on the average or interpolation of neighboring reference samples of the current block, and (ii) the predicted sample can also be derived based on reference samples that exist in a specific (predicted) direction with respect to the predicted sample among the neighboring reference samples of the current block. In the case of (i), it is called a non-directional mode or a non-angle mode, and in the case of (ii), it can be called a directional mode or an angular mode. Also, among the adjacent reference samples, based on the predicted sample of the current block, the predicted sample can also be generated through interpolation between the second adjacent sample and the first adjacent sample that are located in the direction opposite to the prediction direction of the intra prediction mode of the current block. In the case described above, it can be called linear interpolation intra prediction (LIP). Also, a chroma predicted sample can be generated based on a luma sample using a linear model. In this case, it can be called the LM mode. Also, a temporary predicted sample of the current block is derived based on the filtered adjacent reference samples, and a predicted sample of the current block is derived by performing a weighted sum of at least one reference sample derived by the intra prediction mode among the existing adjacent reference samples, that is, the non-filtered adjacent reference samples, and the temporary predicted sample. In the case described above, it can be called PDPC (Position dependent intra prediction). Also, intra prediction coding can be performed by selecting the reference sample line with the highest prediction accuracy from among the adjacent multiple reference sample lines of the current block and using the reference sample located in the prediction direction on the corresponding line, and indicating (signaling) the used reference sample line to the decoding device.In the above-described case, it can be called multi-reference line (MRL) intra prediction or MRL-based intra prediction. Also, the current block can be divided into vertical or horizontal sub-partitions, and intra prediction can be performed based on the same intra prediction mode, and adjacent reference samples can be derived and used in units of the sub-partitions. That is, in this case, the intra prediction mode for the current block is also applied to the sub-partitions, and by deriving and using adjacent reference samples in units of the sub-partitions, in some cases, the intra prediction performance can be improved. Such a prediction method can be called intra sub-partitions (ISP) or ISP-based intra prediction. The intra prediction methods described above can be called intra prediction types, distinguished from the intra prediction modes in Tables of Contents 1 and 2. The intra prediction type can be called by various terms such as an intra prediction technique or an additional intra prediction mode. For example, the intra prediction type (or, an additional intra prediction mode, etc.) can include at least one of the above-described LIP, PDPC, MRL, and ISP. A general intra prediction method excluding specific intra prediction types such as the above-described LIP, PDPC, MRL, and ISP can be called a normal intra prediction type. The normal intra prediction type can be generally applied when the above-described specific intra prediction types are not applied, and prediction can be performed based on the above-described intra prediction mode. On the other hand, if necessary, post-processing filtering for the derived prediction samples can also be performed.

[0108] Specifically, the intra prediction procedure can include an intra prediction mode / type determination step, an adjacent reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. Also, if necessary, a post-processing filtering step for the derived prediction samples can be performed.

[0109] A restored picture modified by an in-loop filtering procedure is generated, and the modified restored picture is output from the decoding device as a decoded picture, and is also stored in the decoded picture buffer or memory of the encoding device / decoding device, and can be used as a reference picture in the inter prediction procedure during picture encoding / decoding later. The in-loop filtering procedure includes, as described above, a deblocking filtering procedure, a SAO (sample adaptive offset) procedure, and / or an ALF (adaptive loop filter) procedure, etc. In this case, one or a part of the deblocking filtering procedure, SAO (sample adaptive offset) procedure, ALF (adaptive loop filter) procedure, and bilateral filter procedure may be sequentially applied, or all of them may be sequentially applied. For example, after the deblocking filtering procedure is applied to the restored picture, the SAO procedure may be performed. Or, for example, after the deblocking filtering procedure is applied to the restored picture, the ALF procedure may be performed. This is also performed in the same way in the encoding device.

[0110] Deblocking filtering is a filtering technique that removes distortions occurring at the boundaries between blocks in the restored picture. The deblocking filtering procedure can, for example, derive a target boundary in the restored picture, determine bS (boundary strength) for the target boundary, and perform deblocking filtering for the target boundary based on the bS. The bS can be determined based on, for example, the prediction modes of two adjacent blocks of the target boundary, the motion vector difference, whether the reference pictures are the same, and the presence or absence of non-zero valid coefficients.

[0111] SAO is a method for compensating the offset difference between the restored picture and the original picture in sample units, and can be applied based on types such as, for example, Band Offset and Edge Offset. According to SAO, samples can be classified into different categories by each SAO type, and an offset value can be added to each sample based on the category. The filtering information for SAO can include information on whether SAO can be applied, SAO type information, SAO offset value information, etc. SAO can also be applied to the restored picture after the deblocking filtering is applied.

[0112] ALF (Adaptive Loop Filter) is a technique for filtering in sample units based on filter coefficients according to the filter shape for the restored picture. The encoding device can determine, through comparison between the restored picture and the original picture, whether ALF can be applied, the ALF shape and / or the ALF filtering coefficient, etc., and can signal them to the decoding device. That is, the filtering information for ALF can include information on whether ALF can be applied, ALF filter shape information, ALF filtering coefficient information, etc. ALF can also be applied to the restored picture after the deblocking filtering is applied.

[0113] Figure 5 shows an example of the ALF filter shape.

[0114] Fig. 5(a) shows a 7×7 diamond filter shape, and (b) shows a 5×5 diamond filter shape. In Fig. 5, Cn within the filter shape indicates the filter coefficient. In the said Cn, when n is the same, it indicates that the same filter coefficient can be assigned. In this document, the position and / or unit where the filter coefficient is assigned according to the filter shape of the ALF can be called a filter tab. At this time, one filter coefficient can be assigned to each filter tab, and the form in which the filter tabs are arranged can correspond to the filter shape. The filter tab located at the center of the filter shape can be called the center filter tab. The same filter coefficient can be assigned to two filter tabs with the same n value existing at positions corresponding to each other with reference to the center filter tab. For example, in the case of a 7×7 diamond filter shape, it includes 25 filter tabs, and since the filter coefficients from C0 to C11 are assigned in a centrally symmetric form, only 13 filter coefficients can be used to assign filter coefficients to the said 25 filter tabs. Also, for example, in the case of a 5×5 diamond filter shape, it includes 13 filter tabs, and since the filter coefficients from C0 to C5 are assigned in a centrally symmetric form, only 7 filter coefficients can be used to assign filter coefficients to the said 13 filter tabs. For example, in order to reduce the data amount of information regarding the signaled filter coefficients, among the 13 filter coefficients for the 7×7 diamond filter shape, 12 filter coefficients are (explicitly) signaled, and 1 filter coefficient can be (implicitly) derived. Also, for example, among the 7 filter coefficients for the 5×5 diamond filter shape, 6 filter coefficients are (explicitly) signaled, and 1 filter coefficient can be (implicitly) derived.

[0115] Fig. 6 is a flowchart for explaining a filtering-based encoding method in an encoding device. The method in Fig. 6 includes steps from S600 to S630.

[0116] In step S600, the encoding device generates a reconstructed picture. Step S600 is performed based on the above-described reconstructed picture (or, reconstructed sample) generation procedure.

[0117] In step S610, the encoding device determines whether in-loop filtering is applied (across virtual boundaries) based on in-loop filtering related information. Here, the in-loop filtering includes at least one of the above-described deblocking filtering, SAO, or ALF.

[0118] In step S620, the encoding device generates a modified reconstructed picture (modified reconstructed sample) based on the determination in step S610. Here, the modified reconstructed picture (modified reconstructed sample) may be a filtered reconstructed picture (filtered reconstructed sample).

[0119] In step S630, the encoding device encodes video / picture information including in-loop filtering related information based on the in-loop filtering procedure.

[0120] FIG. 7 is a flowchart for explaining a filtering-based decoding method in a decoding device. The method of FIG. 7 includes steps S700 to S730.

[0121] In step S700, the decoding device obtains video / picture information including in-loop filtering related information from the bitstream. Here, the bitstream is based on the encoded video / picture information transmitted from the encoding device.

[0122] In step S710, the decoding device generates a reconstructed picture. Step S710 is performed based on the above-described reconstructed picture (or, reconstructed sample) generation procedure.

[0123] In step S720, the decoding device determines whether in-loop filtering is to be applied (across the virtual boundary) based on the in-loop filtering related information. Here, the in-loop filtering includes at least one of the aforementioned deblocking filtering, SAO, or ALF.

[0124] In step S730, the decoding device generates a modified reconstructed picture (modified reconstructed samples) based on the determination in step S720. Here, the modified reconstructed picture (modified reconstructed samples) can be the filtered reconstructed picture (filtered reconstructed samples).

[0125] As described above, the in-loop filtering procedure can be applied to the reconstructed picture. In this case, a virtual boundary can be defined to further enhance the subjective / objective visual quality of the reconstructed picture, and the in-loop filtering procedure can also be applied across the virtual boundary. The virtual boundary includes discontinuous edges such as, for example, 360-degree video, VR video, or PIP (picture in picture). For example, the virtual boundary exists at a predetermined and agreed-upon position, and its existence and / or position can be signaled. As an example, the virtual boundary is located at the fourth sample line above the CTU row (specifically, for example, above the fourth sample line above the CTU row). As another example, information regarding its existence and / or position can also be signaled via HLS. The HLS includes SPS, PPS, picture header, slice header, etc. as described above.

[0126] Hereinafter, the high-level syntax signaling and semantics regarding the embodiments of this document will be described.

[0127] One embodiment of this document includes a method of controlling a loop filter. This method of controlling the loop filter can be applied to a reconstructed picture. The in-loop filter (loop filter) can be used for decoding an encoded bitstream. The loop filter includes the aforementioned deblocking, SAO, and ALF. The SPS includes flags related to each of deblocking, SAO, and ALF. The flag indicates whether each tool is available for coding of a CLVS (coded layer video sequence) and a CVS (coded video sequence) that refers to the SPS.

[0128] In one example, when the loop filter is available for coding a picture within a CVS, the application of the loop filter is controlled so as not to cross a specific boundary. For example, the loop filter can be controlled so as not to cross a sub-picture boundary, the loop filter can be controlled so as not to cross a tile boundary, the loop filter can be controlled so as not to cross a slice boundary, and / or the loop filter can be controlled so as not to cross a virtual boundary.

[0129] The in-loop filtering related information includes the information, syntax, syntax elements, and / or semantics described in this document (or the embodiments included therein). The in-loop filtering related information includes information regarding whether the in-loop filtering procedure (in whole or in part) is available across a specific boundary (e.g., a virtual boundary, a sub-picture boundary, a slice boundary, and / or a tile boundary). The video information included in the bitstream includes high level syntax (HLS), and the HLS includes the in-loop filtering related information. Based on the determination as to whether the in-loop filtering procedure is applied across a specific boundary, modified (or filtered) reconstructed samples (reconstructed picture) are generated. In one example, when the in-loop filtering procedure is disabled for all blocks / boundaries, the modified reconstructed samples may be the same as the reconstructed samples. In another example, the modified reconstructed samples include the modified reconstructed samples derived based on in-loop filtering. However, in this case, a part of the reconstructed samples (e.g., the reconstructed samples across the virtual boundary) may not be in-loop filtered based on the determination. For example, the reconstructed samples across a specific boundary (including at least one of a virtual boundary, a sub-picture boundary, a slice boundary, and / or a tile boundary where the execution of in-loop filtering is enabled) can be in-loop filtered, but the reconstructed samples across other boundaries (including at least one of a virtual boundary, a sub-picture boundary, and / or a tile boundary where the execution of in-loop filtering is disabled) may not be in-loop filtered.

[0130] In one example, in relation to whether the in-loop filtering procedure is performed across the virtual boundary, the in-loop filtering related information includes the SPS virtual boundary presence flag, the picture header virtual boundary presence flag, the information regarding the number of virtual boundaries, the information regarding the position of the virtual boundaries, etc.

[0131] In the embodiments included in this document, the information regarding the position of the virtual boundary includes the information regarding the x - coordinate of the vertical virtual boundary and / or the information regarding the y - coordinate of the horizontal virtual boundary. Specifically, the information regarding the position of the virtual boundary includes the information regarding the x - coordinate of the vertical virtual boundary and / or the y - coordinate of the horizontal virtual boundary in terms of luma sample units. Also, the information regarding the position of the virtual boundary includes the information regarding the number of information (syntax elements) regarding the x - coordinate of the vertical virtual boundary existing in the SPS. Also, the information regarding the position of the virtual boundary includes the information regarding the number of information (syntax elements) regarding the y - coordinate of the horizontal virtual boundary existing in the SPS. Or, the information regarding the position of the virtual boundary includes the information regarding the number of information (syntax elements) regarding the x - coordinate of the vertical virtual boundary existing in the picture header. Also, the information regarding the position of the virtual boundary includes the information regarding the number of information (syntax elements) regarding the y - coordinate of the horizontal virtual boundary existing in the picture header.

[0132] The following table shows the exemplary syntax and semantics of the SPS (sequence parameter set) according to this embodiment.

[0133]

Table 1

[0134]

Table 2

[0135] The following table shows the exemplary syntax and semantics of the PPS (picture parameter set) according to this embodiment.

[0136]

Table 3

[0137]

Table 4

[0138] The following table shows the exemplary syntax and semantics of the picture header according to this embodiment.

[0139]

Table 5-1

[0140]

Table 5-2

[0141]

Table 6-1

[0142]

Table 6-2

[0143] The following table shows the exemplary syntax and semantics of the slice header according to this embodiment.

[0144]

Table 7

[0145]

Table 8

[0146] Hereinafter, information related to sub-pictures, information related to virtual boundaries that can be used in in-loop filtering, and their signaling will be described.

[0147] If there is no sub-picture having a boundary that is treated as a picture boundary even though the picture contains a plurality of sub-pictures, the advantages of using sub-pictures cannot be utilized. In one embodiment of this document, video / video information for video coding includes information for treating sub-pictures as pictures, which is called a picture handling flag (e.g., subpic_treated_as_pic_flag[i]).

[0148] To signal the layout of sub-pictures, a flag related to whether sub-pictures exist (e.g., subpic_present_flag) is signaled. This may be called the sub-picture existence flag. When the value of subpic_present_flag is 1, information regarding the number of sub-pictures that divide the picture (e.g., sps_num_subpics_minus1) is signaled. In one example, the number of sub-pictures that divide the picture may be the same as sps_num_subpics_minus1 + 1 (add 1 to sps_num_subpics_minus1). The possible values of sps_num_subpics_minus1 include 0, which means that only one sub-picture exists within the picture. When the picture contains only one sub-picture, since the sub-picture itself is a picture, the signaling of sub-picture related information is regarded as a redundant procedure.

[0149] In existing embodiments, when a picture contains only one sub-picture and there is sub-picture signaling, the value of the picture handling flag (e.g., subpic_treated_as_pic_flag[i]) and / or the value of the flag related to whether loop filtering is performed across the sub-picture (e.g., loop_filter_across_enabled_flag) can be 0 or 1. Here, when the value of subpic_treated_as_pic_flag[i] is 0, a problem occurs that conflicts with the case where the sub-picture boundary is the picture boundary. This requires additional duplicate procedures to confirm to the decoder that the picture boundary is the sub-picture boundary.

[0150] When a picture is generated based on the merging procedure of two or more sub-pictures, all sub-pictures used in the merging procedure must be independently coded sub-pictures (sub-pictures with the value of the picture handling flag (subpic_treated_as_pic_flag[i]) being 1). This is because when merging a sub-picture that is not an independently coded sub-picture (referred to as the "first sub-picture" in this paragraph), problems after merging may occur due to the blocks within the first sub-picture being coded with reference to reference blocks existing outside the first sub-picture.

[0151] Also, when a picture is partitioned into sub-pictures, the sub-picture ID signaling may or may not exist. When the sub-picture ID signaling exists, the sub-picture ID signaling exists in (is included in) the SPS, PPS, and / or the picture header (PH). When the sub-picture ID signaling does not exist in the SPS, it includes the case where a bitstream is generated as a result of the sub-picture merging procedure. Therefore, when the sub-picture ID signaling is not included in the SPS, it is preferable that all sub-pictures are independently coded.

[0152] In a video coding procedure in which a virtual boundary is used, information regarding the position of the virtual boundary can be signaled in the SPS or the picture header. Signaling information regarding the position of the virtual boundary in the SPS means that there is no change in the position within the CLVS. However, when reference picture resampling (RPR) is enabled for the CLVS, pictures within the CLVS have different sizes. Here, reference picture resampling (also called adaptive resolution change (ARC)) is performed for the normal coding operation of pictures having different resolutions (spatial resolutions). For example, reference picture resampling includes upsampling and downsampling. Through reference picture resampling, high coding efficiency for the adaptation of bitrate and spatial resolution is achieved. Considering reference picture resampling, it is necessary to ensure that the position of the virtual boundary is all within one picture.

[0153] In the existing ALF procedure, the k-th order exponential Golomb code with k = 3 is used to signal the absolute values of the luma and chroma ALF coefficients. However, the k-th order exponential Golomb coding causes a considerable computational overhead and complexity and becomes a problem.

[0154] The embodiments described in the following paragraphs propose solutions for solving the aforementioned problems. The embodiments may be applied independently. Or, at least two or more embodiments may be combined and applied.

[0155] In one embodiment of this document, when subpicture signaling exists and a picture has only one subpicture, the only one subpicture is an independently coded subpicture. For example, when a picture has only one subpicture, the only one subpicture is an independently coded subpicture, and the value of the picture handling flag (e.g., subpic_treated_as_pic_flag[i]) for the only one subpicture is 1. Thereby, duplicate procedures related to the subpicture can be omitted.

[0156] In one embodiment of this document, when subpicture signaling exists, the number of subpictures may be more than one. In one example, when subpicture signaling exists (e.g., the value of subpics_present_flag is 1), the information regarding the number of subpictures (e.g., sps_num_subpics_minus1) is greater than 0, and the number of subpictures can be sps_num_subpics_minus1 + 1 (add 1 to sps_num_subpics_minus1). In another example, the information regarding the number of subpictures is sps_num_subpics_minus2, and the number of subpictures can be sps_num_subpics_minus2 + 2 (add 2 to sps_num_subpics_minus2). In yet another example, the subpicture presence flag subpics_present_flag can be replaced with the information regarding the number of subpictures sps_num_subpics_minus1. Thus, the subpicture signaling can exist when sps_num_subpics_minus1 is greater than 0.

[0157] In one embodiment of this document, when a picture is divided into sub-pictures, at least one of the sub-pictures can be a separately coded sub-picture. Here, the value of the picture handling flag (e.g., subpic_treated_as_pic_flag[i]) for the separately coded sub-picture is 1.

[0158] In one embodiment of this document, the sub-pictures of a picture based on the merging procedure of two or more sub-pictures can be separately coded sub-pictures.

[0159] In one embodiment of this document, when sub-picture ID (identification) signaling exists at other positions (other syntax, other high-level syntax information) than in the SPS, all sub-pictures are separately coded sub-pictures, and the value of the picture handling flag (e.g., subpic_treated_as_pic_flag) for all sub-pictures can be 1. In one example, the sub-picture ID signaling exists in the PPS, and in this case, all sub-pictures can be separately coded sub-pictures. In another example, the sub-picture ID signaling exists in the picture header, and in this case, all sub-pictures can be separately coded sub-pictures.

[0160] In one embodiment of this document, when virtual boundary signaling exists in the SPS for CLVS and reference picture resampling is possible, all horizontal virtual boundary positions are within the minimum picture height of the picture referring to the SPS, and all vertical virtual boundary positions can be within the minimum picture width of the picture referring to the SPS.

[0161] In one embodiment of this document, when reference picture resampling (RPR) is enabled, the virtual boundary signaling is included in the picture header. That is, when reference picture resampling is enabled, the virtual boundary signaling may not be included in the SPS.

[0162] In one embodiment of this document, fixed length coding (FLC) with the number of bits (or bit length) associated therewith is used to signal ALF data. In one example, the information regarding the ALF data includes information regarding the bit length of the ALF luma coefficient absolute value (e.g., alf_luma_coeff_abs_len_minus1) and / or information regarding the bit length of the ALF chroma coefficient absolute value (e.g., alf_chroma_coeff_abs_len_minus1). For example, the information regarding the bit length of the ALF luma coefficient absolute value and / or the information regarding the bit length of the ALF chroma coefficient absolute value can be ue(v) coded.

[0163] The following table shows an exemplary syntax of the SPS according to this embodiment.

[0164]

Table 9

[0165] The following table shows exemplary semantics regarding the syntax elements included in the said syntax.

[0166]

Table 10

[0167] Next, an exemplary syntax of the SPS according to this embodiment is shown.

[0168]

Table 11

[0169] The following table shows exemplary semantics regarding the syntax elements included in the said syntax.

[0170]

Table 12

[0171] The following table shows an exemplary syntax of the ALF data according to this embodiment.

[0172]

Table 13

[0173] The following table shows exemplary semantics regarding the syntax elements included in the syntax.

[0174]

Table 14

[0175] According to the embodiments of the present document described together with the above table, video coding based on sub-pictures and / or virtual boundaries improves the subjective / objective quality of the video and reduces the consumption of hardware resources required for coding.

[0176] FIG. 8 and FIG. 9 schematically show an example of a video / video encoding method and related components according to the embodiments of the present document.

[0177] The method disclosed in FIG. 8 can be performed by the encoding device disclosed in FIG. 2 or FIG. 9. Specifically, for example, S800 and S810 in FIG. 8 are performed by the residual processing unit 230 of the encoding device in FIG. 9, S820 and / or S830 in FIG. 8 are performed by the filtering unit 260 of the encoding device in FIG. 9, and S840 in FIG. 8 can be performed by the entropy encoding unit 240 of the encoding device in FIG. 9. Also, although not shown in FIG. 8, a prediction sample or prediction-related information is derived by the prediction unit 220 of the encoding device in FIG. 8, and a bitstream is generated from the residual information or prediction-related information by the entropy encoding unit 240 of the encoding device. The method disclosed in FIG. 8 includes the foregoing embodiments in this document.

[0178] As shown in FIG. 8, the encoding device derives a residual sample (S800). The encoding device derives a residual sample for the current block, and the residual sample for the current block is derived based on the original sample and the prediction sample of the current block. Specifically, the encoding device derives the prediction sample of the current block based on the prediction mode. In this case, various prediction methods disclosed in this document, such as inter prediction or intra prediction, are applicable. A residual sample is derived based on the prediction sample and the original sample. For inter prediction, the encoding device derives at least one reference picture, and inter prediction is performed based on the at least one reference picture. A prediction sample is generated based on the inter prediction. The encoding device generates reference picture-related information based on the at least one reference picture.

[0179] The encoding device can derive conversion coefficients. The encoding device can derive conversion coefficients based on a conversion procedure for the residual sample. For example, the conversion procedure includes at least one of DCT, DST, GBT, or CNT.

[0180] The encoding device derives the quantized transform coefficients. The encoding device derives the quantized transform coefficients based on the quantization procedure for the transform coefficients. The quantized transform coefficients have a one-dimensional vector form based on the coefficient scan order.

[0181] The encoding device generates residual information (S810). The encoding device generates residual information based on the residual samples for the current block. The encoding device generates residual information indicating the quantized transform coefficients. The residual information is generated by various encoding methods such as exponential Golomb, CAVLC, CABAC, etc.

[0182] The encoding device generates restored samples. The encoding device generates restored samples based on the residual information. The restored samples are generated by adding the residual samples based on the residual information to the predicted samples. Specifically, the encoding device performs prediction (intra or inter prediction) for the current block and generates restored samples based on the original samples and the predicted samples generated from the prediction.

[0183] The restored samples include restored luma samples and restored chroma samples. Specifically, the residual samples include residual luma samples and residual chroma samples. The residual luma samples are generated based on the original luma samples and the predicted luma samples. The residual chroma samples are generated based on the original chroma samples and the predicted chroma samples. The encoding device derives the transform coefficients (luma transform coefficients) for the residual luma samples and / or the transform coefficients (chroma transform coefficients) for the residual chroma samples. The quantized transform coefficients include the quantized luma transform coefficients and / or the quantized chroma transform coefficients.

[0184] The encoding device determines whether the in-loop filtering procedure is performed across the virtual boundary (S820). Here, the virtual boundary may be the same as the aforementioned virtual boundary. The in-loop filtering procedure includes at least one of a deblocking procedure, an SAO procedure, or an ALF procedure.

[0185] The encoding device generates in-loop filtering related information (S830). The encoding device generates in-loop filtering related information based on the determination in the step S820. Here, the in-loop filtering related information indicates the information used to perform the in-loop filtering procedure. For example, the in-loop filtering related information includes information regarding the virtual boundary described in this document (SPS virtual boundary existence flag, picture header virtual boundary existence flag, information regarding the number of virtual boundaries, information regarding the position of the virtual boundary, etc.).

[0186] The encoding device encodes video / video information (S840). The video information includes residual information, prediction related information, reference picture related information, sub-picture related information, in-loop filtering related information and / or virtual boundary related information (and / or additional virtual boundary related information). The encoded video / video information is output in the form of a bitstream. The bitstream is transmitted to the decoding device via a network or a storage medium.

[0187] The video / video information includes various information according to the embodiments of this document. For example, the video / video information includes the information disclosed in at least one of the aforementioned Tables 1 to 14.

[0188] In an embodiment, the video information includes an SPS (sequence parameter set) and picture header information that refers to the SPS. For example, based on whether reference picture resampling is available, it is determined whether additional virtual boundary related information is included in the SPS or the picture header information. Here, reference picture resampling is performed by the aforementioned RPR (reference picture resampling). The additional virtual boundary related information may simply be referred to as virtual boundary related information. "Additional" is used for the purpose of distinction from virtual boundary related information such as an SPS virtual boundary presence flag and / or a PH virtual boundary presence flag.

[0189] In one embodiment, the additional virtual boundary related information includes the number of virtual boundaries and the positions of the virtual boundaries.

[0190] In one embodiment, the additional virtual boundary related information includes information regarding the number of vertical virtual boundaries, information regarding the positions of the vertical virtual boundaries, information regarding the number of horizontal virtual boundaries, and information regarding the positions of the horizontal virtual boundaries.

[0191] In one embodiment, the video information includes a reference picture resampling availability flag. For example, based on the reference picture resampling availability flag, it is determined whether the reference picture resampling is available.

[0192] In one embodiment, the SPS includes an SPS virtual boundary presence flag related to whether the SPS includes the additional virtual boundary related information. Based on the availability of the reference picture resampling, the value of the SPS virtual boundary presence flag is determined to be 0.

[0193] In one embodiment, based on the availability of the reference picture resampling, the additional virtual boundary related information may not be included in the SPS. In this case, for example, the picture header information includes the additional virtual boundary related information.

[0194] In one embodiment, the current picture includes the subpicture as just one subpicture. The subpicture can be independently coded. The restored sample is generated based on the subpicture, subpicture-related information is generated based on the subpicture, and the video information includes the subpicture-related information.

[0195] In one embodiment, there may be no picture handling flag for the subpicture in the video information. Therefore, the value of the picture handling flag for the subpicture can be set by inference (speculation or prediction) at the decoding stage. In one example, the value of the picture handling flag for the subpicture can be set to 1.

[0196] In one embodiment, the current picture includes the subpicture. In one example, the subpicture is derived based on a merging procedure of two or more independently-coded-subpictures. The restored sample is generated based on the subpicture, subpicture-related information is generated based on the subpicture, and the video information includes the subpicture-related information.

[0197] FIG. 10 and FIG. 11 schematically show an example of a video / video decoding method and related components according to the embodiment(s) of this document.

[0198] The method disclosed in FIG. 10 can be performed by the decoding device disclosed in FIG. 3 or FIG. 11. Specifically, for example, S1000 in FIG. 10 is performed by the entropy decoding unit 310 of the decoding device, S1010 is performed by the residual processing unit 320 and / or the addition unit 340 of the decoding device, and S1020 is performed by the filtering unit 350 of the decoding device. The method disclosed in FIG. 10 includes the foregoing embodiments in this document.

[0199] As shown in FIG. 10, the decoding device receives / acquires video / video information (S1000). The video / video information includes residual information, prediction-related information, reference picture-related information, sub-picture-related information, in-loop filtering-related information, and / or virtual boundary-related information (and / or additional virtual boundary-related information). The decoding device can receive / acquire the video / video information via a bitstream.

[0200] The video / video information includes various information according to the embodiments of this document. For example, the video / video information includes the information disclosed in at least one of Tables 1 to 14 described above.

[0201] The decoding device can derive quantized transform coefficients. The decoding device can derive quantized transform coefficients based on the residual information. The quantized transform coefficients have a one-dimensional vector form based on the coefficient scan order. The quantized transform coefficients include quantized luma transform coefficients and / or quantized chroma transform coefficients.

[0202] The decoding device derives transform coefficients. The decoding device derives transform coefficients based on an inverse quantization procedure for the quantized transform coefficients. The decoding device derives luma transform coefficients by inverse quantization based on the quantized luma transform coefficients. The decoding device derives chroma transform coefficients by inverse quantization based on the quantized chroma transform coefficients.

[0203] The decoding device generates / derives residual samples. The decoding device derives residual samples based on an inverse conversion procedure for the conversion coefficients. The decoding device derives residual luma samples by an inverse conversion procedure based on luma conversion coefficients. The decoding device derives residual chroma samples by an inverse conversion procedure based on chroma conversion coefficients.

[0204] The decoding device derives at least one reference picture based on reference picture related information. The decoding device performs a prediction procedure based on the at least one reference picture. Specifically, the decoding device generates prediction samples of the current block based on a prediction mode. In this case, various prediction methods disclosed in this document, such as inter prediction or intra prediction, are applied. The decoding device generates prediction samples for the current block in the current picture based on the prediction procedure. For example, the decoding device performs an inter prediction procedure based on the at least one reference picture and generates prediction samples based on the inter prediction procedure.

[0205] The decoding device generates / derives restored samples (S1010). For example, the decoding device generates / derives restored luma samples and / or restored chroma samples. The decoding device generates restored luma samples and / or restored chroma samples based on the residual information. The decoding device generates restored samples based on the residual information. The restored samples include restored luma samples and / or restored chroma samples. The luma component of the restored samples corresponds to the restored luma samples, and the chroma component of the restored samples corresponds to the restored chroma samples. The decoding device generates predicted luma samples and / or predicted chroma samples by a prediction procedure. The decoding device generates restored luma samples based on the predicted luma samples and the residual luma samples. The decoding device generates restored chroma samples based on the predicted chroma samples and the residual chroma samples.

[0206] The decoding device generates a modified (filtered) restored sample (S1020). The decoding device generates a modified restored sample based on an in-loop filtering procedure for the restored sample. The decoding device generates a modified restored sample based on in-loop filtering related information. The decoding device utilizes a deblocking procedure, an SAO procedure, and / or an ALF procedure to generate the modified restored sample.

[0207] In one embodiment, the video information includes an SPS and picture header information that refers to the SPS. For example, based on whether reference picture resampling is available, it is determined whether additional virtual boundary related information is included in the SPS or the picture header information. Here, the reference picture resampling is performed by the aforementioned RPR. The additional virtual boundary related information may simply be referred to as virtual boundary related information. "Additional" is used for the purpose of distinction from virtual boundary related information such as an SPS virtual boundary presence flag and / or a PH virtual boundary presence flag.

[0208] In one embodiment, the additional virtual boundary related information includes the number of virtual boundaries and the positions of the virtual boundaries.

[0209] In one embodiment, the additional virtual boundary related information includes information regarding the number of vertical virtual boundaries, information regarding the positions of the vertical virtual boundaries, information regarding the number of horizontal virtual boundaries, and information regarding the positions of the horizontal virtual boundaries.

[0210] In one embodiment, the video information includes a reference picture resampling available flag. For example, based on the reference picture resampling available flag, it is determined whether the reference picture resampling is available.

[0211] In one embodiment, the SPS includes an SPS virtual boundary presence flag related to whether the SPS includes the additional virtual boundary related information. Based on the availability of the reference picture sampling, the value of the SPS virtual boundary presence flag can be determined to be 0.

[0212] In one embodiment, based on the availability of the reference picture sampling, the additional virtual boundary related information may not be included in the SPS. In this case, for example, the picture header information includes the additional virtual boundary related information.

[0213] In one embodiment, the current picture includes sub-pictures as just one sub-picture. The sub-pictures can be coded independently. The restored samples are generated based on the sub-pictures, sub-picture related information is generated based on the sub-pictures, and the video information includes the sub-picture related information.

[0214] In one embodiment, there may be no picture handling flag for the sub-picture in the video information. Therefore, the value of the picture handling flag for the sub-picture is set by inference (speculation or prediction) by the decoding device. In one example, the value of the picture handling flag for the sub-picture is set to 1.

[0215] In one embodiment, the current picture includes sub-pictures. In one example, the sub-pictures are derived based on a merging procedure of two or more independently-coded sub-pictures. The restored samples are generated based on the sub-pictures, sub-picture related information is generated based on the sub-pictures, and the video information includes the sub-picture related information.

[0216] When there are residual samples for the current block, the decoding device can receive information regarding the residual for the current block. The information regarding the residual can include transform coefficients regarding the residual samples. The decoding device can derive residual samples (or a residual sample array) for the current block based on the residual information. Specifically, the decoding device can derive quantized transform coefficients based on the residual information. The quantized transform coefficients can have a one-dimensional vector form based on the coefficient scan order. The decoding device can derive transform coefficients based on an inverse quantization procedure for the quantized transform coefficients. The decoding device can derive residual samples based on the transform coefficients.

[0217] The decoding device can generate restored samples based on (intra) prediction samples and residual samples, and can derive a restored block or a restored picture based on the restored samples. Specifically, the decoding device can generate restored samples based on the sum of (intra) prediction samples and residual samples. Thereafter, as described above, the decoding device can apply in-loop filtering procedures such as deblocking filtering and / or SAO procedures to the restored picture, if necessary, to improve subjective / objective image quality.

[0218] For example, the decoding device can decode a bitstream or encoded information and obtain video information including all or part of the aforementioned information (or syntax elements). Also, the bitstream or encoded information can be stored in a computer-readable storage medium and can be caused to perform the aforementioned decoding method.

[0219] In the foregoing embodiments, the method has been described based on a flowchart as a series of steps or blocks, but the corresponding embodiments are not limited to the order of the steps, and a certain step may occur in a different order from the steps described above or simultaneously with different steps. Also, those skilled in the art can understand that the steps shown in the flowchart are not exclusive, and different steps may be included, or one or more steps of the flowchart may be deleted without affecting the scope of the embodiments of this document.

[0220] The method according to the embodiments of this document described above can be embodied in the form of software, and the encoding device and / or decoding device according to this document can be included in a device that performs video processing, such as a TV, computer, smartphone, set-top box, display device, etc.

[0221] In this document, when an embodiment is embodied in software, the method described above can be embodied by modules (processes, functions, etc.) that perform the functions described above. The modules can be stored in a memory and executed by a processor. The memory may be inside or outside the processor and may be connected to the processor by various well-known means. The processor can include an ASIC (application-specific integrated circuit), other chip sets, logic circuits, and / or data processing devices. The memory can include a ROM (read-only memory), RAM (random access memory), flash memory, memory card, storage medium, and / or other storage devices. That is, the embodiments described in this document can be embodied and performed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each drawing can be embodied and performed on a computer, processor, microprocessor, controller, or chip. In this case, the information for the embodiment (e.g., information on instructions) or algorithms can be stored in a digital storage medium.

[0222] In addition, the decoding device and encoding device to which the embodiments of this document are applied may be included in a multimedia broadcast transceiver, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conferencing device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, an on-demand video (VoD) service providing device, an OTT video (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a videophone video device, a transportation means terminal (e.g., a vehicle terminal including an autonomous driving vehicle, an airplane terminal, a ship terminal, etc.) and a medical video device, etc., and may be used to process video signals or data signals. For example, the OTT video (Over the top video) device may include a game console, a Blu-ray player, an Internet access TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recoder), etc.

[0223] In addition, the processing method to which the embodiments of this document are applied can be produced in the form of a program executed by a computer and can be stored in a recording medium readable by the computer. Multimedia data having a data structure according to the embodiments of this document can also be stored in a recording medium readable by the computer. The recording medium readable by the computer includes all types of storage devices and distributed storage devices in which data readable by the computer is stored. The recording medium readable by the computer can include, for example, Blu-ray Disc (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage devices. Further, the recording medium readable by the computer includes a medium embodied in the form of a carrier wave (for example, transmission via the Internet). Also, a bitstream generated by an encoding method can be stored in a recording medium readable by the computer or transmitted via a wired or wireless communication network.

[0224] In addition, the embodiments of this document can be embodied as a computer program product by program code, and the program code can be executed by a computer according to the embodiments of this document. The program code can be stored on a carrier readable by a computer.

[0225] FIG. 12 shows an example of a content streaming system to which the embodiments disclosed in this document can be applied.

[0226] Referring to FIG. 12, the content streaming system to which the embodiments of this document are applied can include, generally, an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.

[0227] The encoding server compresses the content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream, and serves to transmit this to the streaming server. As another example, when a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.

[0228] The bitstream can be generated by an encoding method or a method for generating a bitstream to which the embodiments of this document are applicable, and the streaming server can temporarily store the bitstream in the process of transmitting or receiving the bitstream.

[0229] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server serves as a medium to inform the user of what services are available. If the user requests a desired service from the web server, the web server transmits this to the streaming server, and the streaming server transmits multimedia data to the user. At this time, the content streaming system can include another control server, and in this case, the control server serves to control commands / responses between each device in the content streaming system.

[0230] The streaming server can receive content from a media repository and / or an encoding server. For example, when it comes to receiving content from the encoding server, the content can be received in real time. In this case, in order to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.

[0231] In the example of the user device, there may be a mobile phone, a smart phone, a laptop computer, a digital broadcast terminal, a PDA (personal digital assistants), a PMP (portable multimedia player), a navigation device, a slate PC, a tablet PC, an ultrabook, a wearable device (for example, a smartwatch, a smart glass, an HMD (head mounted display)), a digital TV, a desktop computer, a digital signage, and the like.

[0232] Each server in the content streaming system can be operated as a distributed server. In this case, the data received by each server can be distributedly processed.

[0233] The claims described in this specification can be combined in various ways. For example, the technical features of the method claims in this specification can be combined and implemented as a device, and the technical features of the device claims in this specification can be combined and implemented as a method. Also, the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented as a device, and the technical features of the method claims in this specification and the technical features of the device claims can be combined and implemented as a method.

Claims

1. In a video decoding method performed by a decoding apparatus, a step of acquiring video information including prediction-related information and residual information via a bitstream; a step of generating prediction samples for a current block based on the prediction-related information; a step of generating residual samples for the current block based on the residual information; a step of generating restored samples for the current block based on the prediction samples and the residual samples; a step of generating modified restored samples based on an in-loop filtering procedure for the restored samples, comprising: the video information includes an SPS (sequence parameter set) and picture header information referring to the SPS; based on whether reference picture resampling is available, it is determined whether additional virtual boundary-related information is included in the SPS or in the picture header information; when the reference picture resampling is available, the additional virtual boundary-related information is included in the picture header information; the additional virtual boundary-related information includes information on the number of vertical virtual boundaries, information on the positions of the vertical virtual boundaries, information on the number of horizontal virtual boundaries, and information on the positions of the horizontal virtual boundaries, a video decoding method.

2. In a video encoding method performed by an encoding apparatus, a step of deriving prediction samples for a current block based on a prediction mode; a step of generating prediction-related information for the prediction mode; a step of deriving residual samples for the current block; a step of generating residual information based on the residual samples; a step of determining whether an in-loop filtering procedure is performed across virtual boundaries; a step of generating in-loop filtering-related information for restored samples of the current block based on the determination; a step of encoding video information including the prediction-related information, the residual information, and the in-loop filtering-related information, comprising: the video information includes an SPS (sequence parameter set) and picture header information referring to the SPS; Based on whether reference picture sampling is available, it is determined whether additional virtual boundary related information is included in the SPS or in the picture header information. Based on the case where the reference picture sampling is available, the additional virtual boundary related information is included in the picture header information. The additional virtual boundary related information includes information on the number of vertical virtual boundaries, information on the positions of the vertical virtual boundaries, information on the number of horizontal virtual boundaries, and information on the positions of the horizontal virtual boundaries, a video encoding method. **Claim 3** In a method for transmitting data for a video, A step of obtaining a bitstream for the video, where the bitstream A step of deriving prediction samples for a current block based on a prediction mode; A step of generating prediction related information for the prediction mode; A step of deriving residual samples for the current block; A step of generating residual information based on the residual samples; A step of determining whether an in-loop filtering procedure is performed across a virtual boundary; A step of generating in-loop filtering related information for the restored samples of the current block based on the determination; A step generated based on the steps of encoding video information including the prediction related information, the residual information, and the in-loop filtering related information; A step of transmitting the data including the bitstream, including; The video information includes an SPS (sequence parameter set) and picture header information referring to the SPS. Based on whether reference picture sampling is available, it is determined whether additional virtual boundary related information is included in the SPS or in the picture header information. Based on the case where the reference picture sampling is available, the additional virtual boundary related information is included in the picture header information. The additional virtual boundary related information includes information on the number of vertical virtual boundaries, information on the positions of the vertical virtual boundaries, information on the number of horizontal virtual boundaries, and information on the positions of the horizontal virtual boundaries, a transmission method.