Image decoding / encoding method and bit stream transmission method

By selectively encoding the size information of the slices, the problem of increased information content in the transmission of high-resolution and high-quality images is solved, the encoding/decoding efficiency is improved, and the transmission and storage costs are reduced.

CN120956893APending Publication Date: 2025-11-14NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511227930.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-03-09
Filing Date
2021-03-08
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

With the increasing demand for high-resolution and high-quality images, existing technologies have led to an increase in the amount of information transmitted in image data, resulting in higher transmission and storage costs. Efficient image compression technologies are needed to solve this problem.

Method used

An improved encoding/decoding method and device are generated by selectively encoding the size information of the slices, including the width and height information indicating the tile columns and tile behavior units, and then transmitted and stored via bitstream.

Benefits of technology

It improves the efficiency of image encoding/decoding, reduces transmission and storage costs, and achieves efficient image information transmission and storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956893A_ABST
    Figure CN120956893A_ABST
Patent Text Reader

Abstract

The invention provides an image decoding / encoding method and a bitstream transmission method. An image decoding method performed by an image decoding device according to the present disclosure may comprise the steps of: acquiring size information indicating a size of a current slice corresponding to at least a part of a current picture from a bitstream; and determining a size of the current slice based on the size information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the original invention patent application No. 202180033162.3 (International Application No.: PCT / KR2021 / 002822, Application Date: March 8, 2021, Invention Title: Image Encoding / Decoding Method and Apparatus for Selectively Encoding Size Information of Rectangular Slices and Method for Transmitting Bit Streams). Technical Field

[0002] This disclosure relates to image encoding / decoding methods and apparatus, and more specifically, to image encoding and decoding methods and apparatus for selectively encoding slice size information, and to a method for transmitting a bitstream generated by the image encoding method / apparatus of this disclosure. Background Technology

[0003] Recently, the demand for high-resolution and high-quality images (such as high-definition (HD) and ultra-high-definition (UHD) images) is increasing across various fields. With the improvement in image data resolution and quality, the amount of information or bits transmitted increases relatively compared to existing image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.

[0004] Therefore, efficient image compression technology is needed to effectively send, store, and reproduce information about high-resolution and high-quality images. Summary of the Invention

[0005] Technical issues

[0006] The purpose of this disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.

[0007] Another object of this disclosure is to provide an image encoding / decoding method and apparatus that improves encoding / decoding efficiency by selectively encoding the size information of slices.

[0008] Another object of this disclosure is to provide a method for transmitting a bitstream generated by an image encoding method or device according to this disclosure.

[0009] Another object of this disclosure is to provide a recording medium storing a bitstream generated by an image encoding method or apparatus according to this disclosure.

[0010] Another object of this disclosure is to provide a recording medium that stores a bitstream that is received, decoded, and used to reconstruct an image by an image decoding device according to this disclosure.

[0011] The technical problems solved by this disclosure are not limited to those described above. Other technical problems not described herein will become clear to those skilled in the art through the following description.

[0012] Technical solution

[0013] An image decoding method performed by an image decoding device according to one aspect of this disclosure may include the following steps: obtaining size information from a bitstream indicating the size of a current slice corresponding to at least a portion of a current frame; and determining the size of the current slice based on the size information. Here, the size information may include width information indicating the width of the current slice in units of tile columns and height information indicating the height of the current slice in units of tile rows, and the step of obtaining the size information from the bitstream may be performed based on whether the current slice belongs to the last tile column or the last tile row of the current frame.

[0014] Additionally, an image decoding device according to one aspect of this disclosure may include a memory and at least one processor. The at least one processor may acquire size information from a bitstream indicating the size of a current slice corresponding to at least a portion of the current frame; and determine the size of the current slice based on the size information. Here, the size information may include width information indicating the width of the current slice in units of tile columns and height information indicating the height of the current slice in units of tile rows, and may be acquired based on whether the current slice belongs to the last tile column or the last tile row of the current frame.

[0015] An image encoding method performed by an image encoding device according to another aspect of this disclosure may include the following steps: determining a current slice corresponding to at least a portion of a current frame; and generating a bitstream including size information of the current slice. Here, the size information may include width information indicating the width of the current slice in units of tile columns and height information indicating the height of the current slice in units of tile rows, and the step of generating the bitstream including the size information of the current slice may be performed based on whether the current slice belongs to the last tile column or the last tile row of the current frame.

[0016] Alternatively, according to another aspect of the transmission method of this disclosure, a bit stream generated by the image encoding device or image encoding method of this disclosure can be transmitted.

[0017] Additionally, according to another aspect of this disclosure, a computer-readable recording medium can store a bitstream generated by the image encoding device or image encoding method of this disclosure.

[0018] The features briefly outlined above are merely exemplary aspects of the following detailed description of this disclosure and do not limit the scope of this disclosure.

[0019] Beneficial effects

[0020] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.

[0021] Furthermore, according to this disclosure, an image encoding / decoding method and apparatus can be provided to improve encoding / decoding efficiency by selectively encoding the size information of slices.

[0022] Furthermore, according to this disclosure, a method for transmitting a bitstream generated by an image encoding method or device according to this disclosure can be provided.

[0023] Furthermore, according to this disclosure, a recording medium storing a bitstream generated by an image encoding method or apparatus according to this disclosure can be provided.

[0024] Furthermore, according to this disclosure, a recording medium may be provided that stores a bitstream that is received, decoded, and used to reconstruct an image by an image decoding device according to this disclosure.

[0025] Those skilled in the art will understand that the effects achievable through this disclosure are not limited to those specifically described above, and that other advantages of this disclosure will become clearer from the detailed description. Attached Figure Description

[0026] Figure 1 This is a view schematically illustrating a video encoding system to which embodiments of this disclosure are applicable.

[0027] Figure 2 This is a schematic view illustrating an image encoding device to which embodiments of the present disclosure are applicable.

[0028] Figure 3 This is a schematic view illustrating an image decoding device to which embodiments of the present disclosure are applicable.

[0029] Figure 4 This is a view showing the segmentation structure of an image according to an embodiment.

[0030] Figure 5 This is a view illustrating an implementation of the block segmentation type based on a multi-type tree structure.

[0031] Figure 6 This is a view illustrating the signaling mechanism for block partitioning information in a quadtree with nested multi-type tree structures according to this disclosure.

[0032] Figure 7 This is a view illustrating an implementation of dividing a CTU into multiple CUs.

[0033] Figure 8 This is a view showing a neighboring reference sample according to an embodiment.

[0034] Figures 9 to 10 This is a view illustrating intra-frame prediction according to an implementation method.

[0035] Figure 11 This is a view illustrating a coding method using inter-frame prediction according to an implementation method.

[0036] Figure 12 This is a view illustrating a decoding method using inter-frame prediction according to an implementation method.

[0037] Figure 13 This is a block diagram of CABAC based on an implementation method for encoding a syntax element.

[0038] Figures 14 to 17 This is a view illustrating entropy encoding and entropy decoding according to an implementation method.

[0039] Figure 18 and Figure 19 This is a view illustrating an example of the image decoding and encoding process according to an implementation method.

[0040] Figure 20 This is a view showing the layer structure of an encoded image according to an embodiment.

[0041] Figures 21 to 24 This is a view illustrating implementations of dividing the screen using tiles, slices, and sub-screens.

[0042] Figure 25 This is a view showing an implementation of the syntax for the sequence parameter set.

[0043] Figure 26 This is a view showing an implementation of the syntax for the screen parameter set.

[0044] Figure 27 This is a view showing an implementation of the syntax for the slice header.

[0045] Figure 28 and Figure 29 This is a view illustrating implementation methods of the encoding and decoding methods.

[0046] Figure 30 and Figure 31 This is a view showing another implementation of the screen parameter set.

[0047] Figure 32 This is a view illustrating an implementation of the decoding method.

[0048] Figure 33 and Figure 34 This is a view showing the algorithm used to determine SliceTopLeftTileIdx.

[0049] Figure 35 This is a view illustrating an implementation of the encoding method.

[0050] Figure 36 This is a view illustrating the content streaming system to which embodiments of this disclosure are applicable. Detailed Implementation

[0051] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, this disclosure may be implemented in various different forms and is not limited to the embodiments described herein.

[0052] In describing this disclosure, detailed descriptions of relevant known functions or constructions will be omitted if they unnecessarily obscure the scope of this disclosure. In the accompanying drawings, portions irrelevant to the description of this disclosure are omitted, and similar reference numerals are assigned to similar portions.

[0053] In this disclosure, when a component is "connected," "linked," or "coupled" to another component, it may include not only a direct connection but also an indirect connection with intermediate components. Furthermore, when a component "comprises" or "has" other components, unless otherwise stated, it means that other components may be included, not excluded.

[0054] In this disclosure, the terms first, second, etc., may be used only for the purpose of distinguishing one component from other components and do not limit the order or importance of the components, unless otherwise stated. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.

[0055] In this disclosure, the distinguishing components are intended to clearly describe each feature and do not imply that the components must be separate. That is, multiple components may be integrated and implemented in a single hardware or software unit, or a single component may be distributed and implemented in multiple hardware or software units. Therefore, unless otherwise stated, these implementations, where components are integrated or distributed, are included within the scope of this disclosure.

[0056] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional. Therefore, embodiments consisting of a subset of the components described in the embodiments are also included within the scope of this disclosure. In addition, embodiments that include other components besides those described in the various embodiments are also included within the scope of this disclosure.

[0057] This disclosure relates to the encoding and decoding of images. Unless redefined in this disclosure, the terms used herein may have the general meaning commonly used in the art to which this disclosure pertains.

[0058] In this disclosure, "video" can refer to a set of images over time. A frame typically refers to a unit representing an image at a specific time, while a tile is a coding unit that constitutes part of a frame during the encoding process. A frame can consist of one or more tiles. Additionally, a tile can include one or more Coded Tree Units (CTUs). A frame can consist of one or more tiles. A frame can include one or more tile groups. A tile group can include one or more tiles. A brick can represent a rectangular area of ​​a CTU row within a tile in a frame. A tile can include one or more bricks. A brick can represent a rectangular area of ​​a CTU row within a tile. A tile can be divided into multiple bricks, and each brick can include one or more CTU rows belonging to a tile. A tile that is not divided into multiple bricks can also be considered a brick.

[0059] A "pixel" or "pixel" can refer to the smallest unit that makes up a picture (or image). Additionally, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, or it can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.

[0060] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance blocks (e.g., Cb and Cr). In some cases, "unit" may be used interchangeably with terms such as "sample array," "block," or "region." In general, an M×N block may include a set (or array) of samples (or transform coefficients) with M columns and N rows.

[0061] In this disclosure, "current block" can mean one of "current coding block," "current coding unit," "coding target block," "decoding target block," or "processing target block." When performing prediction, "current block" can mean "current prediction block" or "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block." When performing filtering, "current block" can mean "filter target block."

[0062] Additionally, in this disclosure, unless explicitly stated as a chroma block, "current block" may mean "the luminance block of the current block". "The chroma block of the current block" can be expressed by including an explicit description of a chroma block such as "chroma block" or "current chroma block".

[0063] In this disclosure, the forward slash " / " or "," should be interpreted as indicating "and / or". For example, the expressions "A / B" and "A, B" can mean "A and / or B". Furthermore, "A / B / C" and "A / B / C" can mean "at least one of A, B and / or C".

[0064] In this disclosure, the term "or" should be interpreted as indicating "and / or". For example, expressing "A or B" can include 1) only "A", 2) only "B", and / or 3) both "A and B". In other words, in this disclosure, the term "or" should be interpreted as indicating "additionally or alternatively".

[0065] Overview of Video Encoding Systems

[0066] Figure 1 This is a view showing a video encoding system according to this disclosure.

[0067] The video encoding system according to the embodiment may include a source device 10 and a receiving device 20. The source device 10 may deliver encoded video and / or image information or data to the receiving device 20 in the form of a file or stream via a digital storage medium or network.

[0068] The source device 10 according to an embodiment may include a video source generator 11, an encoding device 12, and a transmitter 13. The receiving device 20 according to an embodiment may include a receiver 21, a decoding device 22, and a renderer 23. The encoding device 12 may be referred to as a video / image encoding device, and the decoding device 22 may be referred to as a video / image decoding device. The transmitter 13 may be included in the encoding device 12. The receiver 21 may be included in the decoding device 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.

[0069] The video source generator 11 can acquire video / images through a process of capturing, compositing, or generating video / images. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc. In this case, the video / image capture process can be replaced by a process of generating related data.

[0070] The encoding device 12 can encode the input video / image. For compression and encoding efficiency, the encoding device 12 can perform a series of processes such as prediction, transformation, and quantization. The encoding device 12 can output the encoded data (encoded video / image information) in the form of a bitstream.

[0071] Transmitter 13 can transmit encoded video / image information or data, output in bitstream form, to receiver 21 of receiving device 20 in the form of a file or stream via digital storage medium or network. Digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 can include elements for generating media files according to a predetermined file format and may include elements for transmission via broadcast / communication networks. Receiver 21 can extract / receive bitstreams from storage medium or network and transmit the bitstreams to decoding device 22.

[0072] The decoding device 22 can perform decoding on video / images by executing a series of processes (such as dequantization, inverse transform, and prediction) corresponding to the operation of the encoding device 12.

[0073] Renderer 23 can render decoded video / images. The rendered video / images can be displayed on a monitor.

[0074] Overview of Image Encoding Devices

[0075] Figure 2 This is a schematic view illustrating an image encoding device to which embodiments of the present disclosure are applicable.

[0076] like Figure 2 As shown, the image source device 100 may include an image segmenter 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame predictor 180, an intra-frame predictor 185, and an entropy encoder 190. The inter-frame predictor 180 and the intra-frame predictor 185 may be collectively referred to as "predictors". The transformer 120, quantizer 130, dequantizer 140, and inverse transformer 150 may be included in a residual processor. The residual processor may also include a subtractor 115.

[0077] In some embodiments, all or at least some of the multiple components configuring the image source device 100 may be configured by a single hardware component (e.g., an encoder or a processor). Additionally, the memory 170 may include a decoded screen buffer (DPB) and may be configured by a digital storage medium.

[0078] Image segmenter 110 can segment an input image (or picture or frame) input to image source device 100 into one or more processing units. For example, a processing unit may be called a coding unit (CU). A coding unit can be obtained by recursively segmenting a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree / binary tree / tritree (QT / BT / TT) structure. For example, a coding unit can be segmented into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the segmentation of coding units, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. The encoding process according to this disclosure can be performed based on the final coding unit that is no longer segmented. The maximum coding unit can be used as the final coding unit, or a deeper coding unit obtained by segmenting the maximum coding unit can be used as the final coding unit. Here, the encoding process may include the prediction, transformation, and reconstruction processes described later. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). Prediction units and transform units can be partitioned or segmented from the final coding unit. Prediction units can be sample prediction units, and transform units can be units used to derive transform coefficients and / or units used to derive residual signals from transform coefficients.

[0079] The predictor (inter-frame predictor 180 or intra-frame predictor 185) can perform prediction on the block to be processed (the current block) and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The predictor can generate various information related to the prediction of the current block and send the generated information to the entropy encoder 190. The information about the prediction can be encoded in the entropy encoder 190 and output as a bitstream.

[0080] Intra-predictor 185 can predict the current block by referencing samples in the current frame. Depending on the intra-prediction mode and / or intra-prediction technique, the reference samples may be located among the neighbors of the current block or may be placed separately. Intra-prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, directional modes may include, for example, 33 or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-predictor 185 can determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.

[0081] Inter-frame predictor 180 can derive the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc. The reference frame including the temporally neighboring block may be referred to as a juxtaposed frame (colPic). For example, inter-frame predictor 180 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 180 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be signaled by encoding motion vector differences and indicators of the motion vector predictors. The motion vector difference can refer to the difference between the motion vector of the current block and the motion vector predictor.

[0082] The predictor can generate a prediction signal based on various prediction methods and techniques described below. For example, the predictor can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. A prediction method that simultaneously applies both intra-frame prediction and inter-frame prediction to predict the current block can be called Combined Intra-Inter-Frame Prediction (CIIP). Alternatively, the predictor can perform Intra-Frame Block Copy (IBC) to predict the current block. Intra-Frame Block Copy can be used for content image / video coding in games, for example, Screen Content Coding (SCC). IBC is a method of predicting the current frame using a previously reconstructed reference block in the current frame at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current frame can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction because the reference block is derived within the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.

[0083] The predicted signal generated by the predictor can be used to generate a reconstructed signal or a residual signal. Subtractor 115 generates a residual signal (residual block or residual sample array) by subtracting the predicted signal (predicted block or predicted sample array) output from the predictor from the input image signal (original block or original sample array). The generated residual signal can be sent to transformer 120.

[0084] Transformer 120 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size or to blocks of variable size instead of square.

[0085] Quantizer 130 can quantize the transform coefficients and send them to entropy encoder 190. Entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 130 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0086] The entropy encoder 190 can perform various encoding methods, such as exponential Columbus coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 190 can encode, together or separately, information required for video / image reconstruction other than quantization transform coefficients (e.g., values ​​of syntax elements). The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the Network Abstraction Layer (NAL) level. The video / image information may also include information about various parameter sets, such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. The signaled information, transmitted information, and / or syntax elements described in this disclosure can be encoded and included in the bitstream through the above encoding process.

[0087] The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be included as internal / external components of the image source device 100. Alternatively, a transmitter may be provided as a component of the entropy encoder 190.

[0088] The quantization transform coefficients output from quantizer 130 can be used to generate residual signals. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantization transform coefficients through dequantizer 140 and inverse transformer 150.

[0089] Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 180 or intra-frame predictor 185 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). If the block to be processed has no residual, such as in the case of applying a skip mode, the prediction block can be used as a reconstructed block. Adder 155 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.

[0090] Filter 160 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 170, specifically in the DPB of memory 170. Various filtering methods can include, for example, deblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc. Filter 160 can generate various filtering-related information and send the generated information to entropy encoder 190, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 190 and output as a bitstream.

[0091] The modified reconstructed frame sent to memory 170 can be used as a reference frame in inter-frame predictor 180. When inter-frame prediction is applied through image source device 100, prediction mismatch between image source device 100 and image decoding device can be avoided and coding efficiency can be improved.

[0092] The DPB of memory 170 can store modified reconstructed frames for use as reference frames in inter-frame predictor 180. Memory 170 can store motion information of blocks used to derive (or encode) motion information in the current frame and / or motion information of already reconstructed blocks in the frame. The stored motion information can be sent to inter-frame predictor 180 and used as motion information for spatially or temporally neighboring blocks. Memory 170 can store reconstructed samples of reconstructed blocks in the current frame and can transmit the reconstructed samples to intra-frame predictor 185.

[0093] Overview of image decoding devices

[0094] Figure 3 This is a schematic view illustrating an image decoding device to which embodiments of the present disclosure are applicable.

[0095] like Figure 3 As shown, the image receiving device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, and an intra-frame predictor 265. The inter-frame predictor 260 and the intra-frame predictor 265 may be collectively referred to as "predictors". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.

[0096] According to an embodiment, all or at least some of the multiple components configuring the image receiving device 200 may be configured by hardware components (e.g., a decoder or a processor). Additionally, the memory 250 may include a decoded screen buffer (DPB) or may be configured by a digital storage medium.

[0097] The image receiving device 200, which has already received a bitstream including video / image information, can perform operations related to... Figure 2 The image is reconstructed by processing corresponding to the processing performed by the image source device 100. For example, the image receiving device 200 can use a processing unit applied in an image encoding device to perform decoding. Therefore, the decoding processing unit can be, for example, an encoding unit. The encoding unit can be obtained by segmenting a coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image receiving device 200 can be reproduced by a reproduction device (not shown).

[0098] Image receiving device 200 can receive images in bitstream form from Figure 2The signal output by the image encoding device. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may also include general constraint information. The image decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements notified / received by signals described in this disclosure can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 210 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values ​​of the syntax elements required for image reconstruction and the quantized values ​​of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the target syntax element, decoding information of neighboring blocks and the target block, or information about symbols / bins decoded in the previous stage, perform arithmetic decoding on the bins based on the determined context model by predicting the occurrence probability of the bins, and generate symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The prediction-related information in the information decoded by the entropy decoder 210 can be provided to the predictors (inter-frame predictor 260 and intra-frame predictor 265), and the residual values ​​of the entropy decoding performed in the entropy decoder 210 (that is, the quantization transform coefficients and related parameter information) can be input to the dequantizer 220. In addition, the filtering information in the information decoded by the entropy decoder 210 can be provided to the filter 240. Furthermore, the receiver (not shown) for receiving signals output from the image encoding device may be further configured as an internal / external element of the image receiving device 200, or the receiver may be a component of the entropy decoder 210.

[0099] Furthermore, the image decoding apparatus according to this disclosure can be referred to as a video / image / screen decoding apparatus. The image decoding apparatus can be divided into an information decoder (video / image / screen information decoder) and a sample decoder (video / image / screen sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame predictor 260, or an intra-frame predictor 265.

[0100] Dequantizer 220 can dequantize the quantized transform coefficients and output transform coefficients. Dequantizer 220 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image encoding device. Dequantizer 220 can obtain transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).

[0101] The inverse transformer 230 can perform inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0102] The predictor can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from the entropy decoder 210, and can determine a specific intra-frame / inter-frame prediction mode (prediction technique).

[0103] Similar to the predictor described in the image source device 100, the predictor can generate a prediction signal based on various prediction methods (techniques) that will be described later.

[0104] Intra-predictor 265 can predict the current block by referring to samples in the current frame. The description of intra-predictor 185 also applies to intra-predictor 265.

[0105] Inter-frame predictor 260 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, inter-frame predictor 260 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.

[0106] Adder 235 generates a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 260 and / or intra-frame predictor 265). If the block to be processed has no residual, such as when a skip mode is applied, the prediction block can be used as a reconstruction block. The description of adder 155 also applies to adder 235. Adder 235 may be referred to as a reconstructor or reconstruction block generator. The generated reconstruction signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can be used for inter-frame prediction of the next frame by filtering as described below.

[0107] Filter 240 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 240 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 250, specifically in the DPB of memory 250. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.

[0108] The (modified) reconstructed frame stored in the DPB of memory 250 can be used as a reference frame in inter-frame predictor 260. Memory 250 can store motion information of blocks used to derive (or decode) motion information in the current frame and / or motion information of already reconstructed blocks in the frame. The stored motion information can be sent to inter-frame predictor 260 to be used as motion information for spatially or temporally neighboring blocks. Memory 250 can store reconstructed samples of reconstructed blocks in the current frame and transmit the reconstructed samples to intra-frame predictor 265.

[0109] In this disclosure, the embodiments described in the filter 160, inter-frame predictor 180 and intra-frame predictor 185 of the image source device 100 can be applied equally or correspondingly to the filter 240, inter-frame predictor 260 and intra-frame predictor 265 of the image receiving device 200.

[0110] Overview of Image Segmentation

[0111] The video / image coding method according to this disclosure can be performed based on the following image segmentation structure. Specifically, the processes of prediction, residual processing (inverse transform, dequantization, etc.), syntax element encoding, and filtering, which will be described later, can be performed based on the CTU, CU (and / or TU, PU) derived from the image segmentation structure. The image can be segmented into block units, and the block segmentation process can be performed in the image segmenter 110 of the encoding device. Segmentation-related information can be encoded by the entropy encoder 190 and sent to the decoding device in the form of a bitstream. The entropy decoder 210 of the decoding device can deduce the block segmentation structure of the current frame based on the segmentation-related information obtained from the bitstream, and based on this, a series of processes (e.g., prediction, residual processing, block / frame reconstruction, in-loop filtering, etc.) can be performed for image decoding.

[0112] The image can be segmented into a sequence of Code Tree Units (CTUs). Figure 4 An example of a screen being segmented into CTUs is shown. A CTU can correspond to a Coding Tree Block (CTB). Alternatively, a CTU can include a coded tree block for luma samples and two coded tree blocks for corresponding chroma samples. For example, for a screen containing three sample arrays, a CTU can include an N×N block of luma samples and two corresponding blocks of chroma samples. The maximum permissible size of the CTU used for encoding and prediction can differ from the maximum permissible size of the CTU used for transformation. For example, even if the maximum size of the luma transformation block is 64×64, the maximum permissible size of the luma block in the CTU can be 128×128.

[0113] Overview of CTU segmentation

[0114] As described above, coding units can be obtained by recursively partitioning coding tree units (CTUs) or maximum coding units (LCUs) according to a quadtree / binary tree / ternary tree (QT / BT / TT) structure. For example, a CTU can first be partitioned into a quadtree structure. Subsequently, the leaf nodes of the quadtree structure can be further partitioned using multiple tree types.

[0115] The quadtree partitioning means that the current CU (or CTU) is equally divided into four. By partitioning according to the quadtree, the current CU can be divided into four CUs with the same width and height. When the current CU is no longer partitioned into a quadtree structure, the current CU corresponds to a leaf node of the quadtree structure. The CUs corresponding to the leaf nodes of the quadtree structure may not be further partitioned and can be used as the final encoding units described above. Alternatively, the CUs corresponding to the leaf nodes of the quadtree structure can be further partitioned using multiple types of tree structures.

[0116] Figure 5This is a view illustrating an implementation of block segmentation types based on multiple tree structures. Segmentation based on multiple tree structures can include two types of partitioning based on binary tree structures and two types of partitioning based on ternary tree structures.

[0117] The two types of partitioning based on the binary tree structure can be vertical binary partitioning (SPLIT_BT_VER) and horizontal binary partitioning (SPLIT_BT_HOR). Vertical binary partitioning (SPLIT_BT_VER) means that the current CU is equally divided into two in the vertical direction. For example... Figure 4 As shown, a vertical binary partition can generate two CUs with the same height as the current CU and a width half the width of the current CU. A horizontal binary partition (SPLIT_BT_HOR) means that the current CU is equally divided into two in the horizontal direction. Figure 5 As shown, by using horizontal binary partitioning, two CUs can be generated with a height that is half the height of the current CU and a width that is the same as the current CU.

[0118] The two types of partitioning based on the ternary tree structure can include vertical ternary partitioning (SPLIT_TT_VER) and horizontal ternary partitioning (SPLIT_TT_HOR). In vertical ternary partitioning (SPLIT_TT_VER), the current CU is partitioned vertically in a 1:2:1 ratio. For example... Figure 5 As shown, a vertical truncation can generate two CUs with the same height as the current CU and a width one-quarter of the current CU's width, and one CU with the same height as the current CU and a width half of the current CU's width. In a horizontal truncation (SPLIT_TT_HOR), the current CU is divided horizontally in a 1:2:1 ratio. Figure 5 As shown, by dividing horizontally into three branches, two CUs with a height of 1 / 4 of the current CU's height and the same width as the current CU, and one CU with a height of half the current CU's height and the same width as the current CU, can be generated.

[0119] Figure 6 This is a view illustrating the signaling mechanism for block partitioning information in a quadtree with nested multi-type tree structures according to this disclosure.

[0120] Here, the CTU is considered the root node of the quadtree and is initially split into a quadtree structure. Signals (e.g., `qt_split_flag`) can be used to indicate whether a quadtree split should be performed on the current CU (the CTU or node (QT_node) of the quadtree). For example, when `qt_split_flag` has a first value (e.g., "1"), the current CU can be split into a quadtree. Alternatively, when `qt_split_flag` has a second value (e.g., "0"), the current CU is not split into a quadtree but becomes a leaf node (QT_leaf_node). Each quadtree leaf node can then be further split into a multi-type tree structure. That is, a leaf node of the quadtree can become a node of a multi-type tree (MTT_node). In the multi-type tree structure, a first flag (e.g., `Mtt_split_cu_flag`) can be used to indicate whether the current node is additionally split. If the corresponding node is additionally split (e.g., if the first flag is 1), a second flag (e.g., `Mtt_split_cu_vertical_flag`) can be signaled to indicate the split direction. For example, the split direction could be vertical when the second flag is 1, and horizontal when the second flag is 0. Then, a third flag (e.g., `Mtt_split_cu_binary_flag`) can be signaled to indicate whether the split type is binary or ternary. For example, the split type could be binary when the third flag is 1, and ternary when the third flag is 0. The nodes of the multi-type tree obtained through binary or ternary splits can be further split into multi-type tree structures. However, the nodes of the multi-type tree may not be split into quadtree structures. If the first flag is 0, the corresponding node of the multi-type tree is no longer split, but becomes a leaf node (`MTT_leaf_node`) of the multi-type tree. The CU corresponding to the leaf node of the multi-type tree can be used as the final encoding unit described above.

[0121] Based on `mtt_split_cu_vertical_flag` and `mtt_split_cu_binary_flag`, the multi-type tree partitioning mode (MttSplitMode) of the CU can be derived as shown in Table 1 below. In the following description, the multi-type tree partitioning mode may be referred to as the multi-tree partitioning type or partitioning type.

[0122] [Table 1]

[0123]

[0124] Figure 7 This is a view illustrating an example of partitioning a CTU into multiple CUs by applying a multi-type tree after applying a quadtree. Figure 7 In the diagram, bold border 710 represents quadtree segmentation, while remaining border 720 represents multi-type tree segmentation. A CU can correspond to a coded block (CB). In an implementation, a CU may include a coded block for luminance samples and two coded blocks for chrominance samples corresponding to the luminance samples. The size of the chrominance component (sample) CB or TB can be derived based on the component ratio of the color format (chrominance format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the image / picture based on the luminance component (sample) CB or TB size. In the case of a 4:4:4 color format, the chrominance component CB / TB size can be set to be equal to the luminance component CB / TB size. In the case of a 4:2:2 color format, the width of the chrominance component CB / TB can be set to half the width of the luminance component CB / TB, and the height of the chrominance component CB / TB can be set to the height of the luminance component CB / TB. In the 4:2:0 color format, the width of the chroma component CB / TB can be set to half the width of the luminance component CB / TB, and the height of the chroma component CB / TB can be set to half the height of the luminance component CB / TB.

[0125] In one implementation, when the size of the CTU is based on 128 luminance sample units, the size of the CU can range from 128×128 to 4×4, which is the same size as the CTU. In one implementation, in the case of a 4:2:0 color format (or chroma format), the chroma CB size can range from 64×64 to 2×2.

[0126] Furthermore, in implementations, the CU size and TU size can be the same. Alternatively, there can be multiple TUs in the CU region. The TU size can typically represent the size of the luminance component (sample) transform block (TB).

[0127] The TU size can be derived based on the maximum permissible TB size maxTbSize, which is a predetermined value. For example, when the CU size is greater than maxTbSize, multiple TUs (TBs) with maxTbSize can be derived from the CU, and transformations / inverse transformations can be performed in units of TUs (TBs). For example, the maximum permissible luminance TB size can be 64×64 and the maximum permissible chrominance TB size can be 32×32. If the width or height of the CB segmented according to the tree structure is greater than the maximum transformation width or height, the CB can be automatically (or implicitly) segmented until the TB size limits in the horizontal and vertical directions are met.

[0128] Additionally, for example, when applying intra-frame prediction, the intra-frame prediction mode / type can be derived on a CU (or CB) basis, and the neighbor reference sample derivation and prediction sample generation process can be performed on a TU (or TB) basis. In this case, there can be one or more TUs (or TBs) in a CU (or CB) region, and multiple TUs or (TBs) can share the same intra-frame prediction mode / type.

[0129] Furthermore, for quadtree coding schemes with nested multi-type trees, the following parameters can be signaled from the encoding device to the decoding device as SPS syntax elements. For example, at least one of the following can be signaled: CTU size (representing the size of the root node of the quadtree), MinQTSize (representing the minimum allowed size of the leaf node of the quadtree), MaxBtSize (representing the maximum allowed size of the root node of the binary tree), MaxTtSize (representing the maximum allowed size of the root node of the ternary tree), MaxMttDepth (representing the maximum allowed depth of the multi-type tree partitioning starting from the leaf node of the quadtree), MinBtSize (representing the minimum allowed size of the leaf node of the binary tree), or MinTtSize (representing the minimum allowed size of the leaf node of the ternary tree).

[0130] As an implementation using the 4:2:0 chroma format, the CTU size can be set to 128×128 luma blocks and two corresponding 64×64 chroma blocks. In this case, MinOTSize can be set to 16×16, MaxBtSize to 128×128, MaxTtSzie to 64×64, MinBtSize and MinTtSize to 4×4, and MaxMttDepth to 4. Quadtree partitioning can be applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can be called leaf QT nodes. The size of quadtree leaf nodes can range from 16×16 (e.g., MinOTSize) to 128×128 (e.g., CTU size). If a leaf QT node is 128×128, it can be partitioned into a binary / ternary tree without additional partitioning. This is because, in this case, even if partitioned, it would exceed MaxBtsize and MaxTtszie (e.g., 64×64). In other cases, leaf QT nodes can be further segmented into multi-type trees. Therefore, a leaf QT node is the root node of the multi-type tree, and a leaf QT node can have a multi-type tree depth (mttDepth) of 0. If the multi-type tree depth reaches MaxMttdepth (e.g., 4), further segmentation is not considered. If the width of a multi-type tree node is equal to MinBtSize and less than or equal to 2×MinTtSize, further horizontal segmentation is not considered. If the height of a multi-type tree node is equal to MinBtSize and less than or equal to 2×MinTtSize, further vertical segmentation is not considered. When segmentation is not considered, the encoding device can skip the signaling of segmentation information. In this case, the decoding device can deduce segmentation information with predetermined values.

[0131] Furthermore, a CTU can include a coded block of luma samples (hereinafter referred to as a "luma block") and two coded blocks of corresponding chroma samples (hereinafter referred to as "chroma blocks"). The above coding tree scheme can be applied equally or separately to the luma and chroma blocks of the current CU. Specifically, the luma and chroma blocks in a CTU can be partitioned into the same block tree structure, and in this case, the tree structure can be represented as SINGLE_TREE. Alternatively, the luma and chroma blocks in a CTU can be partitioned into separate block tree structures, and in this case, the tree structure can be represented as DUAL_TREE. That is, when the CTU is partitioned into a dual-tree, the block tree structure for the luma block and the block tree structure for the chroma block can exist separately. In this case, the block tree structure for the luma block can be called DUAL_TREE_LUMA, and the block tree structure for the chroma components can be called DUAL_TREE_CHROMA. For P and B slice / piece groups, the luma and chroma blocks in a CTU can be restricted to having the same coding tree structure. However, for I-slice / patch groups, luma blocks and chroma blocks can have separate block tree structures. If separate block tree structures are applied, the luma CTB can be segmented into CUs based on a specific coding tree structure, and the chroma CTB can be segmented into chroma CUs based on another coding tree structure. That is, this means that CUs in I-slice / patch groups with separate block tree structures can include either a coded block of the luma component or coded blocks of the two chroma components, and CUs in P or B-slice / patch groups can include blocks of three color components (one luma component and two chroma components).

[0132] Although quadtree coding tree structures with nested multi-type trees have been described, the structures for splitting CUs are not limited to this. For example, BT and TT structures can be interpreted as concepts included in a multi-split tree (MPT) structure, and CUs can be interpreted as being split via QT and MPT structures. In the example of splitting CUs via QT and MPT structures, the splitting structure can be determined by signaling syntax elements (e.g., MPT_split_type) that include information about how many blocks the leaf nodes of the QT structure are split into, and syntax elements (e.g., MPT_split_mode) that include information about which direction (vertical or horizontal) the leaf nodes of the QT structure are split into.

[0133] In another example, the CU can be segmented in a manner different from the QT, BT, or TT structures. That is, unlike the QT structure which segments a lower-depth CU into 1 / 4 of a higher-depth CU, the BT structure which segments a lower-depth CU into 1 / 2 of a higher-depth CU, or the TT structure which segments a lower-depth CU into 1 / 4 or 1 / 2 of a higher-depth CU, in some cases the lower-depth CU can be segmented into 1 / 5, 1 / 3, 3 / 8, 3 / 5, 2 / 3, or 5 / 8 of the higher-depth CU, and the method of segmenting the CU is not limited to these.

[0134] Quadtree coded block structures with multiple tree types can provide highly flexible block partitioning structures. Due to the partitioning types supported in the multiple tree types, different partitioning patterns can potentially produce the same coded block structure in some cases. In encoding and decoding devices, the amount of data for partitioning information can be reduced by limiting the occurrence of such redundant partitioning patterns.

[0135] Furthermore, in the encoding and decoding of video / images according to this disclosure, the image processing unit can have a hierarchical structure. A frame can be divided into one or more tiles, blocks, slices, and / or tile groups. A slice may include one or more blocks. A block may include one or more CTU rows in a tile. A slice may include an integer number of blocks in the frame. A tile group may include one or more tiles. A tile may include one or more CTUs. A CTU may be divided into one or more CUs. A tile may be a rectangular area comprising a specific tile row and a specific tile column composed of multiple CTUs in the frame. Based on the tile raster scan in the frame, a tile group may include an integer number of tiles. A slice header may carry information / parameters applicable to the corresponding slice (blocks within a slice). When the encoding or decoding device has a multi-core processor, the encoding / decoding processes for tiles, slices, blocks, and / or tile groups can be executed in parallel.

[0136] In this disclosure, the names or concepts of slice or tile group are used interchangeably. That is, a tile group header can be called a slice header. Here, a slice can have one of the slice types, including intra-frame (I) slices, prediction (P) slices, and double prediction (B) slices. For blocks in I slices, inter-frame prediction is not used for prediction, and only intra-frame prediction can be used. Of course, even in this case, it is possible to encode and signal the original sample values ​​without prediction. For blocks in P slices, either intra-frame prediction or inter-frame prediction can be used. When using inter-frame prediction, only single prediction can be used. Furthermore, for blocks in B slices, either intra-frame prediction or inter-frame prediction can be used. When using inter-frame prediction, at most double prediction can be used.

[0137] Encoding devices can determine tile / tile group, patch, slice, and maximum and minimum coding unit sizes based on the characteristics of the video image (e.g., resolution) or considering coding efficiency and parallel processing. Additionally, information about this, or information that can be deduced from this, can be included in the bitstream.

[0138] Decoding devices can acquire information indicating whether a CTU within a tile, or a tile / group of tiles, patch, or slice in the current frame, has been divided into multiple coding units. Encoding and decoding devices can improve encoding efficiency by signaling this information under specific conditions.

[0139] A slice header (slice header syntax) may include information / parameters that can be commonly applied to a slice. APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more frames. SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. DPS may include information / parameters associated with combinations of coded video sequences (CVS).

[0140] Additionally, information regarding the segmentation and configuration of tiles / tile groups / tiles / slices can be constructed at the encoding level using high-level syntax and sent to the decoding device as a bitstream.

[0141] Overview of Intra-Frame Prediction

[0142] In the following sections, intra-frame prediction performed by the aforementioned encoding and decoding devices will be described in more detail. Intra-frame prediction can be defined as a prediction used to generate a prediction sample for the current block based on reference samples in the frame to which the current block belongs (hereinafter referred to as the current frame).

[0143] Reference Figure 8 A description is provided. When intra-frame prediction is applied to the current block 801, neighboring reference samples to be used for intra-frame prediction of the current block 801 can be derived. The neighboring reference samples of the current block may include: a total of 2×nh samples, including sample 811 of size nW×nH adjacent to the left boundary of the current block and sample 812 adjacent to the lower left; a total of 2×nW samples, including sample 821 adjacent to the upper boundary of the current block and sample 822 adjacent to the upper right; and a sample 831 adjacent to the upper left of the current block. Alternatively, the neighboring reference samples of the current block may include multiple columns of upper neighboring samples and multiple rows of left neighboring samples.

[0144] In addition, the neighboring reference samples of the current block may include: a total of nH samples 841 of size nW×nH adjacent to the right boundary of the current block, a total of nW samples 851 adjacent to the bottom boundary of the current block, and a sample 842 adjacent to the lower right of the current block.

[0145] However, some neighboring reference samples of the current block may not have been decoded or may be unavailable. In this case, the decoding device can construct neighboring reference samples to be used for prediction by replacing unavailable samples with available samples. Alternatively, neighboring reference samples to be used for prediction can be constructed by interpolation of available samples.

[0146] When deriving neighboring reference samples, (i) the predicted sample can be derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the predicted sample can be derived based on reference samples in the neighboring reference samples of the current block that exist in a specific (prediction) direction relative to the predicted sample. Case (i) can be called non-directional mode or non-angular mode, while case (ii) can be called directional mode or angular mode. Alternatively, a predicted sample can be generated based on the predicted sample of the current block in the neighboring reference samples by interpolation between a first neighboring sample and a second neighboring sample located in a direction opposite to the prediction direction of the intra-prediction mode of the current block. This case can be called Linear Interpolation Intra-Prediction (LIP). Alternatively, a chromaticity predicted sample can be generated based on a luminance sample using a linear model. This case can be called LM mode. Additionally, a temporary predicted sample of the current block can be derived based on filtered neighboring reference samples, and the predicted sample of the current block can be derived by weighted summing the temporary predicted sample with at least one reference sample derived according to the intra-prediction mode from the existing neighboring reference samples (i.e., unfiltered neighboring reference samples). This case can be called Position-Related Intra-Prediction (PDPC). Alternatively, the reference sample line with the highest prediction accuracy can be selected from multiple neighboring reference sample lines of the current block, and the predicted sample can be derived using reference samples in the prediction direction of the corresponding line. In this case, intra-frame prediction coding can be performed by indicating (signaling) the reference sample line used to the decoding device. This can be called multi-reference line (MRL) intra-frame prediction or MRL-based intra-frame prediction. Furthermore, the current block can be divided into vertical or horizontal sub-partitions, intra-frame prediction can be performed based on the same intra-frame prediction mode, and neighboring reference samples can be derived and used on a sub-partition basis. That is, in this case, the intra-frame prediction mode of the current block is applied equally to the sub-partitions, and neighboring reference samples can be derived and used on a sub-partition basis, thereby increasing intra-frame prediction performance in some cases. This prediction method can be called intra-fractional (ISP) or ISP-based intra-frame prediction. This intra-frame prediction method can be called an intra-frame prediction type to distinguish it from intra-frame prediction modes (e.g., DC mode, planar mode, or directional mode). Intra-frame prediction types can be referred to by various terms, such as intra-frame prediction techniques or additional intra-frame prediction modes. For example, intra-prediction types (or additional intra-prediction modes, etc.) may include at least one of LIP, PDPC, MRL, and ISP mentioned above. A general intra-prediction method that excludes specific intra-prediction types such as LIP, PDPC, MRL, and ISP can be called a general intra-prediction type. A general intra-prediction type may refer to a case where a specific intra-prediction type is not applied, and prediction can be performed based on the aforementioned intra-prediction modes. Furthermore, post-filtering can be performed on the derived prediction samples as needed.

[0147] Specifically, the intra-frame prediction process may include an intra-frame prediction mode / type determination step, a neighboring reference sample derivation step, and a prediction sample derivation step based on the intra-frame prediction mode / type. Additionally, post-filtering may be performed on the derived prediction samples as needed.

[0148] In addition to the intra-prediction types mentioned above, affine linear weighted intra-prediction (ALWIP) can be used. ALWIP can be called linear weighted intra-prediction (LWIP), matrix weighted intra-prediction (MIP), or matrix-based intra-prediction. When MIP is applied to the current block, i) neighboring reference samples after averaging, ii) a matrix-vector multiplication process can be performed, and iii) horizontal / vertical interpolation can be further performed as needed to derive the predicted samples for the current block. The intra-prediction mode used for MIP can be constructed differently from the intra-prediction modes used in LIP, PDPC, MRL, ISP intra-prediction, or ordinary intra-prediction. The intra-prediction mode of MIP can be called MIP intra-prediction mode, MIP prediction mode, or MIP mode. For example, the matrix and offset used in matrix-vector multiplication can be set differently depending on the intra-prediction mode of MIP. Here, the matrix can be called the (MIP) weight matrix, and the offset can be called the (MIP) offset vector or (MIP) bias vector. The detailed MIP method will be described below.

[0149] The block reconstruction process based on intra-prediction and intra-prediction units in the coding device can schematically include, for example, the following: S910 can be executed by the intra-predictor 185 of the coding device, and S920 can be executed by the residual processor of the coding device, which includes at least one of a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, and an inverse transformer 150. Specifically, S920 can be executed by the subtractor 115 of the coding device. In S930, prediction information can be derived by the intra-predictor 185 and encoded by the entropy encoder 190. In S930, residual information can be derived by the residual processor and encoded by the entropy encoder 190. The residual information is information about the residual samples. The residual information may include information about the quantized transform coefficients of the residual samples. As described above, the residual samples can be derived into transform coefficients by the transformer 120 of the coding device, and the transform coefficients can be derived into quantized transform coefficients by the quantizer 130. Information about the quantization transform coefficients can be encoded by the entropy encoder 190 through a residual encoding process.

[0150] The encoding device can perform intra-prediction for the current block (S910). The encoding device can derive the intra-prediction mode / type of the current block, derive neighboring reference samples of the current block, and generate prediction samples in the current block based on the intra-prediction mode / type and neighboring reference samples. Here, the processes for determining the intra-prediction mode / type, deriving neighboring reference samples, and generating prediction samples can be performed simultaneously, or any one process can be performed before the other. For example, although not shown, the intra-predictor 185 of the encoding device may include an intra-prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The intra-prediction mode / type determination unit can determine the intra-prediction mode / type of the current block, the reference sample derivation unit can derive neighboring reference samples of the current block, and the prediction sample derivation unit can derive prediction samples of the current block. Furthermore, when performing the following prediction sample filtering process, the intra-predictor 185 may also include a prediction sample filter. The encoding device can determine the mode / type applied to the current block from a plurality of intra-prediction modes / types. The encoding device can compare the RD costs of intra-prediction modes / types and determine the best intra-prediction mode / type for the current block.

[0151] In addition, the encoding device can perform a prediction sample filtering process. Prediction sample filtering can also be called post-filtering. Some or all of the prediction samples can be filtered through the prediction sample filtering process. In some cases, the prediction sample filtering process can be omitted.

[0152] The encoding device can generate residual samples for the current block based on the (filtered) prediction samples (S920). The encoding device can derive residual samples by comparing the prediction samples with the original samples of the current block based on phase comparison.

[0153] The encoding device can encode image information including information about intra-frame prediction (prediction information) and residual information of residual samples (S930). The prediction information may include intra-frame prediction mode information and intra-frame prediction type information. The encoding device can output the encoded image information in the form of a bitstream. The output bitstream can be sent to the decoding device via a storage medium or a network.

[0154] The residual information may include the following residual coding syntax. The coding device can transform / quantize the residual samples to derive quantization transform coefficients. The residual information may include information about the quantization transform coefficients.

[0155] Furthermore, as described above, the encoding device can generate a reconstructed frame (including reconstructed samples and reconstructed blocks). To this end, the encoding device can perform dequantization / inverse transform again on the quantization transform coefficients to derive (modified) residual samples. The residual samples are transformed / quantized and then dequantized / inverse transformed to derive the same residual samples as those derived in the decoding device as described above. The encoding device can generate a reconstructed block including reconstructed samples of the current block based on the predicted samples and the (modified) residual samples. A reconstructed frame of the current frame can be generated based on the reconstructed blocks. As described above, the in-loop filtering process is further applied to the reconstructed frame.

[0156] For example, a video / image decoding process based on intra-frame prediction and an intra-frame predictor in a decoding device can schematically include the following. The decoding device can perform operations corresponding to those performed in the encoding device.

[0157] S1010 to S1030 can be executed by the intra-frame predictor 265 of the decoding device, and the prediction information of S1010 and the residual information of S1040 can be obtained from the bitstream by the entropy decoder 210 of the decoding device. The residual processor of the decoding device, including at least one of a dequantizer 220 and an inverse transformer 230, can derive the residual samples of the current block based on the residual information. Specifically, the dequantizer 220 of the residual processor can perform dequantization based on the quantization transform coefficients derived from the residual information to derive the transform coefficients, and the dequantizer 220 of the residual processor can perform an inverse transform on the transform coefficients to derive the residual samples of the current block. S1050 can be executed by the adder 235 or the reconstructor of the decoding device.

[0158] Specifically, the decoding device can deduce the intra-prediction mode / type of the current block based on the received prediction information (intra-prediction mode / type information) (S1010). The decoding device can deduce the neighboring reference samples of the current block (S1020). The decoding device can generate prediction samples in the current block based on the intra-prediction mode / type and the neighboring reference samples (S1030). In this case, the encoding device can perform a prediction sample filtering process. Prediction sample filtering can be called post-filtering. Some or all of the prediction samples can be filtered by the prediction sample filtering process. In some cases, the prediction sample filtering process can be omitted.

[0159] The decoding device can generate residual samples for the current block based on the received residual information. The decoding device can generate reconstructed samples for the current block based on the predicted samples and residual samples, and derive reconstructed samples including the reconstructed samples (S1040). A reconstructed image of the current image can be generated based on the reconstructed block. As described above, the in-loop filtering process is further applied to the reconstructed image.

[0160] Here, although not shown, the intra-predictor 265 of the decoding device may include an intra-prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The intra-prediction mode / type determination unit can determine the intra-prediction mode / type of the current block based on the intra-prediction mode / type information acquired by the entropy decoder 210. The reference sample derivation unit can derive the neighboring reference samples of the current block, and the prediction sample derivation unit can derive the prediction samples of the current block. Furthermore, when performing the above-described prediction sample filtering process, the intra-predictor 265 may also include a prediction sample filter.

[0161] Intra-luma_mpm_flag may include flag information indicating whether the most probable mode (MPM) or a remaining mode is applied to the current block. When the MPM is applied to the current block, the prediction mode information may also include index information (e.g., intra_luma_mpm_idx) indicating one of the intra-luma_mpm_candidate prediction modes (MPM candidates). Intra-luma_mpm_candidate prediction modes (MPM candidates) can be configured as an MPM candidate list or an MPM list. Additionally, when the MPM is not applied to the current block, the intra-luma_mpm_remainder may also include remaining mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra-luma_prediction modes besides the MPM candidates. The decoding device can determine the intra-prediction mode for the current block based on the intra-prediction mode information. A separate MPM list can be configured for the above MIP.

[0162] Furthermore, intra-prediction type information can be implemented in various forms. For example, intra-prediction type information may include intra-prediction type index information indicating one of the intra-prediction types. As another example, intra-prediction type information may include reference sample line information (e.g., intra_luma_ref_idx) indicating whether MRL is applied to the current block and, if so, which reference sample line to use; ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether ISP is applied to the current block; ISP type information (e.g., intra_subpartitions_split_flag) indicating the partitioning type of the subpartitions when ISP is applied; flag information indicating whether PDCP is applied; or flag information indicating whether LIP is applied. Additionally, intra-prediction type information may include a MIP flag indicating whether MIP is applied to the current block.

[0163] Intra-prediction mode information and / or intra-prediction type information can be encoded / decoded using the encoding methods described in this disclosure. For example, intra-prediction mode information and / or intra-prediction type information can be encoded / decoded based on truncated (Rice) binary code using entropy coding (e.g., CABAC or CAVLC).

[0164] Overview of inter-frame prediction

[0165] The following text will describe the reference. Figure 2 and Figure 3 The description of encoding and decoding details the techniques for inter-frame prediction. In the case of a decoding device, the video / image decoding method and inter-frame predictor based on inter-frame prediction in the decoding device can operate according to the following description. In the case of an encoding device, the video / image encoding method and inter-frame predictor based on inter-frame prediction in the encoding device can operate according to the following description. Furthermore, in the following description, the data encoded by the following description can be stored in the form of a bitstream.

[0166] The predictor of an encoding / decoding device can perform inter-frame prediction on a block-by-block basis to derive prediction samples. Inter-frame prediction can represent a prediction derived by relying on data elements (e.g., sample values, motion information, etc.) of frames other than the current frame. When applying inter-frame prediction to the current block, the prediction block (prediction sample array) of the current block can be derived based on a reference block (reference sample array) specified by a motion vector on a reference frame indicated by a reference frame index. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, the motion information of the current block can be predicted on a block, sub-block, or sample-by-sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information can include motion vectors and reference frame indices. Motion information can also include inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.) information. When applying inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current frame and temporally neighboring blocks in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block can be the same or different. A temporally neighboring block can be referred to as a collated reference block, collated CU, or colCU, and a reference frame including the temporally neighboring block can be referred to as a collated frame (colPic). For example, a candidate list of motion information can be configured based on the neighboring blocks of the current block, and a flag or index information indicating which candidate to select (use) to derive the motion vector of the current block and / or the reference frame index can be signaled. Inter-frame prediction can be performed based on various prediction modes, and for example, in skip mode and merge mode, the motion information of the current block can be equal to the motion information of the selected neighboring block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, the motion vector of the selected neighboring block can be used as a motion vector predictor, and the motion vector difference can be signaled. In this case, the motion vector of the current block can be derived using the sum of the motion vector predictor and the motion vector difference.

[0167] Depending on the inter-frame prediction type (L0 prediction, L1 prediction, Bi prediction, etc.), motion information can include L0 motion information and / or L1 motion information. A motion vector in the L0 direction can be called an L0 motion vector or MVL0, and a motion vector in the L1 direction can be called an L1 motion vector or MVL1. Prediction based on the L0 motion vector is called L0 prediction, prediction based on the L1 motion vector is called L1 prediction, and prediction based on both L0 and L1 motion vectors is called Bi prediction. Here, the L0 motion vector can represent the motion vector associated with a reference frame list L0 (L0), and the L1 motion vector can represent the motion vector associated with a reference frame list L1 (L1). The reference frame list L0 can include frames preceding the current frame in output order as reference frames, and the reference frame list L1 can include frames following the current frame in output order. Frames preceding the current block can be called forward (reference) frames, and frames following the current block can be called backward (reference) frames. The reference frame list L0 can also include frames following the current frame in output order as reference frames. In this scenario, in reference screen list L0, screens preceding the current block are indexed first, followed by screens following the current block. Reference screen list L1 can also include screens preceding the current screen in the output order as reference screens. In this case, in reference screen list L1, screens following the current block can be indexed first, followed by screens preceding the current block. Here, the output order can correspond to the screen order count (POC) order.

[0168] For example, a video / image coding process based on inter-frame prediction and inter-frame predictors in a coding device can be illustrated as follows. (Refer to...) Figure 11A description is provided. The encoding device performs inter-frame prediction for the current block (S1110). The encoding device can deduce the inter-frame prediction mode and motion information of the current block, and generate prediction samples for the current block. Here, the inter-frame prediction mode determination, motion information derivation, and prediction sample derivation processes can be performed simultaneously, and any process can be performed before another process. For example, the inter-frame predictor of the encoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit can determine the prediction mode of the current block, the motion information derivation unit can deduce the motion information of the current block, and the prediction sample derivation unit can deduce the prediction samples of the current block. For example, the inter-frame predictor of the encoding device can search for blocks similar to the current block in a specific region (search region) of a reference frame through motion estimation, and deduce the reference block whose difference from the current block is the smallest or less than or equal to a specific criterion. Based on this, a reference frame index indicating the reference frame where the reference block is located can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode applied to the current block among various prediction modes. The encoding device can compare the RD costs of various prediction modes and determine the best prediction mode for the current block.

[0169] For example, when a skip mode or merge mode is applied to the current block, the encoding device can construct a merge candidate list (described below) and deduce a reference block among those reference blocks indicated by the merge candidates included in the merge candidate list whose difference from the current block is the smallest or less than or equal to a specific criterion. In this case, a merge candidate associated with the deduced reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be deduced using the motion information of the selected merge candidate.

[0170] As another example, when the (A)MVP mode is applied to the current block, the encoding device can construct an (A)MVP candidate list, as described below, and use motion vector predictor (MVP) candidates selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, the motion vector of the reference block derived through motion estimation can be used as the motion vector of the current block, and the MVP candidate with the motion vector having the smallest difference from the motion vector of the current block can be the selected MVP candidate. The motion vector difference (MVD) obtained by subtracting the MVP from the motion vector of the current block can be derived. In this case, information about the MVD can be signaled to the decoding device. Additionally, when the (A)MVP mode is applied, the value of the reference frame index can be constructed as reference frame index information and signaled to the decoding device.

[0171] The encoding device can derive residual samples based on the predicted samples (S1120). The encoding device can derive residual samples by comparing the original samples of the current block with the predicted samples.

[0172] The encoding device encodes image information including prediction information and residual information (S1130). The encoding device can output the encoded image information in the form of a bitstream. The prediction information may include prediction mode information (e.g., skip flag, merge flag, mode index, etc.) and motion information as information related to the prediction process. The motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index) as information for deriving motion vectors. In addition, the motion information may include information about the aforementioned MVD and / or reference frame index information. Furthermore, the motion information may include information indicating whether L0 prediction, L1 prediction, or bi prediction is applied. The residual information is information about the residual samples. The residual information may include information about the quantization transform coefficients of the residual samples.

[0173] The output bitstream can be stored in a (digital) storage medium and sent to the decoding device, or it can be sent to the decoding device via a network.

[0174] Furthermore, as mentioned above, the encoding device can generate reconstructed frames (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This allows the decoding device to derive the same prediction results as those performed in the encoding device, thereby improving encoding efficiency. Therefore, the encoding device can store the reconstructed frames (or reconstructed samples or reconstructed blocks) in memory and use them as reference frames for inter-frame prediction. As mentioned above, the in-loop filtering process can be further applied to the reconstructed frames.

[0175] For example, the video / image decoding process based on inter-frame prediction and inter-frame predictor in a decoding device may schematically include the following.

[0176] The decoding device can perform operations corresponding to those performed by the encoding device. The decoding device can perform predictions on the current block and derive prediction samples based on the received prediction information.

[0177] Specifically, the decoding device can determine the prediction mode of the current block based on the received prediction information (S1210). The decoding device can determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.

[0178] For example, whether the merge mode or (A)MVP mode is applied to the current block can be determined based on the merge flag. Alternatively, one of a variety of inter-frame prediction mode candidates can be selected based on the mode index. Inter-frame prediction mode candidates may include skip mode, merge mode, and / or (A)MVP mode, or may include a variety of inter-frame prediction modes as described below.

[0179] The decoding device derives motion information for the current block based on a determined inter-frame prediction mode (S1220). For example, when a skip mode or merge mode is applied to the current block, the decoding device can construct a merge candidate list and select one of the merge candidates included in the list. The selection can be performed based on the selection information (merge index). The motion information of the selected merge candidate can be used to derive motion information for the current block. The motion information of the selected merge candidate can be used as motion information for the current block.

[0180] As another example, when the (A)MVP mode is applied to the current block, the decoding device can construct an (A)MVP candidate list, as described below, and use motion vector predictor (MVP) candidates selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. Selection can be performed based on the aforementioned selection information (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on information about the MVD, and the motion vector of the current block can be derived based on the MVD and MVP of the current block. Additionally, the reference frame index of the current block can be derived based on reference frame index information. The frame indicated by the reference frame index in the reference frame list of the current block can be derived as the reference frame for inter-frame prediction reference of the current block.

[0181] Furthermore, as mentioned above, the motion information of the current block can be derived without constructing a candidate list. In this case, the motion information of the current block can be derived according to the process described in the prediction mode below. In this case, the construction of the candidate list as described above can be omitted.

[0182] The decoding device can generate a prediction sample for the current block based on the motion information of the current block (S1230). In this case, a reference frame can be derived based on the reference frame index of the current block, and the prediction sample for the current block can be derived using samples of the reference block indicated by the motion vector of the current block on the reference frame. In this case, as described above, in some cases, a prediction sample filtering process for all or some of the prediction samples of the current block can be further performed.

[0183] For example, the inter-frame predictor of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit. The prediction mode determination unit can determine the prediction mode of the current block based on the received prediction mode information. The motion information derivation unit can derive the motion information (motion vector and / or reference frame index) of the current block based on the received motion information. The prediction sample derivation unit can derive the prediction sample of the current block.

[0184] The decoding device generates a residual sample for the current block based on the received residual information (S1240). The decoding device can generate a reconstructed sample for the current block based on the predicted sample and the residual sample, and generate a reconstructed image based on this (S1250). As described above, the in-loop filtering process can be further applied to the reconstructed image.

[0185] As described above, the inter-frame prediction process may include the steps of determining an inter-frame prediction mode, deriving motion information based on the determined prediction mode, and performing prediction (prediction sample generation) based on the derived motion information. As described above, the inter-frame prediction process can be performed in both the encoding and decoding devices.

[0186] Quantization / Dequantization

[0187] As mentioned above, the quantizer of the encoding device can derive the quantized transform coefficients by applying quantization to the transform coefficients, and the dequantizer of the encoding device or the dequantizer of the decoding device can derive the transform coefficients by applying dequantization to the quantized transform coefficients.

[0188] In encoding and decoding moving / still images, the quantization ratio can be changed, and the compression ratio can be adjusted using the changed quantization ratio. From an implementation perspective, considering complexity, instead of directly using the quantization ratio, a quantization parameter (QP) can be used. For example, a quantization parameter with integer values ​​from 0 to 63 can be used, and each quantization parameter value can correspond to the actual quantization ratio. Furthermore, the quantization parameter QP for the luminance component (luminance sample) can be set differently. Y Quantization parameter QP of chromaticity components (chromaticity samples) C .

[0189] During quantization, the transform coefficients C can be received and divided by the quantization ratio Qstep to obtain the quantization transform. In this case, considering computational complexity, the quantization ratio can be multiplied by a scale to become an integer, and a shift operation can be performed according to the value corresponding to the scale. Based on the product of the quantization ratio and the scale value, the quantization scale can be derived. That is, the quantization scale can be derived from QP. By applying the quantization scale to the transform coefficients C, the quantization transform coefficients C' can be derived.

[0190] Dequantization is the inverse of quantization. The reconstructed transform coefficients C' can be obtained by multiplying the quantization transform coefficients C' by the quantization ratio Qstep. Furthermore, a level scale can be derived from the quantization parameters, and this level scale can be applied to the quantization transform coefficients C' to derive the reconstructed transform coefficients C'. Due to losses during the transform and / or quantization processes, the reconstructed transform coefficients C' may differ slightly from the initial transform coefficients C. Therefore, dequantization can be performed in the same manner as in the decoding device, even within the encoding device.

[0191] Furthermore, adaptive frequency-weighted quantization (IFQ) can be applied, where the quantization intensity is adjusted according to frequency. IFQ refers to a method that applies different quantization intensities based on frequency. In IFQ, a predefined quantization scaling matrix can be used to apply different quantization intensities based on frequency. That is, the quantization / dequantization process described above can be performed based on the quantization scaling matrix. For example, different quantization scaling matrices can be used depending on the size of the current block and / or whether the prediction mode applied to the current block is inter-frame prediction or intra-frame prediction to generate the residual signal for the current block. The quantization scaling matrix can be called a quantization matrix or a scaling matrix. The quantization scaling matrix can be predefined. Additionally, for frequency-adaptive scaling, the frequency quantization scaling information used for the quantization scaling matrix can be constructed / encoded in the encoding device and signaled to the decoding device. The frequency quantization scaling information can be called quantization scaling information. The frequency quantization scaling information can include scaling list data (scaling_list_data). The quantization scaling matrix can be derived (modified) based on the scaling list data. Furthermore, the frequency quantization scaling information can include a presence flag indicating the existence of the scaling list data. Alternatively, when the scaling list data is signaled at a higher level (e.g., SPS), information indicating whether the scaling list data has been modified at a lower level (e.g., PPS or tile group header, etc.) may also be included.

[0192] Transform / Inverse Transform

[0193] As described above, the encoding device can derive residual blocks (residual samples) based on blocks predicted via intra / inter / IBC prediction (predicted blocks), and derive quantized transform coefficients by applying transform and quantization to the derived residual samples. Information about the quantized transform coefficients (residual information) can be included and encoded in the residual coding syntax and output as a bitstream. The decoding device can obtain and decode the information about the quantized transform coefficients (residual information) from the bitstream to derive the quantized transform coefficients. The decoding device can derive residual samples based on the quantized transform coefficients through dequantization / inverse transform. As described above, either quantization / dequantization or transform / inverse transform can be skipped. When a transform / inverse transform is skipped, the transform coefficients can be referred to as coefficients or residual coefficients; for consistency, they can still be referred to as transform coefficients. A transform skip flag (e.g., transform_skip_flag) can be used to signal whether a transform / inverse transform has been skipped.

[0194] Transformation / inverse transformation can be performed based on a transform kernel. For example, a multiple transform selection (MTS) scheme for performing transformation / inverse transformation can be applied. In this case, some from a set of multiple transform kernels can be selected and applied to the current block. Transformation kernels can be referred to by various terms such as transform matrix or transform type. For example, a transform kernel set can refer to a combination of vertical transform kernels (vertical transform kernels) and horizontal transform kernels (horizontal transform kernels).

[0195] Transform / inverse transforms can be performed on a CU or TU basis. That is, the transform / inverse transform can be applied to residual samples in a CU or a TU. The CU size can be equal to the TU size, or multiple TUs can exist within a CU region. Furthermore, the CU size typically indicates the size of the luma component (sample) CB. The TU size typically indicates the size of the luma component (sample) TB. The chroma component (sample) CB or TB size can be derived based on the luma component (sample) CB or TB size according to the component ratios of the color format (chroma format) (e.g., 4:4:4, 4:4:4, 4:2:2, 4:2:0, etc.). The TU size can be derived based on maxTbSize. For example, when the CU size is greater than maxTbSize, multiple TUs (TBs) with maxTbSize can be derived from the CU, and transform / inverse transforms can be performed on a TU (TB) basis. maxTbSize can be considered to determine whether various intra-frame prediction types (such as ISP) should be applied. Information about maxTbSize can be predetermined, or information about maxTbSize can be generated and encoded in the encoding device and then signaled to the encoding device.

[0196] Entropy coding

[0197] As referenced above Figure 2 As stated above, all or some of the video / image information can be entropy encoded by the entropy encoder 190, referencing... Figure 3 All or some of the described video / image information can be entropy decoded by entropy decoder 310. In this case, the video / image information can be encoded / decoded on a syntactic element basis. In this disclosure, encoding / decoding information can include encoding / decoding using the methods described in this paragraph.

[0198] Figure 13 This is a block diagram of CABAC used to encode a syntax element. In the CABAC encoding process, firstly, when the input signal is a syntax element rather than a binary value, it can be transformed into a binary value through binarization. When the input signal is already a binary value, binarization can be bypassed. Here, the binary numbers 0 or 1 that constitute the binary value can be called bins. For example, when the binarized binary string (bin string) is 110, each of 1, 1, and 0 can be called a bin. A bin of a syntax element can represent the value of the corresponding syntax element.

[0199] Binarized bins can be input into either a regular encoding engine or a bypass encoding engine. A regular encoding engine can assign a context model that reflects the probability values ​​to the corresponding bin and encode the bin based on the assigned context model. A regular encoding engine can encode each bin and then update the probability model for that bin. Bins encoded in this way are called context-coded bins. A bypass encoding engine can bypass the process of estimating the probabilities for the input bin and the process of updating the probability model applied to the corresponding bin after encoding. A bypass encoding engine can improve encoding speed by applying a uniform probability distribution (e.g., 50:50) instead of assigning context to encode the input bin. Bins encoded in this way are called bypass bins. A context model can be assigned and updated for each context-coded (regularly encoded) bin, and the context model can be indicated based on ctxidx or ctxInc. ctxidx can be derived based on ctxInc. Specifically, for example, the context index ctxidx, which indicates the context model of each regular coding bin, can be derived as the sum of the context index increment (ctxInc) and the context index offset (ctxIdxOffset). Here, ctxInc can be derived differently for each bin. ctxIdxOffset can be represented by the lowest value of ctxIdx. The lowest value of ctxIdx can be called the initial value initValue of ctxIdx. ctxIdxOffset is a value typically used to distinguish the context model from other syntax elements; the context model of a syntax element can be distinguished / derived based on ctxinc.

[0200] During entropy encoding, it can be determined whether encoding is performed using a regular encoding engine or a bypass encoding engine, and the encoding path can be switched. Entropy decoding can be performed in the reverse order of the same process as entropy encoding.

[0201] For example, the entropy coding described above can be as follows: Figure 14 and Figure 15 Execute it as described above. (Reference) Figure 14 and Figure 15 An encoding device (entropy encoder) can perform entropy coding on image / video information. This image / video information may include segmentation-related information, prediction-related information (e.g., inter-frame / intra-frame prediction difference information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, in-loop filtering-related information, or various related syntax elements. Entropy coding can be performed on a per-syntax-element basis. Figure 14 Steps S1410 to S1420 can be performed by Figure 2 The entropy encoder 190 of the encoding device is executed.

[0202] The encoding device can perform binarization (S1410) on the target syntax element. Here, the binarization can be based on various binarization methods such as truncated Rice binarization, fixed-length binarization, etc., and the binarization method for the target syntax element can be predefined. The binarization process can be performed by the binarization unit 191 in the entropy encoder 190.

[0203] The encoding device can perform entropy encoding on the target syntax element (S1420). The encoding device can perform encoding based on regular encoding (context-based) or bypass encoding on the bin string of the target syntax element using entropy encoding techniques such as CABAC (context-adaptive arithmetic coding) or CAVLC (context-adaptive variable-length coding), and its output can be included in the bitstream. The entropy encoding process can be executed by the entropy encoding processor 192 in the entropy encoder 190. As described above, the bitstream can be sent to the decoding device via a (digital) storage medium or a network.

[0204] refer to Figure 16 and Figure 17 The decoding device (entropy decoder) can decode the encoded image / video information. The image / video information may include segmentation-related information, prediction-related information (e.g., inter-frame / intra-frame prediction difference information, intra-frame prediction mode information, inter-frame prediction mode information, etc.), residual information, in-loop filtering-related information, etc., or may include various related syntax elements. Entropy coding can be performed on a syntax element-by-syntax basis. Steps S1610 to S1620 can be performed by... Figure 3 The entropy decoder 210 of the decoding device is executed.

[0205] The decoding device can perform binarization on the target syntax element (S1610). Here, binarization can be performed based on various binarization methods such as truncated Rice binarization or fixed-length binarization, and the binarization method for the target syntax element can be predefined. The decoding device can derive available bin strings (bin string candidates) for the available values ​​of the target syntax element through the binarization process. The binarization process can be performed by the binarization unit 211 in the entropy decoder 210.

[0206] The decoding device can perform entropy decoding (S1620) on the target syntax element. When decoding and parsing the bins of the target syntax element sequentially from the input bits in the bitstream, the decoding device can compare the derived bin string with the available bin strings for the corresponding syntax element. If the derived bin string equals one of the available bin strings, the value corresponding to the corresponding bin string can be derived as the value of the corresponding syntax element. If not, the above process can be repeated after further parsing the next bit in the bitstream. Through this process, variable-length bits are used to signal the corresponding information instead of the start or end bits of specific information (specific syntax element) in the bitstream. In this way, relatively few bits can be assigned low values, and overall encoding efficiency can be improved.

[0207] The decoding device can perform context-based or bypass-based decoding on bins in a bin string from a bitstream based on entropy coding techniques such as CABAC or CAVLC. The entropy decoding process can be executed by the entropy decoding processor 212 in the entropy decoder 210. The bitstream can include various information for image / video decoding as described above. As mentioned above, the bitstream can be sent to the decoding device via a (digital) storage medium or a network.

[0208] In this disclosure, a table including syntax elements (syntax table) can be used to indicate signaling of information from an encoding device to a decoding device. The order of the syntax elements in the table including the syntax elements used in this disclosure can indicate the order in which syntax elements are parsed from the bitstream. The encoding device can construct and encode the syntax table such that the decoding device parses the syntax elements in the parsing order, and the decoding device can parse and decode the syntax elements of the corresponding syntax table from the bitstream according to the parsing order, and obtain the values ​​of the syntax elements.

[0209] General image / video encoding process

[0210] In image / video coding, the frames that make up an image / video can be encoded / decoded according to the decoding order. The frame order corresponding to the output order of the decoded frames can be set to be different from the decoding order, and based on this, not only forward prediction but also backward prediction can be performed during inter-frame prediction.

[0211] Figure 18 An example of an illustrative screen decoding process to which embodiments of this disclosure apply is shown. Figure 18In this process, S1810 can be executed in the entropy decoder 210 of the decoding device; S1820 can be executed in the predictor including the intra-frame predictor 265 and the inter-frame predictor 260; S1830 can be executed in the residual processor including the dequantizer 220 and the inverse transformer 230; S1840 can be executed in the adder 235; and S1850 can be executed in the filter 240. S1810 can include the information decoding process described in this disclosure; S1820 can include the inter-frame / intra-frame prediction process described in this disclosure; S1830 can include the residual processing process described in this disclosure; S1840 can include the block / frame reconstruction process described in this disclosure; and S1850 can include the in-loop filtering process described in this disclosure.

[0212] refer to Figure 18 The image decoding process can schematically include a process of obtaining image / video information from the bitstream (through decoding) (S1810), an image reconstruction process (S1820 to S1840), and an in-loop filtering process for the reconstructed image (S1850). The image reconstruction process can be performed based on prediction samples and residual samples obtained through inter-frame / intra-frame prediction (S1820) and residual processing (S1830) (dequantization and inverse transform of quantization transform coefficients) described in this disclosure. For the reconstructed image generated by the image reconstruction process, a modified reconstructed image can be generated through the in-loop filtering process. The modified reconstructed image can be used as the decoded image output, stored in the decoded image buffer or memory 250 of the decoding device, and used as a reference image in the inter-frame prediction process when decoding the image later. In some cases, the in-loop filtering process can be omitted. In this case, the reconstructed image can be used as the decoded image output, stored in the decoded image buffer or memory 250 of the decoding device, and used as a reference image in the inter-frame prediction process when decoding the image later. The in-loop filtering process (S1850) may include a deblocking filtering process, a sample adaptive offset (SAO) process, an adaptive loop filter (ALF) process, and / or a bilateral filtering process, some or all of which may be omitted as described above. Furthermore, one or more of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and / or the bilateral filter process may be applied sequentially, or all of them may be applied sequentially. For example, the SAO process may be performed after the deblocking filtering process is applied to the reconstructed frame. Alternatively, for example, the ALF process may be performed after the deblocking filtering process is applied to the reconstructed frame. This can be performed similarly even in an encoding device.

[0213] Figure 19 An example of an illustrative screen encoding process to which embodiments of this disclosure apply is shown. Figure 19 In the above reference, S1910 can be found. Figure 2 The coding apparatus described herein includes a predictor comprising an intra-frame predictor 185 or an inter-frame predictor 180. S1920 may be executed in a residual processor comprising a transformer 120 and / or a quantizer 130, and S1930 may be executed in an entropy encoder 190. S1910 may include the inter-frame / intra-frame prediction process described herein, S1920 may include the residual processing process described herein, and S1930 may include the information encoding process described herein.

[0214] refer to Figure 19 The image encoding process can schematically include not only the process of encoding information used for image reconstruction (e.g., prediction information, residual information, segmentation information, etc.) and outputting that information as a bitstream, but also the process of generating a reconstructed image of the current image and (optionally) applying in-loop filtering to the reconstructed image, as shown in the following example. Figure 2 As described, the encoding device can derive (modified) residual samples from the quantized transform coefficients using dequantizer 140 and inverse transformer 150, and generate a reconstructed frame based on the predicted sample as the output of S1910 and the (modified) residual samples. The reconstructed frame generated in this way can be equal to the reconstructed frame generated in the decoding device. The modified reconstructed frame can be generated for the reconstructed frame through an in-loop filtering process, can be stored in the decoded frame buffer or memory 170, and can be used as a reference frame in the inter-frame prediction process when encoding frames later, similar to the case in the decoding device. As mentioned above, in some cases, some or all of the in-loop filtering process can be omitted. When the in-loop filtering process is performed, the (in-loop) filtering-related information (parameters) can be encoded in the entropy encoder 190 and output as a bitstream. The decoding device can perform the in-loop filtering process based on the filtering-related information using the same method as the encoding device.

[0215] This in-loop filtering process reduces noise generated during image / video encoding (such as block artifacts and ringing artifacts) and improves subjective / objective visual quality. Furthermore, by performing the in-loop filtering process in both the encoding and decoding devices, the encoding and decoding devices can derive the same prediction results, increasing the reliability of image encoding and reducing the amount of data transmitted for image encoding.

[0216] As described above, the image reconstruction process can be performed not only in the decoding device but also in the encoding device. Reconstructed blocks can be generated based on intra-frame prediction / inter-frame prediction on a block-by-block basis, and a reconstructed image including these blocks can be generated. When the current image / slice / patch group is an I-frame / slice / patch group, blocks included in the current image / slice / patch group can be reconstructed based solely on intra-frame prediction. Furthermore, when the current image / slice / patch group is a P-frame / slice / patch group or a B-frame / slice / patch group, blocks included in the current image / slice / patch group can be reconstructed based on either intra-frame prediction or inter-frame prediction. In this case, inter-frame prediction can be applied to some blocks in the current image / slice / patch group, and intra-frame prediction can be applied to the remaining blocks. The color components of the image can include luminance and chrominance components; unless explicitly defined in this disclosure, the methods and implementations of this disclosure can be applied to both luminance and chrominance components.

[0217] Examples of coding layers and structures

[0218] Video / images encoded according to this disclosure can be processed, for example, according to the encoding layers and structures described below.

[0219] Figure 20 This is a view showing the layer structure of an encoded image. Encoded images can be classified into the Video Coding Layer (VCL) for image decoding and its own processing, the lower system for transmitting and storing encoded information, and the Network Abstraction Layer (NAL) that exists between the VCL and the lower system and is responsible for network adaptation functions.

[0220] In VCL, VCL data that includes compressed image data (slice data) can be generated, or additional enhancement information (SEI) messages that include information such as picture parameter set (PPS), sequence parameter set (SPS), or video parameter set (VPS) can be generated for the decoding process of the image.

[0221] In NAL, header information (NAL cell header) can be added to the raw byte sequence payload (RBSP) generated in VCL to generate NAL cells. In this case, RBSP refers to the slice data, parameter set, and SEI message generated in VCL. The NAL cell header can include NAL cell type information specified based on the RBSP data included in the corresponding NAL cell.

[0222] As shown in the figure, NAL units can be classified into VCL NAL units and non-VCL NAL units based on the RBSP generated in the VCL. VCL NAL units can refer to NAL units that include information about the image (slice data), while non-VCL NAL units can refer to NAL units that include information needed to decode the image (parameter set or SEI message).

[0223] VCL NAL units and non-VCL NAL units can be accompanied by header information and transmitted over a network according to the data standard of the lower system. For example, NAL units can be modified to a predetermined standard data format such as H.266 / VVC file format, RTP (Real-Time Transport Protocol), or TS (Transport Streaming) and sent over various networks.

[0224] As described above, in a NAL unit, the NAL unit type can be specified according to the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and notified by a signal.

[0225] For example, based on whether the NAL unit includes information about the image (slice data), it can be roughly classified into VCLNAL unit type and non-VCL NAL unit type. VCL NAL unit type can be classified according to the characteristics and type of the image included in the VCL NAL unit, and non-VCL NAL unit type can be classified according to the type of parameter set.

[0226] Below are examples of NAL cell types specified based on the type of parameter set / information included in non-VCL NAL cell types.

[0227] -DCI (Decoding Capability Information) NAL Unit: Includes the type of NAL unit for DCI.

[0228] -VPS (Video Parameter Set) NAL Unit: Includes the type of NAL unit for the VPS.

[0229] -SPS (Sequence Parameter Set) NAL Unit: The type of NAL unit that includes SPS.

[0230] -PPS (Picture Parameter Set) NAL Unit: Includes the types of NAL units for PPS.

[0231] -APS (Adaptive Parameter Set) NAL Unit: The type of NAL unit including APS.

[0232] -PH (Picture Header) NAL Unit: Includes the type of NAL unit for PH.

[0233] The aforementioned NAL unit type can have syntax information specific to the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified as the nal_unit_type value.

[0234] Furthermore, as mentioned above, a frame can include multiple slices, and a slice can include a slice header and slice data. In this case, a frame header can be further added to multiple slices within a frame (slice header and slice data set). The frame header (frame header syntax) can include information / parameters that are typically applicable to the frame.

[0235] The slice header (slice header syntax) may include information / parameters commonly applicable to slices. APS (APS syntax) or PPS (PPS syntax) may include information / parameters commonly applicable to one or more slices or frames. SPS (SPS syntax) may include information / parameters commonly applicable to one or more sequences. VPS (VPS syntax) may include information / parameters commonly applicable to multiple layers. DCI (DCI syntax) may include information / parameters commonly applicable to the entire video. DCI may include information / parameters related to decoding capabilities. In this disclosure, the High-Level Syntax (HLS) may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, frame header syntax, or slice header syntax. Furthermore, in this disclosure, the Low-Level Syntax (LLS) may include, for example, slice data syntax, CTU syntax, coding unit syntax, transform unit syntax, etc.

[0236] In this disclosure, the image / video information encoded in the encoding device and signaled to the decoding device in the form of a bitstream can include not only intra-frame segmentation related information, intra / inter-frame prediction information, residual information, and intra-loop filtering information, but also information about slice headers, frame headers, APS, PPS, SPS, VPS, and / or DCI. Additionally, the image / video information may also include general constraint information and / or information about NAL unit headers.

[0237] Use sub-screens, slices, and tiles to divide the screen.

[0238] A frame can be divided into at least one tile row and at least one tile column. A tile can consist of a CTU sequence and can cover a rectangular area of ​​a frame.

[0239] A slice can consist of an integer number of complete tiles or an integer number of consecutive complete CTU rows in a frame.

[0240] For slicing, two modes are supported: one can be called raster scan slicing mode, and the other can be called rectangular slicing mode. In raster scan slicing mode, a slice can include a complete sequence of tiles existing in a frame according to the tile raster scan order. In rectangular slicing mode, a slice can include multiple complete tiles assembled to form a rectangular area of ​​a frame, or multiple consecutive complete CTU rows of a tile assembled to form a rectangular area of ​​a frame. Tiles in a rectangular slice can be scanned in the rectangular area corresponding to the slice according to the tile raster scan order. A sub-frame can include at least one slice that is assembled to cover a rectangular area of ​​the frame.

[0241] To describe the segmentation of the image in more detail, reference will be made to... Figures 21 to 24 Provide a description. Figures 21 to 24 This illustrates an implementation method for dividing a screen using tiles, slices, and sub-screens. Figure 21 An example of an image divided into 12 tiles and three raster scan slices is shown. Figure 22 This example shows a screen divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices. Figure 23 This shows an example of a screen divided into four tiles (two tile columns and two tile rows) and four rectangular slices.

[0242] Figure 24 This shows an example of dividing a screen into sub-screens. Figure 24 In this context, the image can be divided into 12 left tiles covering a slice consisting of 4×4 CTUs and 6 right tiles covering two vertically assembled slices consisting of 2×2 CTUs, resulting in an image divided into 24 slices and 24 sub-images with different regions. Figure 24 In the example, an individual slice corresponds to an individual sub-picture.

[0243] HLS (High-Level Syntax) Signaling and Semantics

[0244] As described above, the HLS can be encoded and / or signaled for video and / or image encoding. As described above, in this disclosure, video / image information can be included in the HLS. Furthermore, image / video encoding methods can be performed based on this image / video information.

[0245] Image header and slice header

[0246] An encoded frame can consist of at least one slice. Parameters describing the encoded frame can be signaled in the frame header (PH), or parameters describing the slice can be signaled in the slice header (SH). The PH can be sent as a NAL unit type. The SH can be provided at the beginning of the NAL unit configuring the slice's payload (e.g., slice data).

[0247] Screen splitting signaling

[0248] In one implementation, a frame can be divided into multiple sub-frames, tiles, and / or slices. Signaling for sub-frames can be provided in the sequence parameter set. Signaling for tiles and rectangular slices can be provided in the frame parameter set. Additionally, signaling for raster scan slices can be provided in the slice header.

[0249] Figure 25 This illustrates an implementation of the syntax for sequence parameter sets. Figure 25 In the syntax of , the syntax elements are as follows.

[0250] Syntax elements can indicate whether subpic information exists. For example, a first value of `subpic_info_present_flag` (e.g., 0) can indicate that subpic information for a Coding Layer Video Sequence (CLVS) does not exist in the bitstream, and that only one subpic exists in the individual frames of the CLVS. A second value of `subpic_info_present_flag` (e.g., 1) can indicate that subpic information for a Coding Layer Video Sequence (CLVS) exists in the bitstream, and that at least one subpic may exist in the individual frames of the CLVS.

[0251] Here, CLVS can refer to a layer that encodes a video sequence. A CLVS can be a sequence of PUs with the same nuh_layer_id as the prediction unit (PU) of a progressively decoded refresh (GDR) frame or an intra-frame random access point (IRAP) frame, which is not output until the reconstructed signal is generated.

[0252] Syntax elements This can indicate the number of subpicks. For example, the value obtained by adding 1 to this can represent the number of subpicks belonging to an individual CLVS. The value of sps_num_subpics_minus1 can have values ​​from 0 to...

[0253] The value of Ceil(pic_width_max_in_luma_samples÷CtbSizeY)*Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1. When the value of sps_num_subpics_minus1 does not exist, the value of sps_num_subpics_minus1 can be deduced to be 0.

[0254] Syntax elements A value of 1 indicates that intra-frame prediction, inter-frame prediction, and in-loop filtering are not performed outside the boundaries of subframes in CLVS.

[0255] Syntax elements A value of 0 indicates that inter-frame prediction or in-loop filtering can be performed outside the boundaries of subpics in CLVS. When the value of sps_independent_subpics_flag is not present, its value can be deduced to be 0.

[0256] Syntax elements This can indicate the horizontal position of the top-left CTU of the i-th subpic, in units of CtbSizeY. The length of the syntax element subpic_ctu_top_left_x[i] can be Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. When subpic_ctu_top_left_x[i] does not exist, its value can be deduced to be 0. Here, pic_width_max_in_luma_samples can be a variable indicating the maximum width of the picture expressed in units of luminance samples. CtbSizeY can be a variable indicating the size of the luminance sample unit of the CTB. CtbLog2SizeY can be a variable indicating the value obtained by taking log2 of the luminance sample unit size of the CTB.

[0257] Syntax elements This indicates the vertical position of the top-left CTU of the i-th subpic, expressed in units of CtbSizeY. The length of the syntax element subpic_ctu_top_left_x[i] can be Ceil(Log2((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. Here, pic_height_max_in_luma_samples can be a variable indicating the maximum height of the picture expressed in units of luminance samples. When subpic_ctu_top_left_y[i] does not exist, its value can be deduced to be 0.

[0258] By using 1 with syntax elements The summed value indicates the width of the first sub-picture, and its unit can be CtbSizeY. The length of subpic_width_minus1[i] can be Ceil(Log2((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. When the value of subpic_width_minus1[i] does not exist, the value of subpic_width_minus1[i] can be calculated as ((pic_width_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_x[i]-1.

[0259] By using 1 with syntax elements The summed value indicates the height of the first subpic, and its unit can be CtbSizeY. The length of subpic_height_minus1[i] can be Ceil(Log2((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)) bits. When subpic_height_minus1[i] does not exist, the value of subpic_height_minus1[i] can be calculated as ((pic_height_max_in_luma_samples+CtbSizeY-1)>>CtbLog2SizeY)-subpic_ctu_top_left_y[i]-1.

[0260] Syntax elements A value of 1 indicates that, except for in-loop filtering operations, the i-th subpic of each encoded frame in CLVS is treated as a frame. A value of 0 indicates that, except for in-loop filtering operations, the i-th subpic of each encoded frame in CLVS is not treated as a frame. When subpic_treated_as_pic_flag[i] does not exist, its value can be set to the value of sps_independent_subpics_flag.

[0261] Syntax elements A value of 1 indicates that in-loop filtering can be performed outside the boundary of the i-th subpic in each encoded frame of CLVS. A value of 0 for loop_filter_across_subpic_enabled_flag[i] indicates that in-loop filtering is not performed outside the boundary of the i-th subpic in each encoded frame of CLVS. When the value of loop_filter_across_subpic_enabled_flag[i] is not present, the value of loop_filter_across_subpic_enabled_flag[i] can be determined as 1 - sps_independent_subpics_flag.

[0262] Figure 26 This is a view illustrating an implementation of the syntax for the screen parameter set. Figure 26 In the syntax of , the syntax elements are as follows.

[0263] Syntax elements The first value (e.g., 0) can indicate that each frame of the reference PPS can be divided into two or more tiles or slices. The second value of no_pic_partition_flag (e.g., 1) can indicate that frame partitioning is not applied to each frame of the reference PPS.

[0264] By combining 5 with syntax elements The summed value indicates the luminance coding block size for each CTU. The value of pps_log2_ctu_size_minus5 can be limited to equal to sps_log2_ctu_size_minus5, which indicates the same value in the sequence parameter set.

[0265] By using 1 with syntax elements The sum of the values ​​indicates the number of tile column widths explicitly provided. The value of num_exp_tile_columns_minus1 can range from 0 to PicWidthInCtbsY-1. When the value of no_pic_partition_flag is 1, the value of num_exp_tile_columns_minus1 can be deduced to be 0.

[0266] By using 1 with syntax elements The sum of these values ​​indicates the number of tile row heights explicitly provided. The value of num_exp_tile_rows_minus1 can range from 0 to PicHeightInCtbsY-1. When no_pic_partition_flag is 1, the value of num_exp_tile_rows_minus1 can be deduced to be 0.

[0267] By using 1 with syntax elements The summed value indicates the width of the i-th tile column in CTB units. Here, i can have a value from 0 to num_exp_tile_columns_minus1-1. tile_column_width_minus1[num_exp_tile_columns_minus1] can be used to derive the width of a tile with an index equal to or greater than num_exp_tile_columns_minus1. The value of tile_column_width_minus1[i] can have a value from 0 to PicWidthInCtbsY-1. When tile_column_width_minus1[i] is not provided from the bitstream, the value of tile_column_width_minus1[0] can be set to the value of PicWidthInCtbsY-1.

[0268] By using 1 with syntax elements The summed value indicates the height of the i-th tile row in CTB units. Here, i can have a value from 0 to num_exp_tile_rows_minus1-1. tile_row_height_minus1[num_exp_tile_rows_minus1] can be used to derive the height of a tile with an index equal to or greater than num_exp_tile_rows_minus1. The value of tile_row_height_minus1[i] can have a value from 0 to PicHeightInCtbsY-1. When tile_row_height_minus1[i] is not provided from the bitstream, the value of tile_row_height_minus1[0] can be set to the value of PicHeightInCtbsY-1.

[0269] Syntax elements A value of 0 indicates that each tile in the slice is scanned in raster scan order and slice information is not signaled via the screen parameter set. A value of 1 for rect_slice_flag indicates that each tile in the slice covers a rectangular area of ​​the screen and slice information is signaled via the screen parameter set. Here, the variable NumTilesInPic can represent the number of tiles present in the screen. When rect_slice_flag is not present in the bitstream, its value can be deduced to be 1. Furthermore, when the value of subpic_info_present_flag is 1, the value of rect_slice_flag can be forced to 1.

[0270] Syntax elements A value of 1 indicates that each subpicture consists of only one rectangular slice. A value of 0 for single_slice_per_subpic_flag indicates that each subpicture consists of at least one rectangular slice. When single_slice_per_subpic_flag is not present in the bitstream, its value can be deduced to be 0.

[0271] By using 1 with syntax elements The sum of these values ​​indicates the number of slices in the image. (Syntax element) A value of 0 indicates that the `tile_idx_delta[i]` syntax element does not exist in the screen parameter set, and all screens in the reference screen parameter set are divided into rectangular slice rows and columns according to the slice raster scan order. A value of 1 for `tile_idx_delta_present_flag` indicates that the `tile_idx_delta[i]` syntax element can exist in the screen parameter set, and all rectangular slices of screens belonging to the reference screen parameter set are specified according to the incremented value of `i` in the order indicated by the value of `tile_idx_delta[i]`. When `tile_idx_delta_present_flag` does not exist, its value can be deduced to be 0.

[0272] By using 1 with syntax elements The summed value indicates the width of the i-th rectangular slice in units of tile columns. The value of `slice_width_in_tiles_minus1[i]` can range from 0 to `NumTileColumns-1`. Here, when i is less than `num_slices_in_pic_minus1` and `NumTileColumns` is 1, the value of `slice_width_in_tiles_minus1[i]` can be deduced to be 0. Here, the variable `NumTileColumns` indicates the number of tile columns present in the current frame. Here, the variable `NumTileRows` indicates the number of tile rows present in the current frame.

[0273] When the value of num_exp_slices_in_tile[i] is 0, it is achieved by adding 1 to the syntax element. The summed value indicates the height of the i-th rectangular slice in units of tile rows. The value of slice_height_in_tiles_minus1[i] can range from 0 to NumTileRows-1. When the value of i is less than num_slices_in_pic_minus1 and the value of slice_height_in_tiles_minus1[i] is not obtained from the bitstream, the value of slice_height_in_tiles_minus1[i] can be derived by the following formula.

[0274] [Formula 1]

[0275] slice_height_in_tiles_minus1[i]=NumTileRows==1?0:slice_height_in_tiles_minus1[i-1]

[0276] SliceTopLeftTileIdx can be a variable that indicates the index of the top-left tile of the slice.

[0277] Syntax elements This can indicate the number of slices whose height is explicitly provided relative to the slices in the tile that includes the i-th slice (e.g., a tile with the same tile index as SliceTopLeftTileIdx[i]). The value of num_exp_slices_in_tile[i] can have values ​​from 0 to RowHeight[SliceTopLeftTileIdx[i] / NumTileColumns]-1. When num_exp_slices_in_tile[i] is not provided from the bitstream, the value of num_exp_slices_in_tile[i] can be deduced to be 0. Here, RowHight[i] can be a variable indicating the height of the i-th tile in CTB units. Here, when the value of num_exp_slices_in_tile[i] is 0, the tile that includes the i-th slice may not be split into multiple tiles.

[0278] By using 1 with syntax elements The summed value indicates the height of the j-th rectangular slice in the tile containing the i-th slice, in CTU units. The value of exp_slice_height_in_ctus_minus1[i][j] can be from 0 to RowHeight[SliceTopLeftTileIdx[i] / NumTileColumns]-1.

[0279] The variable NumSlicesInTile[i] can indicate the number of slices present in the tile that includes the i-th slice.

[0280] Syntax elements This can indicate the difference between the tile index of the tile containing the first CTU in the (i+1)th rectangular slice and the tile index of the tile containing the first CTU in the ith rectangular slice. The value of tile_idx_delta[i] can have values ​​from -NumTilesInPic+1 to NumTilesInPic-1. When the value of tile_idx_delta[i] is not present in the bitstream, the value of tile_idx_delta[i] can be deduced to be 0. When the value of tile_idx_delta[i] is present, the value of tile_idx_delta[i] can be forced to a non-zero value.

[0281] Syntax elements A value of 1 indicates that in-loop filtering can be performed outside the tile boundaries in the frame of the reference frame parameter set. A value of 0 for loop_filter_across_tiles_enabled_flag indicates that in-loop filtering cannot be performed outside the tile boundaries in the frame of the reference frame parameter set.

[0282] In-loop filtering operations can include any of the following: deblocking filter, Sample Adaptive Offset (SAO) filter, or Adaptive Loop Filter (ALF). When loop_filter_across_tiles_enabled_flag is not present in the bitstream, the value of loop_filter_across_tiles_enabled_flag can be deduced to be 1.

[0283] Syntax elements A value of 1 indicates that in-loop filtering can be performed outside the slice boundaries in the frame of the reference frame parameter set. A value of 0 for `loop_filter_across_slice_enabled_flag` indicates that in-loop filtering is not performed outside the slice boundaries in the frame of the reference frame parameter set. In-loop filtering can include any of the following: deblocking filter, Sample Adaptive Offset (SAO) filter, or Adaptive Loop Filter (ALF). When `loop_filter_across_slice_enabled_flag` is not present in the bitstream, its value can be deduced to be 1.

[0284] Figure 27 This is a view illustrating an implementation of the syntax for the slice header. Figure 27 In the syntax of , the syntax elements are as follows.

[0285] Syntax elements This can indicate the subpick ID of the subpick containing the slice. When the value of `slice_subpic_id` exists in the bitstream, the value of the variable `CurrSubpicIdx` can be derived from the value of `SubpicIdVal[CurrSubpicIdx]` of `slice_subpic_id`. Otherwise (if `slice_subpic_id` does not exist in the bitstream), the value of `CurrSubpicIdx` can be derived as 0. The length of `slice_subpic_id` can be `sps_subpic_id_len_minus1+1` bits. Here, `NumSlicesInSubpic[i]` can be a variable indicating the number of slices in the i-th subpick. The variable `CurrSubpicIdx` can indicate the index of the current subpick.

[0286] Syntax elements Indicates the slice address. When slice_address is not provided, the value of slice_address can be deduced to be 0.

[0287] Furthermore, when the value of `rect_slice_flag` is 0, `slice_address` can be equal to the raster scan tile index of the first tile in the slice, and the length of the `slice_address` syntax element can be Ceil(Log2(NumTilesInPic)) bits, with values ​​ranging from 0 to NumTilesInPic-1. Otherwise, (if the value of `rect_slice_flag` is non-zero, for example, 1), the slice address can be the sub-picture level slice index of the slice, and the length of the `slice_address` syntax element can be Ceil(Log2(NumSlicesInSubpic[CurrSubpicIdx])) bits, with values ​​ranging from 0 to NumSlicesInSubpic[CurrSubpicIdx]-1.

[0288] Syntax elements It can have a value of 0 or 1. The decoding device can perform decoding regardless of the value of sh_extra_bit[i]. For this purpose, the encoding device needs to generate a bitstream such that decoding is performed regardless of the value of sh_extra_bit[i]. Here, NumExtraShBits can be a variable indicating the number of bits required to signal further information in the slice header.

[0289] By using 1 with syntax elements The sum of the values ​​can indicate the number of tiles in the slice (if any). The value of num_tiles_in_slice_minus1 can be from 0 to NumTilesInPic-1.

[0290] The variable NumCtusInCurrSlice, which indicates the number of CTUs in the current slice, and the list CtbAddrInCurrSlice[i], which indicates the raster scan address of the i-th CTU in the slice, can be derived as follows (where i has a value from 0 to NumCtusInCurrSlice-1).

[0291] [Table 2]

[0292]

[0293] The variables SubpicLeftBoundaryPos, SubpicTopBoundaryPos, SubpicRightBoundaryPos, and SubpicBotBoundaryPos can be derived using the following algorithm.

[0294] [Table 3]

[0295]

[0296] Improvements to screen splitting signaling

[0297] The aforementioned signaling related to screen segmentation has the problem of signaling unnecessary information when the slice is rectangular. For example, when the slice is quadrilateral (e.g., rectangular), the width of an individual slice can be signaled in units of a tile. However, when the top-left tile of a slice is the tile in the last tile column, the slice width may not be a value other than a single tile unit. For example, in this case, the slice width may only have a width value derived from a single tile unit. Therefore, the width of such a slice may not be signaled or may be limited to a single tile unit.

[0298] Similarly, when the slice is a quadrilateral (e.g., rectangular) slice, the width of individual slices can be signaled in chunk units. However, when the chunk at the top left position of a slice is the chunk of the last chunk row, the height of the slice may not be a value other than one chunk unit. Therefore, the height of such a slice may not be signaled or may be limited to one chunk unit.

[0299] To address the aforementioned problems, the following methods can be applied. The following implementations are applicable when the slices are quadrilateral (e.g., rectangular) slices and the width and / or height of individual slices are signaled in units of tiles. The following methods can be applied individually or in combination with at least one other implementation.

[0300] Method 1. When the first tile of a rectangular slice (e.g., the top-left tile) is located in the last tile column of the screen, signaling for the slice width may not be provided. In this case, the slice width can be derived as a tile unit.

[0301] For example, the syntax element slice_width_in_tiles_minus1[i] may not exist in the bitstream. The value of the syntax element slice_width_in_tiles_minus1[i] can be derived as 0.

[0302] Method 2. Signaling for the slice width can be provided even when the first tile of a rectangular slice (e.g., the top-left tile) is located in the last tile column of the screen. However, in this case, the slice width can be limited to one tile unit.

[0303] For example, the syntax element slice_width_in_tiles_minus1[i] can exist in the bitstream and therefore be parsed. However, the value of the syntax element slice_width_in_tiles_minus1[i] can be limited to 0.

[0304] Method 3. When the first tile of a rectangular slice (e.g., the top-left tile) is the tile in the last tile row of the image, signaling for the slice's height may not be provided. In this case, the slice's height can be derived as a tile unit.

[0305] For example, the syntax element slice_height_in_tiles_minus1[i] may not exist in the bitstream. The value of the syntax element slice_height_in_tiles_minus1[i] can be deduced to be 0.

[0306] Method 4. When the first tile of a rectangular slice (e.g., the top-left tile) is the tile in the last tile row of the image, signaling for the slice's height can be provided. However, in this case, the slice's height can be limited to one tile unit.

[0307] For example, the syntax element slice_height_in_tiles_minus1[i] can exist in the bitstream and therefore be parsed. However, the value of the syntax element slice_height_in_tiles_minus1[i] can be restricted to 0.

[0308] In one embodiment, the above embodiments are applicable to Figure 28 and Figure 29 The encoding and decoding methods are shown. According to one embodiment, the encoding device can deduce slices and / or tiles in the current frame (S2810). Furthermore, the encoding device can encode the current frame based on the deduced slices and / or tiles (S2820).

[0309] Similarly, the decoding device according to the embodiment can acquire video / image information from the bitstream (S2910). Furthermore, the decoding device can deduce the slices and / or tiles present in the current frame based on the video / image information (including information about slices and / or tiles) (S2920). Additionally, the decoding device can reconstruct and / or decode the current frame based on the slices and / or tiles (S2930).

[0310] For the above processing by the encoding and decoding devices, the information regarding slices and / or tiles may include the aforementioned information and syntax. Video or image information may include HLS. HLS may include information about slices and / or information about tiles. HLS may also include information about sub-frames. Information about slices may include information specifying at least one slice belonging to the current frame. Additionally, information about tiles may include information specifying at least one tile belonging to the current frame. Information about sub-frames may include information specifying at least one sub-frame belonging to the current frame. A tile including at least one slice may exist within a single frame.

[0311] For example, in Figure 29 The S2930 can reconstruct and / or decode the current frame based on derived slices and / or tiles. By segmenting a frame, encoding and decoding efficiency can be achieved in various aspects.

[0312] For example, frames can be segmented for parallel processing and error recovery. In the case of parallel processing, some implementations executing on a multi-core CPU may require segmenting the source frame into tiles and / or slices. Individual slices and / or tiles can be processed in parallel on different cores. This is highly efficient for performing high-resolution real-time video coding that cannot be performed by other methods. Additionally, this segmentation has the advantage of reducing memory constraints by minimizing the information shared between tiles. Parallel architectures are useful due to their segmentation mechanism, as they distribute tiles across different threads while performing parallel processing. For example, in deriving motion information in inter-frame prediction, neighboring blocks existing in different slices and / or tiles can be restricted to not being used. Context information used for encoding information and / or syntax elements can be initialized for each slice and / or tile.

[0313] Error recovery can be achieved by applying Inequality Error Protection (UEP) to encoded blocks and / or slices.

[0314] Implementation Method 1

[0315] The following describes implementations based on methods 1 and 3 described above. These implementations can be applied to improve encoding / decoding techniques, such as the VVC specification.

[0316] In one implementation, it can be as follows: Figure 30 The diagram shows a syntax table for notifying the screen parameter set using signals. In another embodiment, it can be as follows: Figure 31 The settings shown are syntax tables used to notify the screen parameter set with signals.

[0317] exist Figure 30 In the implementation, for i with values ​​from 0 to num_slices_in_pic_minus1-1, when the value of NumTileColumns is greater than 1 and the value of SliceTopLeftTileIdx[i]%NumTileColumns is not NumTileColumns-1, the syntax element slice_width_in_tile_minus1[i] can be obtained sequentially relative to i.

[0318] Additionally, for an i with a value from 0 to num_slices_in_pic_minus1-1, when the value of NumTileRows is greater than 1, the value of tile_idx_delta_present_flag is 1, or the value of SliceTopLeftTileIdx[i]%NumTileColumns is 0 and the value of SliceTopLeftTileIdx[i] / NumTileColumns is not NumTileRows-1, the syntax element slice_height_in_tiles_minus1[i] can be obtained sequentially relative to i.

[0319] exist Figure 30 and Figure 31 In this implementation, the syntax element slice_width_in_tiles_minus1[i] can be a syntax element indicating the width of the i-th rectangular slice. For example, the value obtained by adding 1 to slice_width_in_tiles_minus1[i] can indicate the width of the i-th rectangular slice in units of tile columns. The value of slice_width_in_tiles_minus1[i] can have values ​​from 0 to NumTileColumns-1. When slice_width_in_tiles_minus1[i] is not obtained from the bitstream, the value of slice_width_in_tiles_minus1[i] can be deduced to be 0.

[0320] When the definition of slice_width_in_tiles_minus1[i] is changed as described above, the constraint that "the value of slice_width_in_tiles_minus1[i] is deduced to be 0 when i is less than num_slices_in_pic_minus1 and the value of NumTileColumns is equal to 1" can be omitted. Therefore, as in Figure 31 In this implementation, the constraint "NumTileColumns>1" can be removed from the screen parameter set syntax.

[0321] `slice_height_in_tiles_minus1[i]` can be a syntax element indicating the height of the i-th rectangular slice. For example, when `num_exp_slices_in_tile[i]` is 0, the value obtained by adding 1 to `slice_height_in_tiles_minus1[i]` can indicate the height of the i-th rectangular slice in units of tile rows. The value of `slice_height_in_tiles_minus1[i]` can have values ​​from 0 to `NumTileRows-1`.

[0322] When slice_height_in_tiles_minus1[i] is not obtained from the bitstream, the value of slice_height_in_tiles_minus1[i] can be derived as follows.

[0323] First, when the value of NumTileRow is 1 or the value of SliceTopLeftTileIdx[i]%NumTileColumns is NumTileColumns-1, the value of slice_height_in_tiles_minus1[i] can be deduced to be 0.

[0324] Otherwise (e.g., when the value of NumTileRow is not 1 and the value of SliceTopLeftTileIdx[i]%NumTileColumns is not NumTileColumns-1), the value of slice_height_in_tiles_minus1[i] can be derived as slice_height_in_tiles_minus1[i-1]. For example, the value of slice_height_in_tiles_minus1[i] can be set as slice_height_in_tiles_minus1[i-1], which is the height value of the previous slice. For example, the value of slice_height_in_tiles_minus1[i] can be set equivalently for all slices in a tile.

[0325] When the definition of slice_height_in_tiles_minus1[i] is changed as described above, the existing constraint that "when i is less than num_slices_in_pic_minus1 and the value of NumTileRows is equal to 1, the value of slice_width_in_tiles_minus1[i] is deduced to be 0" can be omitted. Therefore, as in Figure 31In one implementation, the constraint "NumTileRows>1" can be removed from the screen parameter set syntax.

[0326] Implementation Method 2

[0327] The following sections will describe implementations based on methods 2 and 4 described above. These implementations can be applied to improve encoding / decoding techniques, such as the VVC specification.

[0328] In one implementation, slice_width_in_tiles_minus1[i] can be a syntax element indicating the width of the i-th rectangular slice. For example, the value obtained by adding 1 to slice_width_in_tiles_minus1[i] can indicate the width of the i-th rectangular slice in units of tile columns. The value of slice_width_in_tiles_minus1[i] can have values ​​from 0 to NumTileColumns-1. When slice_width_in_tiles_minus1[i] is not obtained from the bitstream, the value of slice_width_in_tiles_minus1[i] can be deduced to be 0.

[0329] At this point, when i is less than num_slices_in_pic_minus1 and the value of NumTileColumns is equal to 1, the value of slice_width_in_tiles_minus1[i] can be deduced to be 0. Additionally, for bitstream consistency, when the first piece of the i-th rectangular slice is the last piece of the slice column, the value of slice_width_in_tiles_minus1[i] can be forced to be 0.

[0330] `slice_height_in_tiles_minus1[i]` can be a syntax element indicating the height of the i-th rectangular slice. For example, when `num_exp_slices_in_tile[i]` is 0, the value obtained by adding 1 to `slice_height_in_tiles_minus1[i]` can indicate the height of the i-th rectangular slice in units of tile rows. The value of `slice_height_in_tiles_minus1[i]` can have values ​​from 0 to `NumTileRows-1`.

[0331] In this case, when i is less than num_slices_in_pic_minus1 and the value of slice_height_in_tiles_minus1[i] is not obtained from the bitstream, the value of slice_height_in_tiles_minus1[i] can be determined based on the value of NumTileRows. For example, this can be determined as shown in the following formula.

[0332] [Equation 2]

[0333] slice_height_in_tiles_minus1[i]=NumTileRows==1?0:slice_height_in_tiles_minus1[i-1]

[0334] Additionally, for bitstream consistency, the value of slice_height_in_tiles_minus1[i] can be forced to 0 when the first tile of the i-th rectangular slice is the last tile of the tile row.

[0335] Encoding and Decoding Methods

[0336] In the following text, an image encoding method performed by an image encoding device according to an embodiment and an image decoding method performed by an image decoding device will be described.

[0337] First, the operation of the decoding device will be described. According to an embodiment, the image decoding device may include a memory and a processor, and the decoding device can perform decoding through the operation of the processor. Figure 32 This is a view illustrating an embodiment of the decoding method according to the embodiment.

[0338] According to the implementation, the decoding device can obtain the syntax element no_pic_partition_flag, which indicates the availability of the current frame segmentation, from the bitstream. As described above, the decoding device can determine the availability of the current frame segmentation based on the value of no_pic_partition_flag (S3210).

[0339] When the current frame is available for segmentation, the decoding device can obtain from the bitstream the syntax element num_exp_tile_rows_minus1 indicating the number of tile rows for segmenting the current frame and the syntax element num_exp_tile_columns_minus1 indicating the number of tile columns, and determine the number of tile rows and tile columns from them as described above (S3220).

[0340] Based on the number of tile columns, the decoding device can obtain the syntax element tile_column_width_minus1, which indicates the width of each tile column that divides the current frame, from the bitstream, and determine the width of each tile column from it as described above (S3230).

[0341] Based on the number of tile rows, the decoding device can obtain the syntax element `tile_row_height_minus1[i]` from the bitstream, which indicates the height of each tile row that divides the current frame, and determine the height of each tile row from it (S3240). Alternatively, the decoding device can calculate the number of tiles dividing the current frame by multiplying the number of tile columns by the number of tile rows.

[0342] Next, the decoding device can obtain the syntax element rect_slice_flag indicating whether the current screen is divided into rectangular slices based on whether the number of tiles that divide the current screen is greater than 1, and determine whether the current screen is divided into rectangular slices from its value as described above (S3250).

[0343] Next, the decoding device can obtain the syntax element num_slices_in_pic_minus1, which indicates the number of slices that divide the current frame, from the bitstream based on whether the current frame is divided into rectangular slices, and determine the number of slices that divide the current frame from it as described above (S3260).

[0344] Next, the decoding device can obtain size information from the bitstream (S3270), which indicates the size of each slice of the current frame from the bitstream, as many as the number of slices that divide the current frame.

[0345] Here, the size information may include the syntax element slice_width_in_tiles_minus1[i] which indicates the width of the slice and the syntax element slice_height_in_tiles_minus1[i] which indicates the height of the slice. slice_width_in_tiles_minus1[i] can indicate the width of the slice in units of tile columns, and slice_height_in_tiles_minus1[i] can indicate the height of the slice in units of tile rows.

[0346] Here, when the decoding device obtains the size information of the current slice (e.g., the i-th slice) from the bitstream, the decoding device can obtain slice_width_in_tiles_minus1[i] from the bitstream based on whether the top-left tile of the current slice belongs to the last tile column of the current frame.

[0347] For example, when the top-left tile index of the current slice (e.g., SliceTopLeftTileIdx) does not correspond to the tile index of the last column of the tile column belonging to the current frame, slice_width_in_tiles_minus1[i] can be obtained from the bitstream. However, when the top-left tile index of the current slice is the tile index of the last column of the tile column belonging to the current frame, slice_width_in_tiles_minus1[i] can be left unobtained from the bitstream and can be set to 0.

[0348] Similarly, the decoding device can obtain slice_height_in_tiles_minus1[i] from the bitstream based on whether the top-left tile of the current slice belongs to the last tile row of the current frame.

[0349] For example, when the top-left tile index of the current slice is not the tile index corresponding to the last row of tile rows belonging to the current frame, slice_height_in_tiles_minus1[i] can be obtained from the bitstream. However, when the top-left tile index of the current slice is the tile index corresponding to the last row of tile rows belonging to the current frame, slice_height_in_tiles_minus1[i] can be left unobtained from the bitstream and can be set to 0.

[0350] Next, the decoding device can determine the size of each slice of the current frame based on the size information, and decode the determined slices to decode the image. For example, inter-frame prediction or intra-frame prediction can be used to decode the CTU included in the slice with a determined size to decode the slice (S3280).

[0351] Furthermore, SliceTopLeftTileIdx can be a variable indicating the top-left tile of the slice, and can be accessed via... Figure 33 and Figure 34 The algorithm is used to determine this. Figure 33 and Figure 34 The algorithm is a continuous algorithm.

[0352] Next, the operation of the encoding device will be described. The image encoding device according to the embodiment may include a memory and a processor, and the encoding device can perform encoding in a manner corresponding to decoding by the operation of the processor. For example, as... Figure 35As shown, the encoding device can encode the current frame. First, the encoding device can determine the tile columns and tile rows of the current frame (S3510). Next, it can determine the slices of the segmented image (S3520). Next, the encoding device can generate a bit stream including predetermined information, which includes slice size information. For example, the encoding device can generate a bit stream that includes no_pic_partition_flag, num_exp_tile_rows_minus1, num_exp_tile_columns_minus1, tile_column_width_minus1, tile_row_height_minus1[i], rect_slice_flag, num_slices_in_pic_minus1, slice_width_in_tiles_minus1[i], and slice_height_in_tiles_minus1[i] as syntax elements obtained from the bit stream by the decoding device.

[0353] At this point, size information can be included in the bitstream based on whether the current slice belongs to the last tile column or the last tile row of the current frame. For example, the encoding device can encode the bitstream based on whether the top-left tile of the current slice is the last tile column and / or the last tile row, such that slice_width_in_tiles_minus1[i] and slice_height_in_tiles_minus1[i] correspond to the description of the decoding device. Here, the current slice can be a rectangular slice.

[0354] Application and Implementation Methods

[0355] Although the exemplary methods of this disclosure described above are represented as a series of operations for clarity of description, they are not intended to limit the order in which the steps are performed, and these steps may be performed simultaneously or in different orders if necessary. To implement the methods according to this disclosure, the described steps may further include other steps, including steps in addition to some steps, or may include additional steps in addition to some steps.

[0356] In this disclosure, an image encoding device or an image decoding device that performs a predetermined operation (step) can perform an operation (step) that confirms the execution conditions or circumstances of the corresponding operation (step). For example, if it is described that a predetermined operation is performed when predetermined conditions are met, the image encoding device or the image decoding device can perform the predetermined operation after determining whether the predetermined conditions are met.

[0357] The various embodiments of this disclosure are not a list of all possible combinations and are intended to describe representative aspects of this disclosure; the matters described in the various embodiments may be applied independently or in combination of two or more.

[0358] Various embodiments of this disclosure can be implemented in hardware, firmware, software, or a combination thereof. When this disclosure is implemented in hardware, it can be implemented using application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, etc.

[0359] Furthermore, the image decoding and image encoding devices applying the embodiments of this disclosure can be included in multimedia broadcasting transmitting and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, OTT (over-the-top) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, medical video devices, etc., and can be used to process video signals or data signals. For example, OTT video devices can include game consoles, Blu-ray players, internet access televisions, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.

[0360] Figure 36 This is a view illustrating a content streaming system to which embodiments of the present disclosure can be applied.

[0361] like Figure 36 As shown, the content streaming system using the embodiments of this disclosure may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.

[0362] An encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, which is then sent to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.

[0363] The bitstream can be generated by an image encoding method or image encoding device applying the embodiments of this disclosure, and the stream server can temporarily store the bitstream during the sending or receiving of the bitstream.

[0364] A streaming server sends multimedia data to a user's device based on a request from a web server, and the web server acts as a medium for informing the user of the service. When a user requests a service from the web server, the web server can deliver it to the streaming server, and the streaming server can send the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between devices in the content streaming system.

[0365] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.

[0366] Examples of user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet computers, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.

[0367] In a content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed.

[0368] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or commands stored thereon and executable on a device or computer.

[0369] Industrial applicability

[0370] The embodiments disclosed herein can be used to encode or decode images.

Claims

1. An image decoding method performed by an image decoding device, the image decoding method comprising the following steps: Obtain size information from the bitstream that indicates the size of the current slice corresponding to at least a portion of the current frame; as well as The size of the current slice is determined based on the size information. The current slice is a rectangular slice and includes at least one tile. The size information includes width information indicating the width of the current slice in units of tile columns and height information indicating the height of the current slice in units of tile rows. The step of obtaining the size information is performed based on whether the current slice belongs to the last tile column or the last tile row of the current frame. Specifically, the width information of the current slice is skipped from the bitstream because the top left tile of the current slice belongs to the last tile column of the current image.

2. The image decoding method according to claim 1, wherein, Based on the fact that the top-left tile of the current slice belongs to the last tile column of the current frame, the width information of the current slice obtained from the bitstream is skipped, and the width information is determined to be a predetermined value, which is 0. The predetermined value indicates a tile column.

3. The image decoding method according to claim 1, wherein, The height information of the current slice is skipped from the bitstream because the top-left tile of the current slice belongs to the last tile row of the current frame.

4. The image decoding method according to claim 1, wherein, Based on the fact that the top-left tile of the current slice does not belong to the last tile row of the current frame, the height information of the current slice is obtained from the bitstream.

5. The image decoding method according to claim 1, wherein, Based on the fact that the top-left tile of the current slice belongs to the last tile row of the current frame, the height information of the current slice obtained from the bitstream is skipped, and the height information is determined to be a predetermined value, which is 0. The predetermined value indicates a tile row.

6. The image decoding method according to claim 1, in, The step of obtaining the size information is performed based on the number of slices that divide the current frame. The number of slices that divide the current frame is determined by the following: Based on the segmentation information obtained from the bitstream, the availability of segmentation for the current frame is determined; Based on the availability of the segmentation of the current screen, the number of tile rows and the number of tile columns are determined, and the tile rows and the tile columns segment the current screen. The width of each tile column that divides the current frame is determined based on the number of tile columns. The height of each tile row that divides the current frame is determined based on the number of tile rows. The number of tiles used to segment the current image determines whether the current image is divided into rectangular slices; and Based on whether the current frame is divided into rectangular slices, the number of slices that divide the current frame is obtained from the bitstream.

7. An image encoding method performed by an image encoding device, the image encoding method comprising the following steps: Determine the current slice that corresponds to at least a portion of the current frame; as well as The size information of the current slice is encoded into a bitstream. The current slice is a rectangular slice and includes at least one tile. The size information includes width information indicating the width of the current slice in units of tile columns and height information indicating the height of the current slice in units of tile rows. The step of encoding the size information is performed based on whether the current slice belongs to the last tile column or the last tile row of the current frame. Specifically, the width information of the current slice is skipped from being encoded in the bitstream because the top-left tile of the current slice belongs to the last tile column of the current frame.

8. A method for transmitting a bitstream generated by an image encoding method, the image encoding method comprising the following steps: Determine the current slice that corresponds to at least a portion of the current frame; as well as The size information of the current slice is encoded into the bitstream. The current slice is a rectangular slice and includes at least one tile. The size information includes width information indicating the width of the current slice in units of tile columns and height information indicating the height of the current slice in units of tile rows. The step of encoding the size information is performed based on whether the current slice belongs to the last tile column or the last tile row of the current frame. Specifically, the width information of the current slice is skipped from being encoded in the bitstream because the top-left tile of the current slice belongs to the last tile column of the current frame.