Image / video encoding method and apparatus based on slice type

By introducing the concept of slice type in image/video coding and using flags to indicate the existence of inter-frame prediction information, intra-frame prediction and inter-frame prediction are optimized, solving the problems of large information volume and high transmission cost in high-resolution image/video coding, and improving coding efficiency and signaling efficiency.

CN120956928APending Publication Date: 2025-11-14LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511225499.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-11-05
Filing Date
2020-11-05
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as large information volume and high transmission and storage costs in the encoding of high-resolution, high-quality images/videos, especially in image/video broadcasting of virtual reality and immersive media, where the signaling efficiency of inter-frame prediction and intra-frame prediction is low.

Method used

By introducing the concept of slice type in image/video coding, and using a first or second flag to indicate the presence of information required for inter-frame prediction operations, the execution of intra-frame and inter-frame prediction is optimized, unnecessary signaling is reduced, and coding efficiency is improved.

Benefits of technology

It improves image/video compression efficiency, enhances the efficiency of inter-frame and intra-frame prediction, reduces unnecessary signaling, and improves the overall efficiency of the coding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120956928A_ABST
    Figure CN120956928A_ABST
Patent Text Reader

Abstract

The invention provides an image / video encoding method and apparatus based on slice type. A video decoding method performed by a video decoding apparatus according to the present document comprises the steps of: obtaining image information from a bitstream, the image information including a picture header associated with a current picture, and the current picture including a plurality of slices; parsing at least one of the first flag and the second flag from the picture header; generating a prediction sample by performing at least one of intra prediction and inter prediction on a current block in the current picture based on at least one of the first flag and the second flag; generating a reconstructed sample based on the prediction sample; and generating a reconstructed picture based on the reconstructed sample, in which the first flag may indicate whether the information necessary for the inter prediction operation exists in the picture header, and the second flag may indicate whether the information necessary for the inter prediction operation exists in the picture header.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of patent application No. 202080084681.8 (PCT / KR2020 / 015400), filed on June 7, 2022, with an application date of November 5, 2020, entitled "Image / Video Compilation Method and Apparatus Based on Slice Type". Technical Field

[0002] This technology relates to a method and apparatus for encoding images / videos based on information about slice type. Background Technology

[0003] Recently, there has been an increasing demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher Ultra High Definition (UHD) images / videos, across various fields. As image / video resolution or quality increases, a relatively larger amount of information or bits is transmitted compared to regular image / video data. Therefore, if image / video data is transmitted via media such as existing wired / wireless broadband lines or stored in traditional storage media, the costs for transmission and storage can easily increase.

[0004] Furthermore, there is growing interest and demand for virtual reality (VR) and artificial reality (AR) content, as well as immersive media such as holograms; and the broadcasting of images / videos that exhibit characteristics different from actual images / videos (e.g., game images / videos) is also increasing.

[0005] Therefore, highly efficient image / video compression technology is needed to effectively compress and send, store, or play high-resolution, high-quality images / videos that exhibit the various characteristics described above. Summary of the Invention

[0006] Technical issues

[0007] This document provides a method and apparatus for improving image / video coding efficiency.

[0008] This document also provides a method and apparatus for efficiently performing inter-frame prediction and / or intra-frame prediction in image / video coding.

[0009] This document also provides a method and apparatus for efficiently signaling slice type-related information when sending image / video information.

[0010] This document also provides a method and apparatus for omitting signaling that is not necessary for inter-frame prediction and / or intra-frame prediction when transmitting image / video information.

[0011] Technical solution

[0012] According to embodiments of this document, a video decoding method performed by a video decoding device is provided. The method includes: obtaining image information from a bitstream, wherein the image information includes an image header associated with a current image, and the current image includes a plurality of slices; parsing at least one of a first flag or a second flag from the image header; generating a prediction sample by performing at least one of intra-frame prediction or inter-frame prediction on a current block in the current image based on at least one of the first flag or the second flag; generating a reconstruction sample based on the prediction sample; and generating a reconstructed image based on the reconstruction sample, wherein the first flag indicates whether information necessary for the inter-frame prediction operation exists in the image header, and wherein the second flag indicates whether information necessary for the inter-frame prediction operation exists in the image header.

[0013] According to another embodiment of this document, a video encoding method performed by a video encoding device is provided. The method includes: determining a prediction mode for a current block in a current image, wherein the current image includes a plurality of slices; generating prediction samples based on the prediction mode; generating reconstruction samples for the current block based on the prediction samples; generating at least one of first information or second information based on the prediction mode; and encoding image information including at least one of the first information or second information, wherein the first information and the second information are included in an image header associated with the current image, wherein the first information indicates whether information necessary for inter-frame prediction operations exists in the image header, and wherein the second information indicates whether information necessary for inter-frame prediction operations exists in the image header.

[0014] According to another embodiment of this document, a computer-readable digital storage medium is provided, the computer-readable digital storage medium containing information that causes a decoding device to perform a video decoding method, the decoding method comprising: obtaining image information, wherein the image information includes an image header associated with a current image, and the current image includes a plurality of slices; parsing at least one of a first flag or a second flag from the image header; generating a prediction sample by performing at least one of intra-frame prediction or inter-frame prediction on a current block in the current image based on at least one of the first flag or the second flag; generating a reconstruction sample based on the prediction sample; and generating a reconstructed image based on the reconstruction sample, wherein the first flag indicates whether information necessary for the inter-frame prediction operation exists in the image header, and wherein the second flag indicates whether information necessary for the inter-frame prediction operation exists in the image header.

[0015] Beneficial effects

[0016] According to the embodiments in this document, the overall image / video compression efficiency can be improved.

[0017] According to the embodiments in this document, inter-frame prediction and / or intra-frame prediction can be performed efficiently when encoding images / videos.

[0018] According to the embodiments in this document, when sending image / video information, information related to the slice type can be efficiently notified by signaling.

[0019] According to the embodiments in this document, when sending image / video information, it is possible to prevent signaling of syntax elements that are not necessary for inter-frame prediction or intra-frame prediction. Attached Figure Description

[0020] Figure 1 An example of a video / image encoding system to which embodiments of the present disclosure may be applied is illustrated schematically.

[0021] Figure 2 This is a schematic diagram illustrating the configuration of a video / image encoding apparatus to which embodiments of the present disclosure may be applied.

[0022] Figure 3 This is a schematic diagram illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure may be applied.

[0023] Figure 4 An example is shown of a hierarchical structure used for encoding images / videos.

[0024] Figure 5 An example illustrating the image decoding process.

[0025] Figure 6 An example illustrating the image encoding process.

[0026] Figure 7 This represents an example of a video / image coding method based on inter-frame prediction.

[0027] Figure 8 This represents an example of a video / image decoding method based on inter-frame prediction.

[0028] Figure 9 and Figure 10 Examples of video / image encoding methods and related components according to embodiments of this document are illustrated schematically.

[0029] Figure 11 and Figure 12 Examples of video / image decoding methods and related components according to embodiments of this document are illustrated schematically.

[0030] Figure 13 Examples of content streaming systems that can apply the embodiments disclosed in this document. Detailed Implementation

[0031] This document relates to video / image coding. For example, the methods / exercises disclosed in this document can be applied to methods disclosed in the Universal Video Coding (VVC) standard. Furthermore, the methods / exercises disclosed in this document can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding 2 (AVS2) standard, or next-generation video / image coding standards (e.g., H.267, H.268, etc.).

[0032] This document presents various embodiments related to video / image encoding, and unless otherwise specified, the above embodiments may also be performed in combination with each other.

[0033] The disclosure of this document may be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the methods disclosed herein. The singular expression includes the expression "at least one," provided that it is clearly interpreted differently. Terms such as "comprising" and "having" are intended to indicate the presence of features, quantities, steps, operations, elements, components, or combinations thereof used in the document, and therefore should be understood that the possibility of having or adding one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.

[0034] Furthermore, each configuration in the accompanying drawings described in this document is a separate illustration for explaining the functionality as distinct features, and does not imply that each configuration is implemented by different hardware or different software. For example, two or more configurations may be combined to form one configuration, and one configuration may also be divided into multiple configurations. Embodiments of combined and / or separated configurations are included within the scope of this disclosure without departing from the spirit of the methods disclosed herein.

[0035] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Furthermore, "A, B" can mean "A and / or B". Additionally, "A / B / C" can mean "at least one of A, B, and / or C". Furthermore, "A / B / C" can mean "at least one of A, B, and / or C".

[0036] Furthermore, in this document, the term "or" should be interpreted as indicating "and / or". For example, expressing "A or B" could include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".

[0037] Furthermore, the parentheses used in this specification may mean "for example". Specifically, when expressing "prediction (intra-frame prediction)", it may indicate that "intra-frame prediction" is presented as an example of "prediction". In other words, the term "prediction" in this specification is not limited to "intra-frame prediction", and may indicate that "intra-frame prediction" is presented as an example of "prediction". Moreover, even when expressing "prediction (i.e., intra-frame prediction)", it may indicate that "intra-frame prediction" is presented as an example of "prediction".

[0038] In this specification, a technical feature described individually in a single drawing may be implemented individually or simultaneously.

[0039] In the following, embodiments of this document will be described in detail with reference to the accompanying drawings. Furthermore, throughout the drawings, the same reference numerals will be used to indicate the same elements, and the same descriptions of the same elements will be omitted.

[0040] Figure 1 The illustration shows an example of a video / image encoding system to which embodiments of the present disclosure may be applied.

[0041] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.

[0042] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.

[0043] Video sources can acquire video / images through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. For example, a video / image capture device may include one or more cameras, a video / image archive including previously captured video / images, etc. For example, a video / image generation device may include a computer, tablet computer, and smartphone, and may generate video / images (electronically). For example, virtual video / images can be generated via a computer, etc. In this case, the video / image capture process can be replaced by a process that generates related data.

[0044] Encoding devices can encode input video / images. For compression and encoding efficiency, encoding devices can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output as a bitstream.

[0045] The transmitter can send encoded images / image information or data, output as a bitstream, to the receiver of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to a decoding device.

[0046] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.

[0047] The renderer can render decoded video / images. The rendered video / images can be displayed on a monitor.

[0048] In this document, video can refer to a series of images over time. An image typically refers to a unit representing an image at a specific time frame, and a slice / tile refers to a unit that constitutes part of an image in terms of encoding. A slice / tile can include one or more coding tree units (CTUs). An image can consist of one or more slices / tiles. An image can consist of one or more groups of tiles. A group of tiles can include one or more tiles. A brick can represent a rectangular area of ​​CTU rows within a tile in an image. A tile can be divided into multiple tiles, each tile consisting of one or more CTU rows within the tile. A tile that is not divided into multiple tiles can also be referred to as a tile. Tile scan can represent a specific order of CTUs that divide an image, where CTUs are ordered consecutively by CTU raster scans within a tile, tiles within a tile are ordered consecutively by raster scans of the tiles within the tile, and tiles in an image are ordered consecutively by raster scans of the tiles within the image. A tile is a rectangular region of a CTU within a specific tile column and a specific tile row in an image. A tile column is a rectangular region of a CTU with a height equal to the height of the image and a width specified by a syntax element in the image parameter set. A tile row is a rectangular region of a CTU with a height specified by a syntax element in the image parameter set and a width equal to the width of the image. A tile scan is a specific ordering of the CTUs that segment the image, where the CTUs are ordered consecutively in a tile by a CTU raster scan, and the tiles in the image are ordered consecutively by a tile raster scan. A slice comprises an integer number of tiles of an image that may be contained within a single NAL unit. A slice may consist of multiple complete tiles or a consecutive sequence of complete tiles of a single tile. In this document, tile groups and slices may be used interchangeably. For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.

[0049] A pixel or cell (pel) can refer to the smallest unit that makes up a picture (or image). Additionally, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.

[0050] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region." In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients in M ​​columns and N rows. Alternatively, a sample may refer to a pixel value in the spatial domain, and when such a pixel value is transformed to the frequency domain, it may refer to a transform coefficient in the frequency domain.

[0051] In some cases, a unit can be used interchangeably with terms such as block or region. Typically, an MxN block can represent a sample or a set of transform coefficients consisting of M columns and N rows. A sample can typically represent a pixel or the value of that pixel, and can also represent a pixel / pixel value for only the luminance component, and even a pixel / pixel value for only the chrominance component. A sample can be used as a term corresponding to pixels or cells that configure a picture (or image).

[0052] Figure 2 This diagram schematically illustrates a configuration of a video / image encoding apparatus to which embodiments of the present disclosure can be applied. Hereinafter, an apparatus referred to as a video encoding apparatus may include an image encoding apparatus.

[0053] Reference Figure 2 The encoding device 200 includes an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image partitioner 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.

[0054] Image partitioner 210 can segment an input image (or picture or frame) input to encoding device 200 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit can be recursively segmented from a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree-binary-truncate (QTBTTT) structure. For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure may be applied first, followed by a binary tree structure and / or a ternary structure. Alternatively, a binary tree structure may be applied first. The encoding process according to this disclosure can be performed based on the final coding unit that is no longer segmented. In this case, the maximum coding unit may be used as the final coding unit based on image characteristics, coding efficiency, etc., or, if necessary, the coding unit may be recursively segmented into deeper coding units, and the coding unit with the optimal size may be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes (described later). As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be separated or partitioned from the aforementioned final encoding unit. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from the transform coefficients.

[0055] Encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to converter 232. In this case, as shown, the unit in encoder 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called subtractor 231. The predictor can perform prediction on the processing target block (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction in the unit of the current block or CU. As described later in the description of the various prediction modes, the predictor can generate various types of information about the prediction (e.g., prediction mode information) and send the generated information to entropy encoder 240. The information about the prediction can be encoded by entropy encoder 240 and output as a bitstream.

[0056] Intra-predictor 222 can refer to samples in the current image to predict the current block. Depending on the prediction mode, the referenced samples may be located near or separated from the current block. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, non-directional modes may include DC mode and planar mode. For example, depending on the level of detail in the prediction direction, the directional modes may include 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-predictor 222 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.

[0057] Inter-frame predictor 221 can derive the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information of neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.

[0058] Predictor 220 can generate prediction signals based on various prediction methods described later. For example, predictor 220 can apply intra-frame prediction or inter-frame prediction to predict a block, and can apply both intra-frame and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Furthermore, the predictor can use an intra-block copy (IBC) prediction mode or a palette mode for predicting blocks. The IBC prediction mode or palette mode can be used for image / video coding of content such as games, for example, screen content coding (SCC). IBC essentially performs prediction in the current image, but it can be performed similarly to inter-frame prediction in that it derives a reference block in the current image. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying a palette mode, sample values ​​in the image can be signaled based on information about the palette table and palette index.

[0059] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate the reconstructed signal or the residual signal.

[0060] Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform technique may include at least one of the following: Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graphic when the relationship information between pixels is illustrated as a graphic. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Additionally, the transform processing can be applied to pixel blocks that are squares of the same size, or it can be applied to blocks of variable size that are not squares.

[0061] Quantizer 233 quantizes the transform coefficients and sends the quantized transform coefficients to entropy encoder 240, which encodes the quantized signal (information about the quantized transform coefficients) and outputs the encoded signal as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the quantized transform coefficients in block form as a one-dimensional vector based on the coefficient scan order, and also generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0062] The entropy encoder 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can also encode, either together or separately, information necessary for video / image reconstruction (e.g., values ​​of syntax elements, etc.) other than the quantized transform coefficients. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream at the Network Abstraction Layer (NAL). The video / image information can also include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Additionally, the video / image information can also include general constraint information. In this document, information and / or syntax elements signaled from / transmitted from the encoding device to the decoding device can be included in the video / image information. The video / image information can be encoded by the aforementioned encoding process and thus included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit (not shown) for transmitting the signal output from the entropy encoder 240 and / or the storage unit (not shown) for storing the signal may be configured as internal / external components of the encoding device 200, or the transmitting unit may also be included in the entropy encoder 240.

[0063] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using dequantizer 234 and inverse transform unit 235. Adder 250 can add the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). For example, when skip mode is applied, the prediction block can be used as a reconstructed block when there is no residual for the processing target block. Adder 250 can be referred to as a restorer or restore block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current image, and can also be used for inter-frame prediction of the next image after filtering, as described below.

[0064] Additionally, Luminance Mapping and Chromaticity Scaling (LMCS) can be applied during image encoding and / or reconstruction processing.

[0065] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 270 (specifically, the DPB of memory 270). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various types of filtering-related information and transmit the generated information to entropy encoder 240, as described later in the description of the various filtering methods. The filtering-related information can be encoded by entropy encoder 240 and output as a bitstream.

[0066] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. When inter-frame prediction is applied via the encoding device, prediction mismatch between the encoding device 200 and the decoding device can be avoided, and encoding efficiency can be improved.

[0067] The DPB of memory 270 can store the corrected reconstructed image for use as a reference image in inter-frame predictor 221. Memory 270 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of blocks in already reconstructed images. The stored motion information can be transmitted to inter-frame predictor 221 to be used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and can transmit these reconstructed samples to intra-frame predictor 222.

[0068] Figure 3 This is a diagram illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure may be applied.

[0069] Reference Figure 3 The decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 may be configured by hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.

[0070] When the input includes a bitstream containing video / image information, the decoding device 300 can reconstruct and... Figure 2The encoding device processes video / image information corresponding to the image. For example, the decoding device 300 can derive units / blocks based on block partitioning information obtained from the bitstream. The decoding device 300 can perform decoding using processing units applied in the encoding device. Therefore, for example, the decoding processing unit can be an encoding unit, and the encoding unit can be segmented from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units can be derived from the encoding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.

[0071] Decoding device 300 can receive data in bitstream form from... Figure 2 The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets, such as adaptation parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements that are signaled / received, as described later in this document, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information within the bitstream based on encoding methods such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC), and output the syntax elements required for image reconstruction and the quantized values ​​of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model by using information about the target syntax element, decoding information about the target block, or information about symbols / bins decoded in a previous stage, and perform arithmetic decoding on the bins by predicting the probability of bin occurrence based on the determined context model, generating symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. Information related to prediction from the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values ​​(i.e., quantized transform coefficients and related parameter information) from the entropy decoding performed by the entropy decoder 310 can be input to the residual processor 320.

[0072] The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, filtering information from the information decoded by the entropy decoder 310 can be provided to the filter 350. Meanwhile, a receiver (not shown) for receiving signals output from the encoding device can be configured as an internal / external component of the decoding device 300, or the receiver can be a component of the entropy decoder 310. Furthermore, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the following: a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0073] Dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can use quantization parameters (e.g., quantization step size information) to perform dequantization on the quantized transform coefficients and obtain the transform coefficients.

[0074] The inverse transformer 322 performs inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0075] Predictor 330 can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from entropy decoder 310 and determine a specific intra-frame / inter-frame prediction mode.

[0076] Predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can apply intra-frame prediction or inter-frame prediction to predict a block, and can apply both intra-frame and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. IBC prediction mode or palette mode can be used for image / video coding of content such as games, for example, screen content coding (SCC). IBC can essentially perform prediction within the current frame, but can be performed similarly to inter-frame prediction, such that a reference block is derived within the current frame. That is, IBC can use at least one inter-frame prediction technique described in this document. Palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.

[0077] Intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples may be located near the current block or separated from it. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Intra-predictor 331 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.

[0078] Inter-frame predictor 332 can derive the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks may include spatially adjacent blocks existing in the current image and temporally adjacent blocks existing in the reference image. For example, inter-frame predictor 332 can construct a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode used for the current block.

[0079] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331). If there is no residual for the target block, such as when a skip mode is applied, the prediction block can be used as the reconstruction block.

[0080] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and as described later, it can also be output by filtering or used for inter-frame prediction of the next image.

[0081] In addition, Luminance Mapping with Chroma Scaling (LMCS) can also be applied to image decoding processing.

[0082] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a corrected reconstructed image by applying various filtering methods to the reconstructed image and store the corrected reconstructed image in memory 360, specifically in the DPB of memory 360. Various filtering methods may include, for example, unblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc.

[0083] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks from which motion information within the current image is derived (decoded) and / or motion information of blocks within reconstructed images. The stored motion information can be transmitted to inter-frame predictor 260 to be used as motion information for spatially adjacent blocks or temporally adjacent blocks. Memory 360 can store reconstructed samples of reconstructed blocks within the current image and transmit the reconstructed samples to intra-frame predictor 331.

[0084] In this document, the embodiments described in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding device 200 can be equally applied to or correspond to the filter 350, inter-frame predictor 332 and intra-frame predictor 331.

[0085] The video / image coding method described in this document can be executed based on the following partitioning structure. Specifically, the prediction, residual processing (inverse transform and dequantization), syntax element encoding, and filtering processes, which will be described later, can be performed based on the CTU and CU (and / or TU and PU) derived from the partitioning structure. The block partitioning process can be performed by the image partitioner 210 of the aforementioned encoding device, and the partitioning-related information can be processed (encoded) by the entropy encoder 240 and can be sent to the decoding device in the form of a bitstream. The entropy decoder 310 of the decoding device can derive the block partitioning structure of the current image based on the partitioning-related information obtained from the bitstream, and based on this, a series of processes for image decoding (e.g., prediction, residual processing, block / image reconstruction, and in-loop filtering) can be performed. The CU size and TU size can be equal to each other, or multiple TUs can exist in the CU region. Meanwhile, the CU size can generally represent the size of the luma component (sample) coding block (CB). The TU size can generally represent the size of the luma component (sample) transform block (TB). The chroma component (sample) CB or TB size can be derived based on the color format (chroma format, e.g., 4:4:4, 4:2:2, 4:2:0, etc.) of the image / picture according to the component ratio and the luminance component (sample) CB or TB size. The TU size can be derived based on maxTbSize. For example, if the CU size is larger than maxTbSize, multiple TUs (TBs) of maxTbSize can be derived, and transform / inverse transform can be performed on a TU (TB) basis. Furthermore, for example, in the case of applying intra-frame prediction, the intra-frame prediction mode / type can be derived on a CU (or CB) basis, and the derivation of neighboring reference samples and the generation of prediction samples can be performed on a TU (or TB) basis. In this case, one or more TUs (or TBs) can exist within a CU (or CB) region, and multiple TUs (or TBs) can share the same intra-frame prediction mode / type.

[0086] Furthermore, when encoding video / images according to this document, the image processing unit can have a hierarchical structure. An image can be divided into one or more tiles, patches, slices, and / or tile groups. A slice can include one or more patches. A patch can include one or more CTU rows within a tile. A slice can include an integer number of patches in the image. A tile group can include one or more tiles. A tile is a rectangular area of ​​CTUs within a specific tile column and a specific tile row in the image. A tile group can include an integer number of tiles according to a tile raster scan in the image. The slice header can carry information / parameters that can be applied to the corresponding slice (the blocks within the slice). If the encoding / decoding device has a multi-core processor, the encoding / decoding processes for tiles, slices, patches, and / or tile groups can be processed in parallel. In this document, slices or tile groups can be used interchangeably. That is, the tile group header can be referred to as the slice header. Here, a slice can have one of the slice types, including intra-frame (I) slices, prediction (P) slices, and bidirectional prediction (B) slices. For prediction of blocks in I-slices, inter-frame prediction may not be used, but intra-frame prediction only can be used. Even in this case, the original sample values ​​can be encoded and signaled without prediction. For blocks in P-slices, either intra-frame or inter-frame prediction can be used, and if inter-frame prediction is used, unidirectional prediction only can be used. Meanwhile, for blocks in B-slices, either intra-frame or inter-frame prediction can be used, and if inter-frame prediction is used, up to bidirectional prediction can be used.

[0087] Depending on the characteristics of the video image (e.g., resolution), or considering coding efficiency or parallel processing, the encoder can determine the patch / patch group, tile, slice, maximum and minimum coding unit size, and can include the corresponding information or information that can summarize the corresponding information in the bitstream.

[0088] The decoder can obtain information indicating whether tiles / tile groups, patches, slices, or CTUs in the current image have been partitioned into multiple coding units. Efficiency can be improved by obtaining (sending) this information only under specific conditions.

[0089] Figure 4 An exemplary hierarchical structure for encoding images / videos is shown.

[0090] refer to Figure 4 The encoded image / video is divided into the VCL (Video Coding Layer), which handles the image / video decoding process and itself; the subsystem that sends and stores encoded information; and the Network Abstraction Layer (NAL), which exists between the VCL and the subsystem and is responsible for network adaptation.

[0091] VCL can generate VCL data that includes compressed image data (slice data), or it can generate supplementary enhancement information (SEI) messages that are additionally necessary for the decoding process of parameters such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS).

[0092] In NAL, NAL cells can be generated by adding header information (NAL cell header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc., generated in VCL. The NAL cell header can include NAL cell type information specified according to the RBSP data included in the corresponding NAL cell.

[0093] As shown in the figure, NAL units can be divided into VCL NAL units and non-VCL NAL units based on the RBSP generated in the VCL. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data), while a non-VCL NAL unit can refer to a NAL unit that contains information necessary for decoding the image (parameter set or SEI message).

[0094] By attaching header information according to the subsystem's data standard, the aforementioned VCL NAL units and non-VCL NAL units can be transmitted over the network. For example, NAL units can be transformed into predetermined standard data formats such as H.266 / VVC file format, Real-time Transport Protocol (RTP), Transport Stream (TS), etc., and transmitted over various networks.

[0095] As described above, in a NAL cell, the NAL cell type can be specified according to the RBSP data structure included in the corresponding NAL cell, and information about this NAL cell type can be stored in the NAL cell header and signaled.

[0096] For example, depending on whether the NAL unit includes information about the image (slice data), NAL units can be roughly classified into VCL NAL unit types and non-VCL NAL unit types. VCL NAL unit types can be classified based on the characteristics and type of the image included in the VCL NAL unit, and non-VCL NAL unit types can be classified based on the type of parameter set.

[0097] The following is an example of a NAL cell type specified based on a parameter set type included in a non-VCL NAL cell type.

[0098] - APS (Adaptive Parameter Set) NAL Unit: A type of NAL unit that includes APS.

[0099] - DPS (Decoding Parameter Set) NAL Unit: The type used for NAL units that include DPS.

[0100] - VPS (Video Parameter Set) NAL Unit: The type used for NAL units that include VPS.

[0101] - SPS (Sequence Parameter Set) NAL Unit: The type used for NAL units that include SPS.

[0102] - PPS (Picture Parameter Set) NAL Unit: The type used for NAL units that include PPS.

[0103] - PH (Picture Header) NAL Unit: The type used for NAL units that include PH.

[0104] The aforementioned NAL unit types have syntax information for NAL unit types, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified through the nal_unit_type value.

[0105] Furthermore, as mentioned above, an image can include multiple slices, and a slice can include a slice header and slice data. In this case, an image header can be further added to the multiple slices (a set of slice headers and slice data) in an image. The image header (image header syntax) can include information / parameters that can be commonly applied to the image. The slice header (slice header syntax) can include information / parameters that can be commonly applied to the slice. The Adaptive Parameter Set (APS) or Picture Parameter Set (PPS) can include information / parameters that can be commonly applied to one or more images. The Sequence Parameter Set (SPS) can include information / parameters that can be commonly applied to one or more sequences. The Video Parameter Set (VPS) can include information / parameters that can be commonly applied to multiple layers. The Decoding Parameter Set (DPS) can include information / parameters that can be commonly applied to the overall video. The DPS can include information / parameters related to the concatenation of encoded video sequences (CVS).

[0106] In this document, high-level syntax may include at least one of the following: APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, image header syntax, and slice header syntax.

[0107] Furthermore, for example, information regarding the division and configuration of tiles / tile groups / tiles / slices can be configured by the encoding end through higher-level syntax and can be sent to the decoding device in the form of a bitstream.

[0108] In this paper, the image / video information encoded as a bitstream from the encoding device and signaled to the decoding device can include not only intra-image partition information, intra / inter-frame prediction information, residual information, and intra-loop filtering information, but also information included in the slice header, image header, APS, PPS, SPS, VPS, and / or DPS. Additionally, the image / video information may also include information from the NAL unit header.

[0109] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantization transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or for consistency, they may still be referred to as transform coefficients.

[0110] In this document, quantization transform coefficients and transform coefficients can be referred to as transform coefficients and scaling transform coefficients, respectively. In this context, residual information can include information about the transform coefficients, and this information can be signaled using residual coding syntax. Transform coefficients can be derived based on residual information (or information about the transform coefficients), and scaling transform coefficients can be derived by performing an inverse transform (scaling) on ​​the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaling transform coefficients. This can also be applied / expressed in other parts of this document.

[0111] As described above, the encoding device can perform various encoding methods, such as Exponential Golomb, Context Adaptive Variable-Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC). Furthermore, the decoding device can decode information in the bitstream based on encoding methods such as Exponential Golomb, CAVLC, or CABAC, and can output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the transform coefficients used for the residuals. For example, the aforementioned encoding methods can be performed as will be described later.

[0112] In this document, intra-frame prediction can be defined as generating prediction samples for the current block based on reference samples in the image to which the current block belongs (hereinafter, the current image). When applying intra-frame prediction to the current block, neighboring reference samples to be used for intra-frame prediction of the current block can be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary of the current block of size nW x nH and a total of 2 x nH samples adjacent to the left bottom, samples adjacent to the top boundary of the current block and a total of 2 x nW samples adjacent to the right top, and one sample adjacent to the left top of the current block. In addition, the neighboring reference samples of the current block may include multiple columns of top neighbor samples and multiple rows of left neighbor samples. Furthermore, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block of size nW x nH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the right bottom of the current block.

[0113] However, some neighboring reference samples of the current block may not have been decoded or enabled. In this case, the decoding device can configure neighboring reference samples to be used for prediction by replacing disabled samples with enabled ones. Furthermore, the neighboring reference samples to be used for prediction can be configured by interpolating enabled samples.

[0114] If neighboring reference samples are derived, then (i) the predicted sample can be inductively derived based on the average or interpolation of the neighboring reference samples of the current block, and (ii) the predicted sample can be inductively derived based on reference samples present in a specific (prediction) direction of the predicted sample among the neighboring reference samples of the current block. Case (i) can be referred to as non-directional mode or non-angular mode, and case (ii) can be referred to as directional mode or angular mode. Furthermore, the predicted sample can be generated by interpolating the first neighboring sample and the second neighboring sample among the neighboring reference samples, based on the current block's predicted sample being located in a direction opposite to the prediction direction of the intra-prediction mode of the current block. This case can be referred to as linear interpolation intra-prediction (LIP). Furthermore, chroma predicted samples can be generated based on luminance samples using a linear model. This case can be referred to as LM mode. Additionally, a temporary predicted sample of the current block can be derived based on filtered neighboring reference samples, and the predicted sample of the current block can be derived by calculating a weighted sum of the temporary predicted sample and at least one reference sample derived according to the intra-prediction mode (i.e., an unfiltered neighboring reference sample) among the existing neighboring reference samples. The above situation can be called orientation-dependent intra-prediction (PDPC). Furthermore, prediction samples can be derived by selecting the appropriate line and using reference samples located in the prediction direction on the reference sample line with the highest prediction accuracy among the neighboring multi-reference sample lines of the current block. In this case, intra-prediction coding can be performed in the method of indicating (signaling) the reference sample line used to the decoding device. This situation can be called multi-reference line (MRL) intra-prediction or MRL-based intra-prediction. Additionally, intra-prediction can be performed by dividing the current block into vertical or horizontal sub-partitions based on the same intra-prediction mode, and neighboring reference samples can be derived and used on a sub-partition basis. That is, in this case, since the intra-prediction mode for the current block is applied equally to the sub-partitions, and neighboring reference samples are derived and used on a sub-partition basis, intra-prediction performance can be improved in some cases. Such a prediction method can be called intra-partition (ISP) or ISP-based intra-prediction. The above intra-prediction methods can be called intra-prediction types to distinguish them from intra-prediction modes. Intra-prediction types can be referred to by various terms such as intra-prediction techniques or additional intra-prediction modes. For example, an intra-prediction type (or additional intra-prediction mode) can include at least one of LIP, PDPC, MRL, or ISP mentioned above. A general intra-prediction method that excludes a specific intra-prediction type such as LIP, PDPC, MRL, or ISP can be called a normal intra-prediction type. Normal intra-prediction types can typically be applied when no specific intra-prediction type is applied, and prediction can be performed based on the intra-prediction modes mentioned above. Additionally, post-filtering can be performed on the derived prediction samples as needed.

[0115] Specifically, the intra-frame prediction process may include steps such as determining the intra-frame prediction mode / type, deriving neighboring reference samples, and deriving prediction samples based on the intra-frame prediction mode / type. Furthermore, a post-filtering step may be performed on the derived prediction samples, if necessary.

[0116] In addition to the prediction types mentioned above, affine linear weighted intra-prediction (ALWIP) can also be used. ALWIP can be referred to as linear weighted intra-prediction (LWIP), matrix weighted intra-prediction (MIP), or matrix-based intra-prediction. When applying MIP to the current block, i) by using neighboring reference samples to which an averaging process has been performed, ii) a matrix-vector multiplication process can be performed, and iii) if necessary, the prediction samples for the current block can be derived by further performing horizontal / vertical interpolation. The intra-prediction mode used for MIP can be configured differently from the LIP, PDPC, MRL, or ISP intra-prediction described above, or the intra-prediction mode can be used for normal intra-prediction. The intra-prediction mode used for MIP can be referred to as MIP intra-prediction mode, MIP prediction mode, or MIP mode. For example, depending on the intra-prediction mode used for MIP, the matrix and offset used for matrix-vector multiplication can be configured differently. Here, the matrix can be referred to as the (MIP) weighted matrix, and the offset can be referred to as the (MIP) offset vector or (MIP) bias vector.

[0117] In image / video coding, the images that make up an image / video can be encoded / decoded according to the decoding order. The image order can be set differently from the decoding order, corresponding to the output order of the decoded images. Based on this, backward prediction and forward prediction can also be performed in inter-frame prediction.

[0118] Figure 5 An example illustrating the image decoding process.

[0119] Figure 5 An example of a schematic chip decoding process is shown, illustrating an embodiment of which this document can be applied. Figure 5 In the middle, it can be on top. Figure 3 S500 is executed in the entropy decoder 310 of the decoding device described herein; S510 can be executed in the predictor 330; S520 can be executed in the residual processor 320; S530 can be executed in the adder 340; and S540 can be executed in the filter 350. S500 may include the information decoding process described herein; S510 may include the inter-frame / intra-frame prediction process described herein; S520 may include the residual processing process described herein; S530 may include the block / picture reconstruction process described herein; and S540 may include the in-loop filtering process described herein.

[0120] refer to Figure 5 , such as regarding Figure 3 As illustrated in the specification, the image decoding process can schematically include (through decoding) a process for obtaining image / video information from the bitstream S500, image reconstruction processes S510 to S530, and an in-loop filtering process S540 for reconstructing the image. The image reconstruction process can be performed based on residual samples and prediction samples obtained through the inter-frame / intra-frame prediction S510 and residual processing S520 (dequantization, inverse transform for quantization transform coefficients) processes described in this document. A modified reconstructed image can be generated through the in-loop filtering process on the reconstructed image generated by the image reconstruction process. This modified image can be output as a decoded image and can also be stored in the decoded image buffer or memory 360 of the decoding device and used as a reference image in subsequent inter-frame prediction processes for image decoding. Depending on the situation, the in-loop filtering process can be skipped, and in this case, the reconstructed image can be output as a decoded image and can also be stored in the decoded image buffer or memory 360 of the decoding device and used as a reference image in subsequent inter-frame prediction processes for image decoding. The in-loop filtering process S540 may include the deblocking filtering process, the Sample Adaptive Offset (SAO) process, the Adaptive Loop Filter (ALF) process, and / or the bilateral filtering process as described above, and may skip all or some of them. Furthermore, one or more of the deblocking filtering process, the Sample Adaptive Offset (SAO) process, the Adaptive Loop Filter (ALF) process, and the bilateral filtering process may be applied sequentially, or all of them may be applied sequentially. For example, after applying the deblocking filtering process to the reconstructed image, the SAO process may be performed on it. Alternatively, for example, after applying the deblocking filtering process to the reconstructed image, the ALF process may be performed on it. This can also be performed in the encoding device.

[0121] Figure 6 An example illustrating the image encoding process.

[0122] Figure 6 An example of a schematic image encoding process that can be applied to embodiments of this document is shown. Figure 6 In the middle, it can be on top. Figure 2 S600 is executed in the predictor 220 of the coding apparatus described herein; S610 may be executed in the residual processor 230; and S620 may be executed in the entropy encoder 240. S600 may include the inter-frame / intra-frame prediction process described herein; S610 may include the residual processing process described herein; and S620 may include the information encoding process described herein.

[0123] refer to Figure 6 For example, regarding Figure 2As illustrated in the description, the image encoding process can schematically include the process of generating a reconstructed image of the current image, the process of applying in-loop filtering to the reconstructed image (optional), and the process of encoding information used for image reconstruction (e.g., prediction information, residual information, partitioning information, etc.) and outputting it as a bitstream. The encoding device can derive (modified) residual samples from the quantization transform coefficients using dequantizer 234 and inverse transformer 235, and can generate a reconstructed image based on the (modified) residual samples and the prediction samples as the output of S600. The reconstructed image generated in this way can be the same as the reconstructed image generated in the decoding device described above. Similar to the case of the decoding device, a modified reconstructed image can be generated through the in-loop filtering process for the reconstructed image, which can be stored in the decoded image buffer or memory 270 and used as a reference image in the inter-frame prediction process of subsequent image encoding. As mentioned above, all or part of the in-loop filtering process can be skipped depending on the situation. When performing an in-loop filtering process, the filtering-related information (parameters) can be encoded in the entropy encoder 240 and output as a bit stream, and the decoding device can perform the in-loop filtering process in the same way as the encoding device based on the filtering-related information.

[0124] This in-loop filtering process reduces noise generated during image / video encoding, such as block artifacts and ringing artifacts, and improves subjective / objective visual quality. Furthermore, when the in-loop filtering process is performed in both the encoding and decoding devices, both devices can produce the same prediction results, improving the reliability of image encoding and reducing the amount of data to be transmitted for image encoding.

[0125] As described above, the image reconstruction process can be performed in both the encoding and decoding devices. Based on intra-frame prediction / inter-frame prediction for each block unit, reconstructed blocks can be generated, and a reconstructed image including these blocks can be generated. When the current image / slice / patch group is an I-type image / slice / patch group, blocks included in the current image / slice / patch group can be reconstructed based solely on intra-frame prediction. Conversely, when the current image / slice / patch group is a P-type or B-type image / slice / patch group, blocks included in the current image / slice / patch group can be reconstructed based on either intra-frame or inter-frame prediction. In this case, inter-frame prediction can be applied to some blocks in the current image / slice / patch group, and intra-frame prediction can be applied to some of the remaining blocks. The color components of the image can include luma and chroma components, and unless explicitly limited by this document, the methods and embodiments presented herein can be applied to both luma and chroma components.

[0126] Meanwhile, the video / image coding process based on inter-frame prediction can schematically include, for example, the following.

[0127] Figure 7 This represents an example of a video / image coding method based on inter-frame prediction.

[0128] refer to Figure 7 The encoding device performs inter-frame prediction on the current block (S700). The encoding device can derive the inter-frame prediction mode and motion information for the current block, and generate prediction samples for the current block. Here, the inter-frame prediction mode determination, motion information derivation, and prediction sample generation processes can be performed simultaneously or sequentially. For example, the inter-frame predictor of the encoding device may include a prediction mode determiner, a motion information deriver, and a prediction sample deriver. The prediction mode determiner can determine the prediction mode for the current block, the motion information deriver can derive the motion information for the current block, and the prediction sample deriver can derive the prediction samples for the current block. For example, the inter-frame predictor of the encoding device can search for blocks similar to the current block in a certain region (search region) of the reference image through motion estimation, and derive a reference block whose difference from the current block is the smallest or less than or equal to a certain level. Based on this, a reference image index indicating the reference block is located in the reference image above can be derived, and a motion vector can be derived based on the positional difference between the reference block and the current block. The encoding device can determine the mode from various prediction modes applied to the current block. The encoding device can compare the rate distortion (RD) costs of various prediction modes and determine the best prediction mode for the current block.

[0129] For example, when applying a skip mode or merge mode to the current block, the encoding device can construct a merge candidate list and derive a reference block from the reference blocks included in the merge candidate list whose difference from the current block is the smallest or less than or equal to a certain level. In this case, a merge candidate associated with the derived reference block can be selected, and merge index information indicating the selected merge candidate can be generated and signaled to the decoding device. The motion information of the current block can be derived using the motion information of the selected merge candidate.

[0130] As another example, when applying the (A)MVP mode to the current block, the encoding device can construct an (A)MVP candidate list and use the motion vector of the MVP (motion vector predictor) candidate selected from the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, the motion vector of the reference block derived by the motion estimation described above can be used as the motion vector of the current block, and the MVP candidate with the motion vector having the smallest difference from the motion vector of the current block can be the selected MVP candidate. The MVD (motion vector difference) can be derived, which is the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, information about the MVD can be signaled to the decoding device. Additionally, when applying the (A)MVP mode, the value of the reference picture index can be configured as reference picture index information and signaled separately to the decoding device.

[0131] The encoding device can derive residual samples based on the predicted samples (S710). The encoding device can derive residual samples by comparing the original samples and the predicted samples of the current block.

[0132] The encoding device encodes image information including prediction information and residual information (S720). The encoding device can output the encoded image information in the form of a bitstream. The prediction information may include prediction mode information (e.g., skip flag, merge flag, mode index, etc.) and information about motion as information about the prediction process. The information about motion information may include candidate selection information (e.g., merge index, MVP flag, or MVP index), which is information used to derive motion vectors. In addition, the information about motion information may include information about the aforementioned MVD and / or reference image index information. Furthermore, the information about motion information may include information indicating whether L0 prediction, L1 prediction, or bidirectional prediction is applied. The residual information is information about residual samples. The residual information may include information about the quantization transform coefficients used for the residual samples.

[0133] The output bitstream can be stored in (digital) storage media and transmitted to the decoding device, or it can be transmitted to the decoding device via a network.

[0134] Simultaneously, as mentioned above, the encoding device can generate a reconstructed image (including reconstructed samples and reconstructed blocks) based on reference samples and residual samples. This is to derive the same prediction result in the encoding device as the prediction result performed in the decoding device, and the reason is that this can improve encoding efficiency. Therefore, the encoding device can store the reconstructed image (or reconstructed samples, reconstructed blocks) in memory and use it as a reference image for inter-frame prediction. In-loop filtering processes, etc., can be further applied to the reconstructed image as described above.

[0135] The video / image decoding process based on inter-frame prediction can schematically include, for example, the following.

[0136] Figure 8 This represents an example of a video / image decoding method based on inter-frame prediction.

[0137] The decoding device can perform operations corresponding to those already performed in the encoding device. The decoding device can perform predictions on the current block and derive prediction samples based on the received prediction information.

[0138] Specifically, refer to Figure 8 The decoding device can determine the prediction mode for the current block based on the prediction information received from the bitstream (S800). The decoding device can determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.

[0139] For example, it can be determined whether to apply a merge mode to the current block, or the (A)MVP mode can be determined based on a merge flag. Alternatively, an inter-frame prediction mode can be selected from a variety of inter-frame prediction mode candidates based on a merge index. Inter-frame prediction mode candidates can include various inter-frame prediction modes, such as skip mode, merge mode, and / or (A)MVP mode.

[0140] The decoding device derives motion information for the current block based on the determined inter-frame prediction mode (S810). For example, when a skip mode or merge mode is applied to the current block, the decoding device can construct a merge candidate list (described later) and select one of the merge candidates included in that list. The selection can be performed based on the selection information (merge index) described above. The motion information of the current block can be derived using the motion information of the selected merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0141] As another example, when applying (A)MVP mode to the current block, the decoding device can construct an (A)MVP candidate list and use the motion vector of the MVP (Motion Vector Predictor) candidate selected from the MVP (Motion Vector Predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. Selection can be performed based on the aforementioned selection information (MVP flag or MVP index). In this case, the MVD of the current block can be derived based on information about the MVD, and the motion vector of the current block can be derived based on the MVD and MVP of the current block. Furthermore, the reference image index of the current block can be derived based on reference image index information. Images in the reference image list related to the current block indicated by the reference image index can be exported as reference images used for inter-frame prediction of the current block.

[0142] Meanwhile, the motion information of the current block can be exported without constructing a candidate list, and in this case, the construction of the candidate list as described above can be omitted.

[0143] The decoding device can generate a prediction sample for the current block based on the motion information of the current block (S820). In this case, a reference image can be derived based on the reference image index of the current block, and the prediction sample for the current block can be derived using samples of the reference block on the reference image indicated by the motion vector of the current block. In this case, a prediction sample filtering process can be further performed on all or some of the prediction samples of the current block as described later.

[0144] For example, the inter-frame predictor of the encoding device may include a prediction mode determiner, a motion information deriver, and a prediction sample deriver. It may determine the prediction mode for the current block based on the prediction mode information received at the prediction mode determiner, derive the motion information (motion vector and / or reference picture index, etc.) of the current block based on the motion information received at the motion information deriver, and derive the prediction samples of the current block at the prediction sample deriver.

[0145] The decoding device generates residual samples for the current block based on the received residual information (S830). The decoding device can generate reconstructed samples for the current block based on the residual samples and the predicted samples, and generate a reconstructed image based on these reconstructed samples (S840). In the following text, an in-loop filtering process can be applied to the reconstructed image as described above.

[0146] Simultaneously, as described above, High-Level Syntax (HLS) can be encoded / signaled for use in video / image encoding. An encoded picture can consist of one or more slices. The parameters of the encoded picture are described using signaled notations in the picture header, and the parameters of the slice are described using signaled notations in the slice header. The picture header is carried in the form of the NAL unit itself. The slice header exists at the beginning of the NAL unit, which includes the slice's payload (i.e., slice data).

[0147] Each image is associated with an image header. Images can be composed of different types of slices: intra-coded slices (i.e., I-slices) and inter-coded slices (i.e., P-slices and B-slices). Therefore, the image header can include the necessary syntax elements for the intra-coded and inter-coded slices of the image. For example, the syntax of the image header can be as shown in Table 1 below.

[0148] [Table 1]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157] Among the syntax elements in Table 1, syntax elements whose titles include “intra_slice” (e.g., pic_log2_diff_min_qt_min_cb_intra_slice_luma) are syntax elements used in the I slice of the corresponding image, and syntax elements related to syntax elements whose titles include “inter_slice” (e.g., pic_log2_diff_min_qt_min_cb_inter_slice, mvp, mvd, mmvd, and merge) (e.g., pic_temporal_mvp_enabled_flag) are syntax elements used in the P slice and / or B slice of the corresponding image.

[0158] In other words, the picture header includes all the syntax elements necessary for intra-coded slices and inter-coded slices for each individual picture. However, this is only useful relative to pictures that include mixed-type slices (pictures that include all intra-coded and inter-coded slices). In general, since pictures do not include mixed-type slices (i.e., general pictures include either intra-coded slices only or inter-coded slices only), it is not necessary to perform signaling for all data (syntax elements used in intra-coded slices and syntax elements used in inter-coded slices).

[0159] The following figures are provided to illustrate detailed examples in this document. Since the names of detailed devices or detailed signals / information are presented exemplarily, the technical features of this document are not limited to the detailed names used in the following figures.

[0160] This document provides the following methods to solve the above problems. Each method can be applied individually or in combination.

[0161] 1. A signal can be used to indicate the presence of a flag in the image header that specifies the syntax element required for intra-coded slices. This flag can be called the intra_signaling_present_flag.

[0162] a) When intra_signaling_present_flag equals 1, the syntax elements required for intra-coded slices are present in the image header. Similarly, when intra_signaling_present_flag equals 0, the syntax elements required for intra-coded slices are not present in the image header.

[0163] (b) When the picture associated with the picture header has at least one intra-coded slice, the value of intra_signaling_present_flag in the picture header should be equal to 1.

[0164] c) Even when the picture associated with the picture header does not have an intra-coded slice, the value of intra_signaling_present_flag in the picture header can be equal to 1.

[0165] d) When an image has one or more sub-images that contain only intra-coded slices and it is expected that one or more sub-images can be extracted and merged with a sub-image that contains one or more inter-coded slices, the value of intra_signaling_present_flag should be set to 1.

[0166] 2. A signal can be used to indicate the presence of a flag in the image header that specifies the syntax element required for inter-frame-only coded slices. This flag can be called `inter_signaling_present_flag`.

[0167] a) When inter_signaling_present_flag equals 1, the syntax elements required for inter-frame coding slicing are present in the image header. Similarly, when inter_signaling_present_flag equals 0, the syntax elements required for inter-frame coding slicing are not present in the image header.

[0168] b) When the image associated with the image header has at least one inter-frame coded slice, the value of inter_signaling_present_flag in the image header should be equal to 1.

[0169] c) Even when the picture associated with the picture header does not have an inter-frame coded slice, the value of inter_signaling_present_flag in the picture header can be equal to 1.

[0170] d) When an image has one or more sub-images that contain only inter-coded slices and it is expected that one or more sub-images can be extracted and merged with a sub-image that contains one or more intra-coded slices, the value of inter_signaling_present_flag should be set to 1.

[0171] 3. The aforementioned flags (intra_signaling_present_flag and inter_signaling_present_flag) can be signaled in other parameter sets, such as the Picture Parameter Set (PPS), instead of in the Picture header.

[0172] 4. Another alternative for using signals to notify the above-mentioned signs may be as follows.

[0173] a) You can define two variables, IntraSignalingPresentFlag and InterSignalingPresentFlag, to specify whether the syntax elements required for intra-coded slices and inter-coded slices exist in the image header, respectively.

[0174] (b) A flag called mixed_slice_types_present_flag in the image header can be signaled. When mixed_slice_types_present_flag equals 1, the values ​​of IntraSignalingPresentFlag and InterSignalingPresentFlag are set to 1.

[0175] c) When `mixed_slice_types_present_flag` equals 0, an additional flag called `intra_slice_only_flag` can be signaled in the image header, and this applies below. If `intra_slice_only_flag` equals 1, then the value of `IntraSignalingPresentFlag` is set to 1 and the value of `InterSignalingPresentFlag` is set to 0. Otherwise, the value of `IntraSignalingPresentFlag` is set to 0 and the value of `InterSignalingPresentFlag` is set to 1.

[0176] 5. A fixed-length syntax element in the image header can be signaled using a signal called slice_types_idc, specifying the following information.

[0177] a) Whether the image associated with the image header contains intra-frame encoded slices. For this type, the value of slice_types_idc can be set to 0.

[0178] b) Whether the image associated with the image header contains slices that are only inter-frame encoded. The value of slice_types_idc can be set to 1.

[0179] c) Whether the image associated with the image header can contain intra-coded slices and inter-coded slices. The value of slice_types_idc can be set to 2.

[0180] Note that when slice_types_idc has a value of 2, it is still possible for the image to contain slices that are coded only within the frame or slices that are coded only between the frames.

[0181] d) Other values ​​of slice_types_idc can be reserved for future use.

[0182] 6. For the slice_types_idc semantics in the image header, the following constraints can be further specified.

[0183] a) When the picture associated with the picture header has one or more intra-coded slices, the value of slice_types_idc should not be equal to 1.

[0184] b) When the picture associated with the picture header has one or more inter-frame coded slices, the value of slice_types_idc should not be equal to 0.

[0185] 7. The slice_types_idc can be notified by a signal in other parameter sets such as the Picture Parameter Set (PPS) instead of in the picture header.

[0186] As an example, the encoding and decoding devices may use Tables 2 and 3 below as the syntax and semantics of the image header based on methods 1 and 2 as described above.

[0187] [Table 2]

[0188]

[0189]

[0190]

[0191] [Table 3]

[0192]

[0193] Referring to Tables 2 and 3, if the value of `intra_signaling_present_flag` is 1, it indicates that syntax elements used only in intra-coded slices are present in the picture header. If the value of `intra_signaling_present_flag` is 0, it indicates that syntax elements used only in intra-coded slices are not present in the picture header. Therefore, if the picture associated with the picture header includes one or more slices of the slice type with I slices, the value of `intra_signaling_present_flag` becomes 1. Furthermore, if the picture associated with the picture header does not include slices of the slice type with I slices, the value of `intra_signaling_present_flag` becomes 0.

[0194] If the value of `inter_signaling_present_flag` is 1, it indicates that syntax elements used only in inter-frame encoded slices are present in the picture header. If the value of `inter_signaling_present_flag` is 0, it indicates that syntax elements used only in inter-frame encoded slices are not present in the picture header. Therefore, if the picture associated with the picture header includes one or more slices of slice type P-slice and / or B-slice, the value of `intra_signaling_present_flag` becomes 1. Furthermore, if the picture associated with the picture header does not include slices of slice type P-slice and / or B-slice, the value of `intra_signaling_present_flag` becomes 0.

[0195] Furthermore, in cases where the image includes one or more sub-images that can be merged with one or more sub-images that include slices with inter-frame coding, the values ​​of intra_signaling_present_flag and inter_signaling_present_flag are both set to 1.

[0196] For example, if the current image includes slices that are only inter-frame encoded (P slices and / or B slices), the encoding device can set the value of inter_signaling_present_flag to 1 and the value of intra_signaling_present_flag to 0.

[0197] As another example, in the case where the current image includes intra-frame encoded slices (I slices), the encoding device can set the value of inter_signaling_present_flag to 0 and the value of intra_signaling_present_flag to 1.

[0198] As another example, if the current image includes at least one inter-coded slice or at least one intra-coded slice, the encoding device may set the values ​​of inter_signaling_present_flag and intra_signaling_present_flag to a total of 1.

[0199] When the value of `intra_signaling_present_flag` is set to 0, the encoding device can generate image information in which the syntax elements necessary for intra-frame slices are excluded or omitted, and only the syntax elements necessary for inter-frame slices are included in the image header. If the value of `inter_signaling_present_flag` is set to 0, the encoding device can generate image information in which the syntax elements necessary for inter-frame slices are excluded or omitted, and only the syntax elements necessary for intra-frame slices are included in the image header.

[0200] If the value of `inter_signaling_present_flag` obtained from the image header in the image information is 1, the decoding device can determine that at least one inter-coded slice is included in the corresponding image, and can parse the syntax elements necessary for intra-frame prediction from the image header. If the value of `inter_signaling_present_flag` is 0, the decoding device can determine that the corresponding image includes an intra-coded slice only, and can parse the syntax elements necessary for intra-frame prediction from the image header. If the value of `intra_signaling_present_flag` obtained from the image header in the image information is 1, the decoding device can determine that at least one intra-coded slice is included in the corresponding image, and can parse the syntax elements necessary for intra-frame prediction from the image header. If the value of `intra_signaling_present_flag` is 0, the decoding device can determine that the corresponding image includes an inter-coded slice only, and can parse the syntax elements necessary for inter-frame prediction from the image header.

[0201] As another embodiment, the encoding device and the decoding device may use Tables 4 and 5 below as the syntax and semantics of the image header based on the methods 5 and 6 described above.

[0202] [Table 4]

[0203]

[0204]

[0205]

[0206]

[0207] [Table 5]

[0208]

[0209] Referring to Tables 4 and 5, if the value of slice_types_idc is 0, it indicates that all slices in the image associated with the image header are of type I slices. If the value of slice_types_idc is 1, it indicates that all slices in the image associated with the image header are of type P or B slices. If the value of slice_types_idc is 2, it indicates that the slices in the image associated with the image header are of type I, P, and / or B slices.

[0210] For example, if the current image includes intra-frame coded slices only, the encoding device can set the value of slice_types_idc to 0 and include the syntax elements necessary for decoding intra-frame coded slices in the image header. That is, in this case, the syntax elements necessary for inter-frame coded slices are not included in the image header.

[0211] As another example, if the current image includes inter-frame-only coded slices, the encoding device can set the value of slice_types_idc to 1 and include the syntax elements necessary for decoding inter-frame-only slices in the image header. That is, in this case, the syntax elements necessary for intra-frame slices are not included in the image header.

[0212] As another example, if the current picture includes at least one inter-frame coded slice and at least one intra-frame coded slice, the encoding device may determine the value of slice_types_idc to be 2, and may include all of the syntax elements necessary for decoding the inter-frame slice and the syntax elements necessary for decoding the intra-frame slice in the picture header.

[0213] If the value of slice_types_idc obtained from the image header in the image information is 0, the decoding device can determine that the corresponding image includes intra-coded slices only, and can parse the necessary syntax elements for decoding the intra-coded slices from the image header. If the value of slice_types_idc is 1, the decoding device can determine that the corresponding image includes inter-coded slices only, and can parse the necessary syntax elements for decoding the inter-coded slices from the image header. If the value of slice_types_idc is 2, the decoding device can determine that the corresponding image includes at least one intra-coded slice and at least one inter-coded slice, and can parse the necessary syntax elements for decoding the intra-coded slices and the inter-coded slices from the image header.

[0214] As another embodiment, the encoding and decoding devices may use a flag indicating whether a picture includes intra-frame coded slices and inter-frame coded slices. If the flag is true, that is, if the value of the flag is 1, then all intra-frame slices and inter-frame slices can be included in the corresponding picture. In this case, Tables 6 and 7 below can be used as the syntax and semantics of the picture header.

[0215] [Table 6]

[0216]

[0217]

[0218]

[0219]

[0220] [Table 7]

[0221]

[0222] Referring to Tables 6 and 7, if the value of mixed_slice_signaling_present_flag is 1, it indicates that the image associated with the corresponding image header has one or more slices of different types. If the value of mixed_slice_signaling_present_flag is 0, it means that the image associated with the corresponding image header includes data that is only associated with a single slice type.

[0223] The variables InterSignalingPresentFlag and IntraSignalingPresentFlag indicate whether the syntax elements required for intra-coded slices and inter-coded slices are present in the corresponding picture header, respectively. If the value of mixed_slice_signaling_present_flag is 1, then the values ​​of IntraSignalingPresentFlag and InterSignalingPresentFlag are set to 1.

[0224] If the value of intra_slice_only_flag is set to 1, it means that the value of IntraSignalingPresentFlag is set to 1 and the value of InterSignalingPresentFlag is set to 0. If the value of intra_slice_only_flag is set to 0, it means that the value of IntraSignalingPresentFlag is set to 0 and the value of InterSignalingPresentFlag is set to 1.

[0225] If the image associated with the image header has one or more slices of type I, the value of IntraSignalingPresentFlag is set to 1. If the image associated with the image header has one or more slices of type P or B, the value of InterSignalingPresentFlag is set to 1.

[0226] For example, if the current image includes intra-frame encoded slices, the encoding device can set the value of mixed_slice_signaling_present_flag to 0, the value of intra_slice_only_flag to 1, the value of IntraSignalingPresentFlag to 1, and the value of InterSignalingPresentFlag to 0.

[0227] As another example, if the current image includes slices that are only encoded by inter-frames, the encoding device can set the value of mixed_slice_signaling_present_flag to 0, the value of intra_slice_only_flag to 0, the value of IntraSignalingPresentFlag to 0, and the value of InterSignalingPresentFlag to 1.

[0228] As another example, if the current image includes at least one intra-coded slice and at least one inter-coded slice, the encoding device may set the values ​​of mixed_slice_signaling_present_flag, IntraSignalingPresentFlag, and InterSignalingPresentFlag to 1.

[0229] If the value of `mixed_slice_signaling_present_flag` obtained from the image header in the image information is 0, the decoding device can determine whether the corresponding image includes an intra-coded-only slice or an inter-coded slice. In this case, if the value of `intra_slice_only_flag` obtained from the image header is 0, the decoding device can parse the syntax elements necessary for decoding an inter-coded-only slice from the image header. If the value of `intra_slice_only_flag` is 1, the decoding device can parse the syntax elements necessary for decoding an intra-coded-only slice from the image header.

[0230] If the value of mixed_slice_signaling_present_flag obtained from the image header in the image information is 1, the decoding device can determine that the corresponding image includes at least one intra-coded slice and at least one inter-coded slice, and can parse the syntax elements necessary for decoding the inter-coded slice and the syntax elements necessary for decoding the intra-coded slice from the image header.

[0231] Figure 9 and Figure 10 Examples of video / image encoding methods and related components according to embodiments of this document are illustrated schematically.

[0232] Figure 9 The publicly disclosed video / image coding methods can be derived from... Figure 2 and Figure 10 The (video / image) encoding device 200 disclosed herein performs the operation. Specifically, for example, Figure 9 S900 and S910 can be executed by the predictor 220 of the encoding device 200; S920 can be executed by the adder 250 of the encoding device 200; and S930 and S940 can be executed by the entropy encoder 240 of the encoding device 200. Figure 9 The video / image encoding methods disclosed herein may include the embodiments described above.

[0233] Specifically, refer to Figure 9 and Figure 10The predictor 220 of the encoding device can determine the prediction mode of the current block in the current image (S900). The current image may include multiple slices. The predictor 220 of the encoding device can generate prediction samples (prediction blocks) of the current block based on the prediction mode (S910). Here, the prediction mode may include inter-frame prediction mode and intra-frame prediction mode. When the prediction mode of the current block is inter-frame prediction mode, the prediction samples can be generated by the inter-frame predictor 221 of the predictor 220. When the prediction mode of the current block is intra-frame prediction mode, the prediction samples can be generated by the intra-frame predictor 222 of the predictor 220.

[0234] The residual processor 230 of the encoding device can generate residual samples and residual information based on the predicted samples and the original image (original block, original sample). Here, the residual information is information about the residual samples and may include information about the (quantization) transform coefficients used for the residual samples.

[0235] The adder (or reconstructor) of the encoding device can generate reconstructed samples (reconstructed picture, reconstructed block, reconstructed sample array) by adding the residual samples generated by the residual processor 230 and the predicted samples generated by the inter-frame predictor 221 or the intra-frame predictor 222 (S920).

[0236] Simultaneously, the entropy encoder 240 of the encoding device can generate at least one of the following based on the prediction mode: first information indicating whether information necessary for inter-frame prediction operations in the decoding process exists in the image header associated with the current image, or second information indicating whether information necessary for intra-frame prediction operations in the decoding process exists in the image header associated with the current image (S930). Here, the first and second information are information included in the image header of the image information, and may correspond to the aforementioned intra_signalling_present_flag, inter_signalling_present_flag, slice_type_idc, mixed_slice_signalling_present_flag, intra_slice_only_flag, IntraSignallingPresentFlag, and / or InterSignallingFlag.

[0237] As an example, when the image header associated with the current image includes information necessary for inter-frame prediction operations during the decoding process due to the inclusion of inter-frame coded slices in the current image, the entropy encoder 240 of the encoding device may determine the value of the first information to be 1. Furthermore, when the image header associated with the current image includes information necessary for intra-frame prediction operations during the decoding process due to the inclusion of intra-frame coded slices in the current image, the entropy encoder 240 of the encoding device may determine the value of the second information to be 1. In this case, the first information may correspond to `inter_signaling_present_flag`, and the second information may correspond to `intra_signaling_present_flag`. The first information may be referred to as a first flag, information regarding the presence of syntax elements for inter-frame slices in the image header, a flag regarding the presence of syntax elements for inter-frame slices in the image header, information regarding whether a slice in the current image is an inter-frame slice, or a flag regarding whether a slice is an inter-frame slice. The second information can be referred to as a second flag, information about whether the syntax element used for intra-frame slicing exists in the picture header, a flag indicating whether the syntax element used for intra-frame slicing exists in the picture header, information about whether a slice in the current picture is an intra-frame slice, or a flag indicating whether a slice is an intra-frame slice.

[0238] Simultaneously, when the image header includes information necessary for intra-frame prediction operations due to the inclusion of intra-frame-only coded slices, the entropy encoder 240 of the encoding device can determine the value of the first information as 0 and the value of the second information as 1. Furthermore, when the image header includes information necessary for inter-frame prediction operations due to the inclusion of inter-frame-only coded slices, the first information can be determined as 1 and the second information as 0. Accordingly, if the value of the first information is 0, all slices in the current image can have the I-slice type. If the value of the second information is 0, all slices in the current image can have the P-slice type or the B-slice type. Here, the information necessary for intra-frame prediction operations can include syntax elements for decoding intra-frame slices, while the information necessary for inter-frame prediction operations can include syntax elements for decoding inter-frame slices.

[0239] As another example, if all slices in the current image have an I slice type, the entropy encoder 240 of the encoding device can determine the value of the slice type information as 0; if all slices in the current image have a P slice type or a B slice type, the entropy encoder 240 of the encoding device can determine the value of the slice type information as 1. If all slices in the current image have an I slice type, a P slice type, and / or a B slice type (i.e., the slice types of the slices in the image are mixed), the entropy encoder 240 of the encoding device can determine the value of the slice type information as 2. In this case, the slice type information can correspond to slice_type_idc.

[0240] As another example, if all slices in the current image have the same slice type, the entropy encoder 240 of the encoding device can determine the value of the slice type information as 0, while if the slices in the current image have different slice types, the entropy encoder 240 of the encoding device can determine the value of the slice type information as 1. In this case, the slice type information can correspond to the mixed_slice_signaling_present_flag.

[0241] If the value of the slice type information is determined to be 0, information regarding whether an intra-slice is included in the slice can be included in the corresponding picture header. This information can correspond to the `intra_slice_only_flag`. If all slices in the picture are of type I slice, the entropy encoder 240 of the encoding device can determine the value of the information regarding whether an intra-slice is included in the slice to be 1, the value of the information regarding the presence of syntax elements for intra-slices in the picture header to be 1, and the value of the information regarding the presence of syntax elements for inter-slices in the picture header to be 0. If all slices in the picture are of type P slice and / or type B slice, the entropy encoder 240 of the encoding device can determine the value of the information regarding whether an intra-slice is included in the slice to be 0, the value of the information regarding the presence of syntax elements for intra-slices in the picture header to be 0, and the value of the information regarding the presence of syntax elements for inter-slices in the picture header to be 1.

[0242] The entropy encoder 240 of the encoding device can encode image information including the aforementioned first information, second information, slice type information, residual information, and prediction-related information (S940). For example, the image information may include partition-related information, information about the prediction mode, residual information, in-loop filtering-related information, the first information, the second information, slice type information, etc., and includes various syntax elements related to them. In the example, the image information may include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). Furthermore, the image information may include various information such as picture header syntax, picture header structure syntax, slice header syntax, and coding unit syntax. The aforementioned first information, second information, information about the slice type, information necessary for intra-frame prediction operations, and information necessary for inter-frame prediction operations can be included in the syntax of the picture header.

[0243] Information encoded by the entropy encoder 240 of the encoding device can be output as a bitstream. The bitstream can be sent to the decoding device via a network or storage medium.

[0244] Figure 11 and Figure 12 Examples of video / image decoding methods and related components according to embodiments of this document are illustrated schematically.

[0245] Figure 11 The publicly disclosed video / image decoding method can be used by Figure 3 and Figure 12 The (video / image) decoding device 300 disclosed herein performs this action. Specifically, for example, it may be performed in the entropy decoder 310 of the decoding device. Figure 11 S1100 and S1120 can be executed in the predictor 330 of the decoding device 300, and S1130 can be executed in the adder 340 of the decoding device 300. Figure 11 The video / image decoding methods disclosed herein may include the embodiments described above.

[0246] refer to Figure 11 and Figure 12 The entropy decoder 310 of the decoding device can obtain image information from the bitstream (S1100). The image information may include an image header associated with the current image. The current image may include multiple slices.

[0247] Simultaneously, the entropy decoder 310 of the decoding device can parse at least one of the following: a first flag indicating whether information necessary for inter-frame prediction operations in the decoding process exists in the image header associated with the current image; or a second flag indicating whether information necessary for intra-frame prediction operations in the decoding process exists in the image header associated with the current image (S1110). Here, the first and second flags may correspond to the aforementioned intra_signalling_present_flag, inter_signalling_present_flag, slice_type_idc, mixed_slice_signalling_present_flag, intra_slice_only_flag, IntraSignallingPresentFlag, and / or InterSignallingPresentFlag. The entropy decoder 310 of the decoding device can parse the syntax elements included in the image header of the image information based on the image header syntax of any one of Tables 2, 4, and 6 above.

[0248] The decoding device can generate a prediction sample by performing at least one of intra-frame prediction and inter-frame prediction on the current block in the current image based on a first flag, a second flag, and information about the slice type (S1120).

[0249] Specifically, the entropy decoder 310 of the decoding device can parse (or obtain) at least one of the information necessary for intra-frame prediction operations and / or information necessary for inter-frame prediction operations for the decoding process from the picture header associated with the current picture, based on a first flag, a second flag, and / or information about the slice type. The predictor 330 of the decoding device can generate prediction samples by performing intra-frame prediction and / or inter-frame prediction based on at least one of the information necessary for intra-frame prediction operations or the information necessary for inter-frame prediction. Here, the information necessary for intra-frame prediction operations may include syntax elements for decoding intra-frame slices, and the information necessary for inter-frame prediction operations may include syntax elements for decoding inter-frame slices.

[0250] As an example, if the value of the first flag is 0, the entropy decoder 310 of the decoding device can determine (or decide) that the syntax elements used for inter-frame prediction do not exist in the picture header, and can parse only the information necessary for intra-frame prediction operations from the picture header. If the value of the first flag is 1, the entropy decoder 310 of the decoding device can determine (or decide) that the syntax elements used for inter-frame prediction exist in the picture header, and can parse the information necessary for inter-frame prediction operations from the picture header. In this case, the first flag may correspond to the inter_signaling_present_flag.

[0251] Furthermore, if the value of the second flag is 0, the entropy decoder 310 of the decoding device can determine (or decide) that the syntax elements used for intra-frame prediction do not exist in the picture header, and can parse only the information necessary for inter-frame prediction operations from the picture header. If the value of the second flag is 1, the entropy decoder 310 of the decoding device can determine (or decide) that the syntax elements used for intra-frame prediction exist in the picture header, and can parse the information necessary for intra-frame prediction operations from the picture header. In this case, the second flag may correspond to the intra_signaling_present_flag.

[0252] If the value of the first flag is 0, the decoding device can determine that all slices in the current image have the type I slice. If the value of the first flag is 1, the decoding device can determine that 0 or more slices in the current image have the type P slice or B slice. In other words, if the value of the first flag is 1, slices with the type P slice or B slice can be included in the current image, or slices with the type P slice or B slice can be excluded from the current image.

[0253] Furthermore, if the value of the second flag is 0, the decoding device can determine that all slices in the current image are of type P slice or B slice. If the value of the second flag is 1, the decoding device can determine that zero or more slices in the current image are of type I slice. In other words, if the value of the second flag is 1, slices of type I slice can be included in the current image, or slices of type I slice can be excluded from the current image.

[0254] As another example, if the value of the slice type information is 0, the entropy decoder 310 of the decoding device can determine that all slices in the current image have the I slice type and can parse only the information necessary for intra-frame prediction operations. If the slice type information is 1, the entropy decoder 310 of the decoding device can determine that all slices in the corresponding image have the P slice type or the B slice type and can parse only the information necessary for inter-frame prediction operations from the image header. If the slice type information is 2, the entropy decoder 310 of the decoding device can determine that the slices in the corresponding image have a slice type that is a mixture of I slice type, P slice type, and / or B slice type, and can parse all of the information necessary for inter-frame prediction operations and the information necessary for intra-frame prediction operations from the image header. In this case, the slice type information can correspond to slice_type_idc.

[0255] As another example, the entropy decoder 310 of the decoding device can determine that all slices in the current image have the same slice type when the value of the slice type information is determined to be 0, and can determine that the slices in the current image have different slice types when the value of the slice type information is determined to be 1. In this case, the slice type information can correspond to the mixed_slice_signalling_present_flag.

[0256] If the value of the information regarding the slice type is determined to be 0, the entropy decoder 310 of the decoding device can parse information from the image header regarding whether an intra-slice is included in the slice. This information can correspond to the `intra_slice_only_flag` as described above. If the information regarding whether an intra-slice is included in the slice is 1, then all slices in the image can have the I-slice type.

[0257] If the value regarding whether intra-frame slices are included in the slice is 1, the entropy decoder 310 of the encoding device can parse only the information necessary for intra-frame prediction operations from the picture header. If the value regarding whether intra-frame slices are included in the slice is 0, the entropy decoder 310 of the decoding device can parse only the information necessary for inter-frame prediction operations from the picture header.

[0258] If the value of the information about the slice type is 1, then the entropy decoder 310 of the decoding device can parse all of the information required for inter-frame prediction operations and intra-frame prediction operations from the picture header.

[0259] Meanwhile, the residual processor 320 of the decoding device can generate residual samples based on the residual information obtained by the entropy decoder 310.

[0260] The adder 340 of the decoding device can generate reconstructed samples based on the predicted samples generated by the predictor 330 and the residual samples generated by the residual processor 320 (S1130). In addition, the adder 340 of the decoding device can generate a reconstructed image (reconstructed block) based on the reconstructed samples (S1140).

[0261] After this, in-loop filtering processes such as ALF, SAO, and / or deblocking filtering can be applied to the reconstructed images as needed to improve subjective / objective video quality.

[0262] Although the methods have been described in the above embodiments based on flowcharts in which steps or blocks are listed in sequence, the steps of this disclosure are not limited to a particular order, and a step may be performed in different steps, in different orders, or simultaneously with respect to the above order. Furthermore, those skilled in the art will understand that the steps in the flowcharts are not exclusive, and one or more steps in the flowcharts may be included or may be deleted without affecting the scope of this disclosure.

[0263] The methods mentioned above according to this disclosure can be in the form of software, and the encoding and / or decoding devices according to this disclosure can be included in an apparatus for performing image processing (e.g., TV, computer, smartphone, set-top box, display device, etc.).

[0264] When the embodiments of this disclosure are implemented in software, the methods described above can be implemented using modules (processes or functions) that perform the functions mentioned above. Modules can be stored in memory and executed by a processor. Memory can be installed internally or externally to the processor and can be connected to the processor via various known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, embodiments of this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information about the implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.

[0265] Furthermore, the decoding and encoding devices using embodiments of this disclosure can be included in multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, portable cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, image telephony video devices, vehicle-mounted terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, or ship terminals), and medical video devices; and can be used to process image signals or data. For example, OTT video devices can include game consoles, Blu-ray players, networked TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).

[0266] Furthermore, the processing methods applying embodiments of this disclosure can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data with data structures according to embodiments of this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Computer-readable recording media also include media implemented in the form of carrier waves (e.g., transmission over the Internet). Additionally, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks.

[0267] Furthermore, embodiments of this disclosure can be implemented as computer program products based on program code, and the program code can be executed on a computer according to embodiments of this document. The program code can be stored on a computer-readable medium.

[0268] Figure 13 Examples of content streaming systems to which embodiments of the present disclosure may be applied are provided.

[0269] refer to Figure 13 The content streaming system to which embodiments of the present disclosure are applied may generally include an encoding server, a streaming server, a web server, media storage, a user device, and a multimedia input device.

[0270] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generate a bitstream, and send it to a streaming server. As another example, if the multimedia input device, such as a smartphone, camera, or camcorder, directly generates the bitstream, the encoding server can be omitted.

[0271] Bitstreams can be generated using the encoding methods or bitstream generation methods applied in the embodiments of this disclosure. Furthermore, the streaming server can temporarily store the bitstream during transmission or reception.

[0272] A streaming server sends multimedia data to a user's device via a web server based on a user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server forwards the request to the streaming server, which then delivers the multimedia data to the user. In this respect, the content streaming system may include a separate control server, which in this case controls the commands / responses between the various devices within the content streaming system.

[0273] A streaming server can receive content from media storage devices and / or encoding servers. For example, if content is received from an encoding server, it can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to provide a smooth streaming service.

[0274] For example, user equipment may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smartwatches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.

[0275] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.

Claims

1. An image decoding method performed by a decoding device, the method comprising: Receive a bitstream including image information, wherein the image information includes an image header associated with the current image, and multiple slices are included in the current image; Parse the syntax element from the image header that specifies whether all slices in the current image are I-slices; Predictive samples are derived by performing at least one of intra-frame prediction or inter-frame prediction on blocks of the slices in the current image based on the syntax element specifying whether all slices in the current image are the I-slice; and Based on the exported predicted samples, reconstructed samples are generated. Specifically, based on the syntax elements specifying that one or more of the slices in the current image are P-slices or B-slices, the information for inter-frame slicing includes a syntax element representing the maximum hierarchical depth of the coding units generated by multi-type tree splits in the inter-frame slices of the current image, and a syntax element representing the difference between the base-2 logarithm of the minimum size generated by quadtree splits and the base-2 logarithm of the minimum coding block size in the inter-frame slices of the current image is included in the image header. Specifically, the syntax element specifies that all slices in the current image are I slices. The information for intra-frame slices includes a syntax element representing the maximum hierarchical depth of coding units generated by multi-type tree splits in the intra-frame slices of the current image, and a syntax element representing the difference between the base-2 logarithm of the smallest size generated by quadtree splits and the base-2 logarithm of the smallest coding block size in the intra-frame slices of the current image is included in the image header.

2. An image encoding method performed by an encoding device, the method comprising: Determine the type of slice in the current image, wherein the slice is included in the current image; Generate a syntax element specifying whether all slices in the current image are I-slices; and The image information is encoded, wherein the image information includes the syntax element specifying whether all slices in the current image are the I-slice. Specifically, based on the syntax elements specifying that one or more of the slices in the current image are P-slices or B-slices, the information for inter-frame slicing includes a syntax element representing the maximum hierarchical depth of the coding units generated by multi-type tree splits in the inter-frame slices of the current image, and a syntax element representing the difference between the base-2 logarithm of the minimum size generated by quadtree splits and the base-2 logarithm of the minimum coding block size in the inter-frame slices of the current image is included in the image header. Specifically, the syntax element specifies that all slices in the current image are I slices. The information for intra-frame slices includes a syntax element representing the maximum hierarchical depth of coding units generated by multi-type tree splits in the intra-frame slices of the current image, and a syntax element representing the difference between the base-2 logarithm of the smallest size generated by quadtree splits and the base-2 logarithm of the smallest coding block size in the intra-frame slices of the current image is included in the image header.

3. A method for transmitting data for an image, the method comprising: A bitstream for the image is generated, wherein the bitstream is generated based on the following: Determine the type of slice in the current image, wherein the slice is included in the current image. Generate a syntax element specifying whether all slices in the current image are I-slices, and The image information is encoded, wherein the image information includes the syntax element specifying whether all slices in the current image are the I-slice, and Send data including the bit stream. Specifically, based on the syntax elements specifying that one or more of the slices in the current image are P-slices or B-slices, the information for inter-frame slicing includes a syntax element representing the maximum hierarchical depth of the coding units generated by multi-type tree splits in the inter-frame slices of the current image, and a syntax element representing the difference between the base-2 logarithm of the minimum size generated by quadtree splits and the base-2 logarithm of the minimum coding block size in the inter-frame slices of the current image is included in the image header. Specifically, the syntax element specifies that all slices in the current image are I slices. The information for intra-frame slices includes a syntax element representing the maximum hierarchical depth of coding units generated by multi-type tree splits in the intra-frame slices of the current image, and a syntax element representing the difference between the base-2 logarithm of the smallest size generated by quadtree splits and the base-2 logarithm of the smallest coding block size in the intra-frame slices of the current image is included in the image header.