Image / video coding method and apparatus
The proposed video decoding and encoding methods optimize inter and intra prediction processes by parsing prediction information flags, enhancing compression efficiency and reducing signaling, thus addressing the need for efficient image/video coding in high-resolution and immersive media formats.
Patent Information
- Application Number
- JP2025264290
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-11-05
- Filing Date
- 2025-12-18
- Publication Date
- 2026-02-27
AI Technical Summary
The increasing demand for high-resolution and high-quality images/videos, particularly in immersive media formats like VR and AR, necessitates highly efficient image/video compression technologies to reduce transmission and storage costs while optimizing inter and intra prediction processes.
A video decoding method that parses flags from a picture header to determine the necessity of inter and intra prediction information, enabling efficient performance of these predictions and omitting unnecessary signaling, and a video encoding method that determines and encodes prediction mode information for improved coding efficiency.
Enhances overall image/video compression efficiency by allowing efficient inter and intra prediction, reducing unnecessary signaling, and improving coding performance.
Smart Images

Figure 2026034668000001_ABST
Abstract
Description
[Technical Field]
[0001] The present technology relates to a method and apparatus for coding image / video. [Background technology]
[0002] Recently, demand for high-resolution, high-quality images / videos of 4K or 8K or higher, such as UHD (Ultra High Definition) images / videos, is increasing in various fields. As the resolution and quality of image / video data increases, the amount of information or bits transmitted increases relatively compared to existing image / video data, which increases transmission and storage costs when transmitting image data using existing media such as wired or wireless broadband lines or storing image / video data using existing storage media.
[0003] In addition, interest in and demand for immersive media such as VR (Virtual Reality), AR (Artificial Reality) content, and holograms has been increasing in recent years, and the broadcast of images / videos with different image characteristics from real images, such as game images, is increasing.
[0004] Therefore, there is a need for highly efficient image / video compression technology to effectively compress and transmit, store, and play back high-resolution, high-quality image / video information having the above-mentioned various characteristics. Summary of the Invention [Problem to be solved by the invention]
[0005] The technical problem of this document is to provide a method and apparatus for improving image / video coding efficiency.
[0006] Another technical problem of this document is to provide a method and apparatus for efficiently performing inter prediction and / or intra prediction in image / video coding.
[0007] Another technical problem of this document is to provide a method and apparatus for omitting signaling unnecessary for inter prediction and / or intra prediction when transmitting image / video information. [Means for solving the problem]
[0008] According to one embodiment of the present document, a video decoding method performed by a video decoding device includes the steps of: obtaining video information from a bitstream, where the video information includes a picture header associated with a current picture, the current picture including a plurality of slices; parsing a first flag from the picture header indicating whether information required for an inter prediction operation for the decoding process is present in the picture header; parsing a second flag from the picture header indicating whether information required for an intra prediction operation for the decoding process is present in the picture header; and performing at least one of intra prediction or inter prediction on slices in the current picture based on the first flag and the second flag to generate predicted samples.
[0009] According to another embodiment of the present document, a video encoding method performed by a video encoding device includes a step of determining a prediction mode of a current block in a current picture, where the current picture includes a plurality of slices; a step of generating, based on the prediction mode, first information indicating whether information required for an inter prediction operation for the decoding process is present in a picture header associated with the current picture; a step of generating, based on the prediction mode, second information indicating whether information required for an intra prediction operation for the decoding process is present in a picture header associated with the current picture; and a step of encoding video information including the first information and the second information, where the first information and the second information are included in the picture header of the video information.
[0010] According to another embodiment of the present document, there is provided a computer-readable digital storage medium, the digital storage medium including information for causing a decoding device to perform a video decoding method, the decoding method including the steps of: acquiring video information, where the video information includes a picture header associated with a current picture, the current picture including a plurality of slices; parsing a first flag from the picture header indicating whether information necessary for an inter-prediction operation for the decoding process is present in the picture header; parsing a second flag from the picture header indicating whether information necessary for an intra-prediction operation for the decoding process is present in the picture header; and performing at least one of intra-prediction or inter-prediction on slices in the current picture based on the first flag and the second flag to generate predicted samples. [Effects of the Invention]
[0011] According to one embodiment of this document, the overall image / video compression efficiency can be improved.
[0012] According to one embodiment of the present document, when coding an image / video, inter prediction and / or intra prediction can be performed efficiently.
[0013] According to one embodiment of this document, when transmitting images / video, signaling of syntax elements unnecessary for inter-prediction or intra-prediction can be prevented. [Brief explanation of the drawings]
[0014] [Figure 1] 1 illustrates schematically an example of a video / image coding system to which embodiments of the present document may be applied. [Figure 2] 1 is a diagram illustrating the configuration of a video / image encoding device to which an embodiment of this document can be applied. [Figure 3] 1 is a diagram illustrating the configuration of a video / image decoding device to which an embodiment of the present document can be applied. [Figure 4] 1 illustrates an example of an intra-prediction based video / image encoding method. [Figure 5] 1 illustrates an example of an intra-prediction based video / image decoding method. [Figure 6] 1 illustrates an example of an inter-prediction based video / image encoding method. [Figure 7] 1 illustrates an example of an inter-prediction based video / image decoding method. [Figure 8] 1 illustrates an example of a video / image encoding method and associated components according to an embodiment of the present document. [Figure 9] 1 illustrates an example of a video / image encoding method and associated components according to an embodiment of the present document. [Figure 10] 1 illustrates an example of a video / image decoding method and related components according to an embodiment of the present document. [Figure 11] 1 illustrates an example of a video / image decoding method and related components according to an embodiment of the present document. [Figure 12]1 illustrates an example of a content streaming system to which the embodiments disclosed herein can be applied. DETAILED DESCRIPTION OF THE INVENTION
[0015] Because the disclosure of this document can be modified in various ways and can have various embodiments, specific embodiments will be illustrated in the drawings and described in detail. The terms used in this document are used merely to describe specific embodiments and are not intended to limit the technical ideas of this document. The singular expressions "a," "an," "an," "the," and the like include the expression "at least one" unless the context clearly dictates otherwise. In this document, the terms "comprise," "have," and the like are intended to specify the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, and should be understood not to preclude the presence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0016] Meanwhile, each component in the drawings described in this document is illustrated independently for the convenience of describing different characteristic functions, and does not mean that each component is realized by separate hardware or software. For example, two or more components may be combined to form a single component, or a single component may be divided into multiple components. Embodiments in which each component is integrated and / or separated are also included within the scope of this document as long as they do not deviate from the essence of the method disclosed herein.
[0017] Hereinafter, the embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Hereinafter, the same reference numerals will be used to refer to the same components in the drawings, and redundant description of the same components will be omitted.
[0018] FIG. 1 illustrates a schematic diagram of an example video / image coding system to which embodiments of this document may be applied.
[0019] As shown in Figure 1, a video / image coding system includes a first device (source device) and a second device (receiving device). The source device can transmit encoded video / image information or data to the receiving device in file or streaming form via a digital storage medium or a network.
[0020] The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / video encoding device, and the decoding device may be referred to as a video / video decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, which may be a separate device or an external component.
[0021] A video source can acquire video / images through a video / image capture, synthesis, or generation process. A video source can include a video / image capture device and / or a video / image generation device. A video / image capture device can include, for example, one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated via a computer, etc., in which case the video / image capture process can be replaced with a process in which related data is generated.
[0022] An encoding device can encode input video / images. The encoding device can perform a series of procedures such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0023] The transmitter can transmit the encoded video / image information or data output in the form of a bitstream to a receiver of a receiving device via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating a media file in a predetermined file format and elements for transmission via a broadcasting / communication network. The receiver can receive / extract the bitstream and transmit it to a decoding device.
[0024] The decoding device can decode the video / image by performing a series of steps such as inverse quantization, inverse transform, and prediction, which correspond to the operations of the encoding device.
[0025] The renderer can render the decoded video / image, and the rendered video / image can be displayed via a display unit.
[0026] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document may be applied to methods disclosed in the VVC (versatile video coding) standard. The methods / embodiments disclosed in this document may also be applied to methods disclosed in the EVC (essential video coding) standard, the AV1 (AOMedia Video 1) standard, the AVS2 (2nd generation audio video coding standard), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0027] Various embodiments relating to video / image coding are presented in this document, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0028] In this document, video can refer to a collection of a series of images over time. A picture generally refers to a unit that shows an image at a specific time, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile contains one or more coding tree units (CTUs). A picture consists of one or more slices / tiles. A picture consists of one or more tile groups. A tile group contains one or more tiles. A brick may represent a rectangular region of CTU rows within a tile in a picture. A tile may be partitioned into multiple bricks, each of which consists of one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick.A brick scan refers to a specific sequential ordering of CTUs partitioning a picture, in which the CTUs are ordered consecutively in a CTU raster scan within a brick, bricks within a tile are ordered consecutively in a raster scan of the bricks of the tile, and tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. The tile column is a rectangular region of CTUs having a height equal to the height of the picture and a width specified by syntax elements in the picture parameter set.The tile row is a rectangular region of CTUs having a height specified by syntax elements in the picture parameter set and a width equal to the height of the picture. A tile scan indicates a specific sequential ordering of CTUs partitioning a picture, in which the CTUs are ordered consecutively in a CTU raster scan in a tile, whereas tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice includes an integer number of bricks of a picture that may be exclusively contained in a single NAL unit. A slice may consist of either a number of complete tiles or only a consecutive sequence of complete bricks of one tile. In this document, the terms tile group and slice may be used interchangeably.For example, in this document, a tile group / tile group header may be referred to as a slice / slice header.
[0029] A pixel or a pel may refer to the smallest unit constituting one picture (or image). A "sample" may also be used as a term corresponding to a pixel. A sample may generally refer to a pixel or a pixel value, may refer to only a pixel / pixel value of a luma component, or may refer to only a pixel / pixel value of a chroma component. Alternatively, a sample may refer to a pixel value in the spatial domain, or may refer to a transform coefficient in the frequency domain when such a pixel value is transformed into the frequency domain.
[0030] A unit may refer to a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. The term unit may be used interchangeably with terms such as block or area. In general, an M×N block may include samples (or a sample array) consisting of M columns and N rows, or a set (or an array) of transform coefficients.
[0031] In this document, " / " and "," should be interpreted to indicate "and / or." For example, "A / B" is interpreted as "A and / or B," and "A, B" is interpreted as "A and / or B." Additionally, "A / B / C" means "at least one of A, B, and / or C." Also, "A, B, C" means "at least one of A, B, and / or C." (In this document, the term " / " and "," should be interpreted to indicate "and / or." For instance, the expression "A / B" may mean "A and / or B." Further, "A, B" may mean "A and / or B." Further, "A / B / C" may mean "at least one of A, B, and / or C." Also, "A / B / C" may mean "at least one of A, B, and / or C.")
[0032] Additionally, in this document, "or" should be interpreted as "and / or." For example, "A or B" may mean 1) only "A," 2) only "B," or 3) "A and B." Further, in the document, the term "or" should be interpreted to indicate "and / or." For instance, the expression "A or B" may comprise 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted to indicate "additionally or alternatively."
[0033] Furthermore, parentheses used in this document may mean "for example." Specifically, when "prediction (intra prediction)" is used, it means that "intra prediction" is proposed as an example of "prediction." In other words, "prediction" in this document is not limited to "intra prediction," and "intra prediction" is proposed as an example of "prediction." Furthermore, when "prediction (i.e., intra prediction)" is used, it means that "intra prediction" is proposed as an example of "prediction."
[0034] In this document, technical features individually described in one drawing may be embodied individually or simultaneously.
[0035] 2 is a diagram illustrating the configuration of a video / image encoding device to which the embodiments of this document can be applied. Hereinafter, the term "video encoding device" includes the image encoding device.
[0036] As shown in FIG. 2, the encoding apparatus 200 may include an image partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. The image dividing unit 210, the predicting unit 220, the residual processing unit 230, the entropy encoding unit 240, the adding unit 250, and the filtering unit 260 may be configured by one or more hardware components (e.g., an encoder chipset or a processor) depending on the embodiment. Also, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0037] The image division unit 210 may divide an input image (or picture, frame) input to the encoding device 200 into one or more processing units. For example, the processing units may be called coding units (CUs). In this case, the coding units may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) using a quad-tree, binary-tree, ternary-tree (QTBTTT) structure. For example, one coding unit may be divided into multiple coding units of deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and then the binary tree structure and / or the ternary structure may be applied. Alternatively, the binary tree structure may be applied first. The coding procedure according to the present disclosure may be performed based on a final coding unit that is not further divided. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to image characteristics, or the coding unit may be recursively divided into coding units of lower depths as needed, and a coding unit of an optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be divided or partitioned from the final coding unit.The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0038] The term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block can refer to a set of samples or transform coefficients consisting of M columns and N rows. A sample generally refers to a pixel or pixel value, and can refer to only a pixel / pixel value of a luma component or only a pixel / pixel value of a chroma component. A sample can also be used as a term corresponding to one pixel or pel of a picture (or image).
[0039] The encoding apparatus 200 subtracts a prediction signal (predicted block, prediction sample array) output from the inter prediction unit 221 or the intra prediction unit 222 from an input video signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is transmitted to the conversion unit 232. In this case, as shown in the figure, a unit in the encoder 200 that subtracts the prediction signal (predicted block, prediction sample array) from the input video signal (original block, original sample array) is called a subtraction unit 231. The prediction unit 220 may perform prediction on a current block (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit 220 determines whether intra prediction or inter prediction is to be applied for the current block or CU. The prediction unit 220 may generate various information related to prediction, such as prediction mode information, as will be described later in the description of each prediction mode, and transmit the information to the entropy encoding unit 240. The prediction information can be encoded in the entropy encoding unit 240 and output in the form of a bitstream.
[0040] The intra prediction unit 222 may predict the current block by referring to samples in the current picture. The referenced samples may be located adjacent to or distant from the current block depending on the prediction mode. Prediction modes in intra prediction may include a plurality of non-directional modes and a plurality of directional modes. The non-directional modes may include, for example, DC mode and planar mode. The directional modes may include, for example, 33 directional prediction modes or 65 directional prediction modes depending on the granularity of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the settings. The intra prediction unit 222 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0041] The inter prediction unit 221 may derive a predicted block for a current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include information on an inter prediction direction (such as L0 prediction, L1 prediction, or Bi prediction). In the case of inter prediction, the neighboring blocks may include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a collocated CU (col CU), or the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic). For example, the inter predictor 221 may construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index for the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of a skip mode or a merge mode, the inter predictor 221 may use motion information of neighboring blocks as motion information of the current block. In the case of the skip mode, unlike the merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using the motion vector of the neighboring block as a motion vector predictor and signaling the motion vector difference.
[0042] The prediction unit 220 generates a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit 200 can apply intra prediction or inter prediction for predicting a block, or can simultaneously apply intra prediction and inter prediction. This is called combined inter and intra prediction (CIIP). The prediction unit can also use intra block copy (IBC) prediction mode or palette mode for predicting a block. The IBC prediction mode or palette mode can be used for content image / video coding, such as games, as in screen content coding (SCC). IBC basically performs prediction within a current picture, but is similar to inter prediction in that it derives a reference block within the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. Palette mode can be considered an example of intra coding or intra prediction. When palette mode is applied, sample values within a picture can be signaled based on information about a palette table and a palette index.
[0043] The prediction signal generated via the prediction unit (including the inter prediction unit 221 and / or the intra prediction unit 222) can be used to generate a reconstructed signal or a residual signal.
[0044] The transform unit 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a graph-based transform (GBT), or a conditionally non-linear transform (CNT). Here, GBT refers to a transform obtained from a graph representing inter-pixel relationship information. CNT refers to a transform obtained based on a predicted signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to pixel blocks having the same square size or to non-square blocks of variable size.
[0045] The quantization unit 233 quantizes the transform coefficients and transmits the quantized signal to the entropy encoding unit 240. The entropy encoding unit 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs it as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantization unit 233 may rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on a coefficient scan order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0046] The entropy encoding unit 240 can perform various encoding methods, such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoding unit 240 can encode information required for video / image restoration (e.g., values of syntax elements) together with or separately from the quantized transform coefficients. The encoded information (e.g., encoded video / video information) can be transmitted or stored in the form of a bitstream in network abstraction layer (NAL) units. The video / video information can further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information can also include general constraint information. In this document, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device are included in video / image information. The video / image information is encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting and / or a storage unit (not shown) for storing the signal output from the entropy encoding unit 240 may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoding unit 240.
[0047] The quantized transform coefficients output from the quantization unit 233 may be used to generate a prediction signal. For example, a residual signal (residual block or residual sample) is reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via the inverse quantization unit 234 and the inverse transform unit 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter prediction unit 221 or the intra prediction unit 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the current block, such as when skip mode is applied, a predicted block may be used as the reconstructed block. The adder 250 may also be referred to as a reconstruction unit or a reconstructed block generator. The generated reconstructed signal is used for intra prediction of the next block to be processed in the current picture, and may also be used for inter prediction of the next picture after filtering, as described below.
[0048] Meanwhile, luma mapping with chroma scaling (LMCS) can be applied in the picture encoding and / or reconstruction process.
[0049] The filtering unit 260 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 260 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture and store the modified reconstructed picture in the memory 270, specifically, in the DPB of the memory 270. The various filtering methods include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, and a bilateral filter. The filtering unit 260 generates various information related to filtering and transmits it to the entropy encoding unit 240, as will be described later in the description of each filtering method. The entropy encoding unit 240 encodes the filtering information and outputs it in the form of a bitstream.
[0050] The modified reconstructed picture sent to the memory 270 can be used as a reference picture in the inter prediction unit 221. This allows the encoding apparatus to avoid prediction mismatch between the encoding apparatus 100 and the decoding apparatus when inter prediction is applied, and also improves coding efficiency.
[0051] The DPB of the memory 270 may store a modified reconstructed picture to be used as a reference picture in the inter predictor 221. The memory 270 may store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a block in an already reconstructed picture. The stored motion information may be transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 270 may store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 222.
[0052] FIG. 3 is a diagram illustrating the configuration of a video / image decoding device to which the embodiments of this document can be applied.
[0053] As shown in FIG. 3, the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter predictor 331 and an intra predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. Depending on the embodiment, the entropy decoding unit 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be implemented as a single hardware component (e.g., a decoder chipset or processor). The memory 360 may include a decoded picture buffer (DPB) or may be implemented as a digital storage medium. The hardware components may further include a memory 360 as an internal / external component.
[0054] When a bitstream including video / image information is input, the decoding apparatus 300 can reconstruct an image corresponding to the process in which the video / image information was processed by the encoding apparatus of FIG. 3. For example, the decoding apparatus 300 can derive units / blocks based on block division-related information obtained from the bitstream. The decoding apparatus 300 can perform decoding using a processing unit applied by the encoding apparatus. Therefore, the processing unit for decoding can be, for example, a coding unit, and the coding unit can be divided from a coding tree unit or a maximal coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced via a reproduction device.
[0055] The decoding apparatus 300 receives a signal output from the encoding apparatus of FIG. 2 in the form of a bitstream, and the received signal is decoded by the entropy decoding unit 310. For example, the entropy decoding unit 310 may parse the bitstream to derive information (e.g., video / video information) necessary for image restoration (or picture restoration). The video / video information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). The video / video information may also include general constraint information. The decoding apparatus may decode pictures based on the information on the parameter sets and / or the general constraint information. Signaled / received information and / or syntax elements, which will be described later in this document, may be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit 310 decodes information in a bitstream based on a coding method such as exponential-Golomb coding, context-adaptive variable length coding (CAVLC), or context-adaptive arithmetic coding (CABAC), and outputs values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. More specifically, the CABAC entropy decoding method receives bins corresponding to each syntax element in the bitstream, determines a context model using information on the syntax element to be decoded and decoding information on neighboring and current blocks or information on symbols / bins decoded in previous steps, predicts the occurrence probability of bins according to the determined context model, and performs arithmetic decoding of the bins to generate symbols corresponding to the values of each syntax element.In this case, the CABAC entropy decoding method may update the context model using information on the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Prediction information from the information decoded by the entropy decoding unit 310 is provided to a prediction unit (inter prediction unit 332 and intra prediction unit 331), and residual values entropy decoded by the entropy decoding unit 310, i.e., quantized transform coefficients and related parameter information, may be input to the residual processing unit 320.
[0056] The residual processing unit 320 may derive a residual signal (residual block, residual sample, residual sample array). Information related to filtering among the information decoded by the entropy decoding unit 310 is provided to the filtering unit 350. A receiving unit (not shown) for receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, and the receiving unit may be a component of the entropy decoding unit 310. The decoding device according to this document may be referred to as a video / image / picture decoding device, and the decoding device may be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder includes the entropy decoding unit 310, and the sample decoder includes at least one of the inverse quantization unit 321, the inverse transform unit 322, the adder 340, the filtering unit 350, the memory 360, the inter prediction unit 332, and the intra prediction unit 331.
[0057] The inverse quantization unit 321 may inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit 321 may rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding apparatus. The inverse quantization unit 321 may inverse quantize the quantized transform coefficients using a quantization parameter (e.g., quantization step size information) to obtain transform coefficients.
[0058] The inverse transform unit 322 performs inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0059] The prediction unit 330 performs prediction on a current block and generates a predicted block including prediction samples for the current block. The prediction unit 330 may determine whether intra prediction or inter prediction is applied to the current block based on information about the prediction output from the entropy decoding unit 310, and may determine a specific intra / inter prediction mode.
[0060] The prediction unit 330 may generate a prediction signal based on various prediction methods, which will be described later. For example, the prediction unit may apply intra prediction or inter prediction for predicting a block, or may simultaneously apply intra prediction and inter prediction. This may be referred to as combined inter and intra prediction (CIIP). The prediction unit may also use an intra block copy (IBC) prediction mode or a palette mode for predicting a block. The IBC prediction mode or palette mode can be used for content video / movie coding, such as games, as in screen content coding (SCC). IBC basically performs prediction within a current picture, but can be similar to inter prediction in that it derives a reference block within the current picture. That is, IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, information regarding a palette table and a palette index is included in the video / picture information and signaled.
[0061] The intra prediction unit 331 may predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located far away from the current block depending on the prediction mode. Prediction modes in intra prediction include a plurality of non-directional modes and a plurality of directional modes. The intra prediction unit 331 may also determine the prediction mode to be applied to the current block using the prediction modes applied to neighboring blocks.
[0062] The inter prediction unit 332 may derive a predicted block for the current block based on a reference block (reference sample array) identified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter prediction mode, the motion information may be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information includes a motion vector and a reference picture index. The motion information may further include information on the inter prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the neighboring blocks include spatial neighboring blocks present in the current picture and temporal neighboring blocks present in the reference picture. For example, the inter prediction unit 332 may construct a motion information candidate list based on the neighboring blocks and derive a motion vector and / or a reference picture index for the current block based on received candidate selection information. Inter prediction may be performed based on various prediction modes, and the prediction information may include information indicating the inter prediction mode for the current block.
[0063] The adder 340 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the acquired residual signal to a predicted signal (predicted block, predicted sample array) output from a prediction unit (including the inter prediction unit 332 and / or the intra prediction unit 331). When there is no residual for the current block, such as when a skip mode is applied, the predicted block may be used as a reconstructed block.
[0064] The adder 340 may be referred to as a reconstruction unit or a reconstruction block generator. The generated reconstruction signal may be used for intra prediction of a next block to be processed in the current picture, may be output after filtering as described below, or may be used for inter prediction of a next picture.
[0065] Meanwhile, LMCS (luma mapping with chroma scaling) can be applied during picture decoding.
[0066] The filtering unit 350 may apply filtering to the reconstructed signal to improve subjective / objective image quality. For example, the filtering unit 350 may apply various filtering methods to the reconstructed picture to generate a modified reconstructed picture, and may transmit the modified reconstructed picture to the memory 60, specifically, to the DPB of the memory 360. The various filtering methods may include, for example, deblocking filtering, sample adaptive offset, an adaptive loop filter, a bilateral filter, etc.
[0067] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter predictor 332. The memory 360 can store motion information of a block from which motion information in the current picture is derived (or decoded) and / or motion information of a block in an already reconstructed picture. The stored motion information is transmitted to the inter predictor 221 to be used as motion information of a spatially neighboring block or a temporally neighboring block. The memory 360 can store reconstructed samples of reconstructed blocks in the current picture and transmit them to the intra predictor 331.
[0068] In this document, the embodiments described for the filtering unit 260, inter prediction unit 221, and intra prediction unit 222 of the encoding device 200 can also be applied identically or correspondingly to the filtering unit 350, inter prediction unit 332, and intra prediction unit 331 of the decoding device 300, respectively.
[0069] The video / picture coding method according to this document may be performed based on the following partitioning structure. Specifically, procedures such as prediction, residual processing (e.g., inverse transform, inverse quantization), syntax element coding, and filtering, which will be described later, may be performed based on the CTUs and CUs (and / or TUs and PUs) derived based on the partitioning structure. The block partitioning procedure is performed by the image partitioning unit 210 of the encoding device described above, and partition-related information may be encoded by the entropy encoding unit 240 and transmitted to the decoding device in the form of a bitstream. The entropy decoding unit 310 of the decoding device may derive a block partitioning structure for the current picture based on the partitioning-related information obtained from the bitstream, and perform a series of procedures for video decoding (e.g., prediction, residual processing, block / picture reconstruction, in-loop filtering, etc.) based on the block partitioning structure. The CU size and the TU size may be the same, or multiple TUs may exist within a CU region. Meanwhile, the CU size may generally refer to the luma component (sample) CB (coding block) size. The TU size may generally refer to the size of a luma component (sample) TB (transform block). The chroma component (sample) CB or TB size may be derived based on the luma component (sample) CB or TB size according to a component ratio according to the color format of a picture / image (chroma format, for example, 4:4:4, 4:2:2, 4:2:0, etc.). The TU size may be derived based on maxTbSize. For example, if the CU size is larger than maxTbSize, a plurality of TUs (TBs) of the maxTbSize may be derived from the CU, and transform / inverse transform may be performed in units of the TUs (TBs). Also, for example, if intra prediction is applied, the intra prediction mode / type may be derived in units of the CU (or CB), and procedures for deriving neighboring reference samples and generating predicted samples may be performed in units of TUs (or TBs).In this case, one or more TUs (or TBs) can exist within one CU (or CB) region, and in this case, the multiple TUs (or TBs) can share the same intra prediction mode / type.
[0070] Furthermore, in video / image coding according to this document, video processing units may have a hierarchical structure. A picture may be divided into one or more tiles, bricks, slices, and / or tile groups. A slice may include one or more bricks. A brick may include one or more CTU rows within the tile. A slice may include an integer number of bricks in the picture. A tile group may include one or more tiles. A tile may include one or more CTUs. The CTUs may be divided into one or more CUs. A tile is a rectangular region of CTUs within a particular tile column and a particular tile row in a picture. A tile group may include an integer number of tiles according to tile raster scanning within a picture. A slice header may carry information / parameters that can be applied to the corresponding slice (block within the slice). If the encoding / decoding device has a multi-core processor, the encoding / decoding procedures for the tiles, slices, bricks, and / or tile groups can be processed in parallel. In this document, the terms slice and tile group can be used interchangeably. That is, a tile group header can be referred to as a slice header. Here, a slice can have one of slice types including an intra (I) slice, a predictive (P) slice, and a bi-predictive (B) slice. For blocks in an I slice, inter prediction can be used for prediction, and only intra prediction can be used. Of course, even in this case, original sample values can be coded and signaled without prediction. For blocks in a P slice, intra prediction or inter prediction can be used, and if inter prediction is used, only uni prediction can be used.Meanwhile, for blocks in a B slice, intra prediction or inter prediction can be used, and when inter prediction is used, up to bi-prediction can be used.
[0071] The encoder determines the tile / tile group, brick, slice, maximum and minimum coding unit sizes based on the characteristics of the video image (e.g., resolution) or taking into account coding efficiency or parallel processing, and information regarding this or information that can guide this can be included in the bitstream.
[0072] The decoder can obtain information indicating whether the tile / tile group, brick, slice, or CTU within the current picture is divided into multiple coding units, etc. Efficiency can be improved by obtaining (transmitting) such information only under specific conditions.
[0073] Meanwhile, as described above, one picture may include multiple slices, and one slice may include a slice header and slice data. In this case, one picture header may be added to multiple slices (slice header and slice data set) in one picture. The picture header (picture header syntax) may include information / parameters commonly applicable to the picture. The slice header (slice header syntax) may include information / parameters commonly applicable to the slices. An adaptation parameter set (APS) or picture parameter set (PPS) may include information / parameters commonly applicable to one or more pictures. A sequence parameter set (SPS) may include information / parameters commonly applicable to one or more sequences. A video parameter set (VPS) may include information / parameters commonly applicable to multiple layers. A decoding parameter set (DPS) may include information / parameters commonly applicable to the entire video. The DPS may include information / parameters related to the concatenation of a coded video sequence (CVS).
[0074] In this document, the higher level syntax may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0075] Also, for example, information regarding the division and configuration of the tiles / tile groups / bricks / slice can be configured at the encoding end through the higher level syntax and transmitted to the decoding device in the form of a bitstream.
[0076] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When the quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When the transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or may still be referred to as transform coefficients for the sake of uniformity of expression.
[0077] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about transform coefficients, and the information about the transform coefficients may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficients), and scaled transform coefficients may be derived through an inverse transform (scaling) of the transform coefficients. Residual samples may be derived based on an inverse transform (transform) of the scaled transform coefficients. This may be similarly applied / expressed in other parts of this document.
[0078] As described above, the encoding apparatus may perform various encoding methods, such as Exponential Golomb coding, Context-Adaptive Variable Length Coding (CAVLC), Context-Adaptive Binary Arithmetic Coding (CABAC), etc. Also, the decoding apparatus may decode information in a bitstream based on a coding method, such as Exponential Golomb coding, CAVLC, or CABAC, and output values of syntax elements required for image restoration and quantized values of transform coefficients related to residuals. For example, the above-described coding methods may be performed as described below.
[0079] In this document, intra prediction may refer to a prediction that generates prediction samples for a current block based on reference samples in a picture to which the current block belongs (hereinafter, the current picture). When intra prediction is applied to the current block, neighboring reference samples used for intra prediction of the current block may be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary and bottom-left of the current block having a size of nW×nH, a total of 2×nH samples, samples adjacent to the top boundary and top-right of the current block, and one sample adjacent to the top-left of the current block. Alternatively, the neighboring reference samples of the current block may include multiple columns of upper neighboring samples and multiple rows of left neighboring samples. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block, which has a size of nW x nH, a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the bottom-right of the current block.
[0080] However, some of the neighboring reference samples of the current block may not yet be decoded or may not be available. In this case, the decoding apparatus may substitute the unavailable samples as available samples to generate neighboring reference samples to be used for prediction, or may generate neighboring reference samples to be used for prediction through interpolation of available samples.
[0081] When neighboring reference samples are derived, (i) a predicted sample may be derived based on the average or interpolation of neighboring reference samples of the current block, or (ii) the predicted sample may be derived based on a reference sample located in a specific (prediction) direction relative to the predicted sample among the neighboring reference samples of the current block. (i) This may be referred to as a non-directional mode or a non-angular mode, and (ii) this may be referred to as a directional mode or an angular mode. Alternatively, the predicted sample may be generated by interpolating the first and second neighboring samples located in the opposite direction to the prediction direction of the intra-prediction mode of the current block based on the predicted sample of the current block among the neighboring reference samples. This may be referred to as linear interpolation intra-prediction (LIP). Alternatively, a chroma predicted sample may be generated based on a luma sample using a linear model. This may be referred to as LM mode. In addition, a temporary predicted sample of the current block may be derived based on filtered neighboring reference samples, and a predicted sample of the current block may be derived by weighting the temporary predicted sample and at least one reference sample derived according to the intra prediction mode from the existing neighboring reference samples, i.e., non-filtered neighboring reference samples. The above-mentioned case may be referred to as Position Dependent Intra Prediction (PDPC). In addition, intra prediction coding may be performed by selecting a reference sample line with the highest prediction accuracy from multiple neighboring reference sample lines of the current block, deriving a predicted sample using a reference sample located in the prediction direction of the corresponding line, and signaling the used reference sample line to a decoding device.The above-described case may be referred to as multi-reference line (MRL) intra prediction or MRL-based intra prediction. Furthermore, the current block may be divided into vertical or horizontal sub-partitions, and intra prediction may be performed based on the same intra prediction mode. Neighboring reference samples may be derived and used for each sub-partition. That is, in this case, the intra prediction mode for the current block is uniformly applied to the sub-partitions, and neighboring reference samples may be derived and used for each sub-partition, thereby improving intra prediction performance in some cases. This prediction method may be referred to as intra sub-partitions (ISP) or ISP-based intra prediction. The above-described intra prediction method may be distinguished from the intra prediction mode and referred to as an intra prediction type. The intra prediction type may be referred to by various terms, such as an intra prediction technique or an additional intra prediction mode. For example, the intra prediction type (or additional intra prediction mode, etc.) may include at least one of the LIP, PDPC, MRL, and ISP. A general intra prediction method excluding specific intra prediction types, such as the LIP, PDPC, MRL, and ISP, may be referred to as a normal intra prediction type. The normal intra prediction type may be generally applied when the specific intra prediction type is not applicable, and prediction may be performed based on the intra prediction mode described above. Meanwhile, post-processing filtering may be performed on the derived prediction samples as needed.
[0082] Specifically, the intra prediction procedure may include an intra prediction mode / type determination step, a neighboring reference sample derivation step, and an intra prediction mode / type-based prediction sample derivation step. In addition, a post-filtering step may be performed on the derived prediction samples as needed.
[0083] Meanwhile, in addition to the above-mentioned intra prediction types, affine linear weighted intra prediction (ALWIP) can also be used. ALWIP can also be called linear weighted intra prediction (LWIP) or matrix weighted intra prediction or matrix-based intra prediction (MIP). When the MIP is applied to a current block, a prediction sample for the current block can be derived by: i) using neighboring reference samples on which an averaging procedure has been performed, ii) performing a matrix-vector multiplication procedure, and iii) further performing horizontal / vertical interpolation, if necessary. The intra prediction mode used for the MIP can be configured to be different from the intra prediction modes used in the above-mentioned LIP, PDPC, MRL, and ISP intra prediction, and normal intra prediction. The intra prediction mode for the MIP can be called a MIP intra prediction mode, MIP prediction mode, or MIP mode. For example, the matrix and offset used in the matrix-vector multiplication can be set differently depending on the intra prediction mode for the MIP. Here, the matrix may be referred to as a (MIP) weight matrix, and the offset may be referred to as a (MIP) offset vector or a (MIP) bias vector.
[0084] A video / image encoding procedure based on intra prediction may generally include, for example, the following.
[0085] FIG. 4 illustrates an example of an intra-prediction based video / image encoding method.
[0086] Referring to FIG. 4, S400 may be performed by the intra prediction unit 222 of the encoding apparatus, and S410 to S430 may be performed by the residual processing unit 230 of the encoding apparatus. Specifically, S410 may be performed by the subtraction unit 231 of the encoding apparatus, S420 may be performed by the transform unit 232 and quantization unit 233 of the encoding apparatus, and S430 may be performed by the inverse quantization unit 234 and inverse transform unit 235 of the encoding apparatus. In S400, prediction information may be derived by the intra prediction unit 222 and encoded by the entropy encoding unit 240. Residual information may be derived through S410 and S420 and encoded by the entropy encoding unit 240. The residual information is information about the residual sample. The residual information may include information about quantized transform coefficients for the residual sample. As described above, the residual samples are derived as transform coefficients through the transform unit 232 of the encoding device, and the transform coefficients can be derived as quantized transform coefficients through the quantization unit 233. Information about the quantized transform coefficients can be encoded in the entropy encoding unit 240 through a residual coding procedure.
[0087] The encoding apparatus performs intra prediction on a current block (S400). The encoding apparatus may derive an intra prediction mode for the current block, derive neighboring reference samples for the current block, and generate prediction samples within the current block based on the intra prediction mode and the neighboring reference samples. Here, the intra prediction mode determination, neighboring reference sample derivation, and prediction sample generation procedures may be performed simultaneously, or one procedure may be performed before the other procedures. For example, the intra prediction unit 222 of the encoding apparatus may include a prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The prediction mode / type determination unit may determine the intra prediction mode / type for the current block, the reference sample derivation unit may derive neighboring reference samples for the current block, and the prediction sample derivation unit may derive motion samples for the current block. Meanwhile, if a prediction sample filtering procedure (described below) is performed, the intra prediction unit 222 may further include a prediction sample filter unit. The encoding apparatus may determine a mode to be applied to the current block from among a plurality of intra prediction modes. The encoding apparatus may compare the RD costs for the intra prediction modes to determine the optimal intra prediction mode for the current block.
[0088] Meanwhile, the encoding apparatus may also perform a prediction sample filtering procedure, which may be called post-filtering. The prediction sample filtering procedure may filter some or all of the prediction samples. In some cases, the prediction sample filtering procedure may be omitted.
[0089] The encoding apparatus derives residual samples for the current block based on the predicted samples (S410). The encoding apparatus may derive the residual samples by comparing the predicted samples with the original samples of the current block based on phase.
[0090] The encoding apparatus transforms / quantizes the residual samples to derive quantized transform coefficients (S420), and then inverse-quantizes / inverse-transforms the quantized transform coefficients to derive (modified) residual samples (S430). The reason for performing inverse-quantization / inverse-transformation again after transform / quantization is to derive residual samples that are the same as the residual samples derived by the decoding apparatus, as described above.
[0091] The encoding apparatus may generate a reconstruction block including reconstruction samples for the current block based on the prediction samples and the (corrected) residual samples (S440). A reconstruction picture for the current picture may be generated based on the reconstruction block.
[0092] As described above, the encoding apparatus may encode video information including prediction information related to the intra prediction (e.g., prediction mode information indicating a prediction mode) and residual information related to the intra / residual samples, and output the encoded video information in the form of a bitstream. The residual information may include a residual coding syntax. The encoding apparatus may transform / quantize the residual samples to derive quantized transform coefficients. The residual information may include information on the quantized transform coefficients.
[0093] A video / picture decoding procedure based on intra prediction may generally include, for example, the following.
[0094] FIG. 5 illustrates an example of an intra-prediction based video / image decoding method.
[0095] The decoding device may perform operations corresponding to those performed by the encoding device.
[0096] Referring to FIG. 5, steps S500 to S510 may be performed by an intra prediction unit 331 of a decoding device, and the prediction information of S500 and the residual information of S530 may be obtained from a bitstream by an entropy decoding unit 310 of the decoding device. The residual processing unit 320 of the decoding device may derive residual samples for a current block based on the residual information. Specifically, the inverse quantization unit 321 of the residual processing unit 320 may perform inverse quantization on quantized transform coefficients derived based on the residual information to derive transform coefficients, and the inverse transform unit 322 of the residual processing unit may perform inverse transform on the transform coefficients to derive residual samples for the current block. Step S540 may be performed by an adder 340 or a reconstruction unit of the decoding device.
[0097] Specifically, the decoding apparatus may derive an intra prediction mode for a current block based on received prediction information (S500). The decoding apparatus may derive neighboring reference samples for the current block (S510). The decoding apparatus may perform intra prediction based on the intra prediction mode and the neighboring reference samples to generate predicted samples within the current block (S520). In this case, the decoding apparatus may perform a predicted sample filtering procedure. The predicted sample filtering may be referred to as post-filtering. Some or all of the predicted samples may be filtered by the predicted sample filtering procedure. In some cases, the predicted sample filtering procedure may be omitted.
[0098] The decoding apparatus generates residual samples for the current block based on the received residual information (S530). The decoding apparatus generates reconstructed samples for the current block based on the predicted samples and the residual samples, and may derive a reconstructed block including the reconstructed samples (S540). A reconstructed picture for the current picture may be generated based on the reconstructed block.
[0099] Here, the intra prediction unit 331 of the decoding device may include a prediction mode / type determination unit, a reference sample derivation unit, and a prediction sample derivation unit. The prediction mode / type determination unit determines the intra prediction mode for the current block based on prediction mode information acquired from the entropy decoding unit 310 of the decoding device, the reference sample derivation unit derives neighboring reference samples of the current block, and the prediction sample derivation unit derives prediction samples of the current block. Meanwhile, if the above-mentioned prediction sample filtering procedure is performed, the intra prediction unit 331 may further include a prediction sample filter unit.
[0100] The prediction information may include intra prediction mode information and / or intra prediction type information. The intra prediction mode information may include, for example, flag information (e.g., intra_luma_mpm_flag) indicating whether a most probable mode (MPM) or a remaining mode is applied to the current block. If the MPM is applied to the current block, the prediction mode information may further include index information (e.g., intra_luma_mpm_idx) indicating one of the intra prediction mode candidates (MPM candidates). The intra prediction mode candidates (MPM candidates) may be configured as an MPM candidate list or an MPM list. If the MPM is not applied to the current block, the intra prediction mode information may further include remaining mode information (e.g., intra_luma_mpm_remainder) indicating one of the remaining intra prediction modes excluding the intra prediction mode candidates (MPM candidates). A decoding apparatus may determine the intra prediction mode of the current block based on the intra prediction mode information. A separate MPM list can be configured for the above-mentioned MIP.
[0101] The intra prediction type information may be implemented in various forms. For example, the intra prediction type information may include intra prediction type index information indicating one of the intra prediction types. For another example, the intra prediction type information may include at least one of reference sample line information (e.g., intra_luma_ref_idx) indicating whether the MRL is applied to the current block and, if so, which reference sample line is used, ISP flag information (e.g., intra_subpartitions_mode_flag) indicating whether the ISP is applied to the current block, ISP type information (e.g., intra_subpartitions_split_flag) indicating a subpartition split type when the ISP is applied, and flag information indicating whether PDCP is applied or flag information indicating whether LIP is applied. The intra prediction type information may also include an MIP flag indicating whether MIP is applied to the current block.
[0102] The intra prediction mode information and / or the intra prediction type information may be encoded / decoded using the coding method described herein. For example, the intra prediction mode information and / or the intra prediction type information may be encoded / decoded using entropy coding (e.g., CABAC, CAVLC) based on a truncated (rice) binary code.
[0103] Meanwhile, a video / image encoding procedure based on inter prediction may generally include, for example, the following.
[0104] FIG. 6 illustrates an example of an inter-prediction based video / image encoding method.
[0105] Referring to FIG. 6, an encoding apparatus performs inter prediction on a current block (S600). The encoding apparatus may derive an inter prediction mode and motion information of the current block and generate a predicted sample for the current block. Here, the inter prediction mode determination, motion information derivation, and predicted sample generation procedures may be performed simultaneously, or one procedure may be performed before the other procedures. For example, an inter prediction unit of the encoding apparatus may include a prediction mode determination unit, a motion information derivation unit, and a predicted sample derivation unit, in which the prediction mode determination unit may determine a prediction mode for the current block, the motion information derivation unit may derive motion information for the current block, and the predicted sample derivation unit may derive a predicted sample for the current block. For example, the inter prediction unit of the encoding apparatus may search for a block similar to the current block within a certain region (search region) of a reference picture through motion estimation and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion. Based on this, a reference picture index indicating a reference picture in which the reference block is located can be derived, and a motion vector can be derived based on a position difference between the reference block and the current block. The encoding apparatus can determine a mode to be applied to the current block from various prediction modes. The encoding apparatus can compare rate-distortion (RD) costs for the various prediction modes to determine an optimal prediction mode for the current block.
[0106] For example, when a skip mode or a merge mode is applied to the current block, the encoding apparatus may construct a merge candidate list and derive a reference block whose difference from the current block is minimum or equal to or less than a certain criterion among reference blocks indicated by merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding apparatus. Motion information of the current block may be derived using motion information of the selected merge candidate.
[0107] As another example, when the (A)MVP mode is applied to the current block, the encoding apparatus may construct an (A)MVP candidate list and use a motion vector of an MVP (motion vector predictor) candidate selected from the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector pointing to a reference block derived by the motion estimation described above may be used as the motion vector of the current block, and the MVP candidate having the smallest difference from the motion vector of the current block among the MVP candidates may become the selected MVP candidate. A motion vector difference (MVD), which is the difference obtained by subtracting the MVP from the motion vector of the current block, may be derived. In this case, information regarding the MVD may be signaled to the decoding apparatus. Furthermore, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the decoding apparatus.
[0108] The encoding apparatus may derive residual samples based on the predicted samples (S610) by comparing the original samples of the current block with the predicted samples.
[0109] The encoding apparatus encodes video information including prediction information and residual information (S620). The encoding apparatus can output the encoded video information in the form of a bitstream. The prediction information is information related to the prediction procedure and can include prediction mode information (e.g., a skip flag, a merge flag, or a mode index) and information about motion information. The information about the motion information can include candidate selection information (e.g., a merge index, an MVP flag, or an MVP index) that is information for deriving a motion vector. The information about the motion information can also include the above-mentioned information about MVD and / or reference picture index information. The information about the motion information can also include information indicating whether L0 prediction, L1 prediction, or bi-prediction is applied. The residual information is information about the residual sample. The residual information can also include information about quantized transform coefficients for the residual sample.
[0110] The output bitstream can be stored in a (digital) storage medium and then transmitted to the decoding device, or can be transmitted to the decoding device via a network.
[0111] Meanwhile, as described above, the encoding apparatus can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is because the encoding apparatus derives the same prediction result as that performed by the decoding apparatus, thereby improving coding efficiency. Therefore, the encoding apparatus can store the reconstructed picture (or reconstructed samples, reconstructed blocks) in memory and use it as a reference picture for inter prediction. As described above, an in-loop filtering procedure can be further applied to the reconstructed picture.
[0112] A video / picture decoding procedure based on inter prediction may generally include, for example, the following.
[0113] FIG. 7 illustrates an example of an inter-prediction based video / picture decoding method.
[0114] The decoding apparatus may perform operations corresponding to those performed by the encoding apparatus, and may perform prediction on the current block based on the received prediction information to derive prediction samples.
[0115] 7, a decoding apparatus may determine a prediction mode for the current block based on prediction information received from a bitstream (S700). The decoding apparatus may determine which inter-prediction mode is applied to the current block based on prediction mode information in the prediction information.
[0116] For example, it may determine whether a merge mode or an (A)MVP mode is applied to the current block based on a merge flag, or may select one of various inter prediction mode candidates based on the merge index. The inter prediction mode candidates may include various inter prediction modes such as skip mode, merge mode, and / or (A)MVP mode.
[0117] The decoding apparatus derives motion information of the current block based on the determined inter prediction mode (S710). For example, when a skip mode or a merge mode is applied to the current block, the decoding apparatus may construct a merge candidate list (described below) and select one merge candidate from among the merge candidates included in the merge candidate list. The selection may be performed based on the selection information (merge index) described above. Motion information of the selected merge candidate may be derived for the current block using motion information of the selected merge candidate. The motion information of the selected merge candidate may be used as motion information of the current block.
[0118] As another example, when the (A)MVP mode is applied to the current block, the decoding apparatus may construct an (A)MVP candidate list and use a motion vector of an MVP (motion vector predictor) candidate selected from among the MVP candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on information related to the MVD, and the motion vector of the current block may be derived based on the MVP of the current block and the MVD. Also, the decoding apparatus may derive a reference picture index of the current block based on the reference picture index information. A picture pointed to by the reference picture index in the reference picture list for the current block may be derived as a reference picture referenced for inter-prediction of the current block.
[0119] On the other hand, the motion information of the current block can be derived without constructing a candidate list, in which case the candidate list construction as described above can be omitted.
[0120] The decoding apparatus may generate predictive samples for the current block based on the motion information of the current block (S720). In this case, the reference picture may be derived based on a reference picture index of the current block, and the predictive samples of the current block may be derived using samples of a reference block to which the motion vector of the current block points on the reference picture. In this case, as described below, a predictive sample filtering procedure may be further performed on all or some of the predictive samples of the current block, depending on the circumstances.
[0121] For example, the inter-prediction unit of the decoding device may include a prediction mode determination unit, a motion information derivation unit, and a prediction sample derivation unit, and may determine a prediction mode for the current block based on prediction mode information received by the prediction mode determination unit, derive motion information (motion vector and / or reference picture index, etc.) of the current block based on information regarding the motion information received by the motion information derivation unit, and derive a prediction sample of the current block by the prediction sample derivation unit.
[0122] The decoding apparatus generates residual samples for the current block based on the received residual information (S730). The decoding apparatus generates reconstructed samples for the current block based on the predicted samples and the residual samples, and can generate a reconstructed picture based on the reconstructed samples (S740). As described above, an in-loop filtering procedure can then be further applied to the reconstructed picture.
[0123] Meanwhile, as mentioned above, HLS (high level syntax) can be coded / signaled for video / picture coding. A coded picture can consist of one or more slices. Parameters describing a coded picture are signaled in a picture header, and parameters describing a slice are signaled in a slice header. The picture header itself is carried in the form of an NAL unit. A slice header is located at the beginning of an NAL unit that contains the payload of a slice (i.e., slice data).
[0124] Each picture is associated with a picture header. A picture can be composed of different types of slices (intra-coded slices (i.e., I-slices) and inter-coded slices (i.e., P-slices and B-slices)). Therefore, a picture header can include syntax elements required for intra-slices and inter-slices of a picture. For example, the syntax of a picture header is shown in Table 1 below.
[0125] [Table 1-1]
[0126] [Table 1-2]
[0127] [Table 1-3]
[0128] [Table 1-4]
[0129] [Table 1-5]
[0130] [Table 1-6]
[0131] [Table 1-7]
[0132] [Table 1-8]
[0133] Among the syntax elements in Table 1, syntax elements containing "intra_slice" in their names (e.g., pic_log2_diff_min_qt_min_cb_intra_slice_luma) are syntax elements used in the I slice of the corresponding picture, and syntax elements containing "inter_slice" in their names (e.g., pic_log2_diff_min_qt_min_cb_inter_slice) and syntax elements related to mvp, mvd, mmvd, merge, etc. (e.g., pic_temporal_mvp_enabled_flag) are syntax elements used in the P slice and / or B slice of the corresponding picture.
[0134] That is, the picture header includes all syntax elements required for intra-coded slices and inter-coded slices for every single picture. However, this is only useful for pictures that include mixed-type slices (pictures that include both intra-coded and inter-coded slices). In the general case, a picture does not include mixed-type slices (i.e., a general picture includes only intra-coded slices or only inter-coded slices), so it is not necessary to signal all data (syntax elements used in intra-coded slices and syntax elements used in inter-coded slices).
[0135] The following drawings are created to explain a specific example of the present document. The names of specific devices and names of specific signals / information shown in the drawings are presented for illustrative purposes only, and the technical features of the present specification are not limited to the specific names used in the following drawings.
[0136] This document provides the following methods to solve the above-mentioned problems, each of which can be applied individually or in combination with each other.
[0137] 1. A flag in the picture header to specify whether syntax elements that are needed only by intra-coded slices are present in the picture header can be signaled. The flag can be called intra_signalling_present_flag.
[0138] a) When intra_signalling_present_flag is equal to 1, syntax elements needed by intra-coded slices are present in the picture header. Similarly, when intra_signalling_present_flag is equal to 0, syntax elements needed by intra-coded slices are not present in the picture header.
[0139] b) If the picture associated with the picture header has at least one intra-coded slice, the value of intra_signalling_present_flag in the picture header is equal to 1 (The value of intra_signalling_present_flag in a picture header shall be equal to 1 on the picture associated with the picture header has at least one intra-coded slice).
[0140] c) The value of intra_signalling_present_flag in a picture header may be equal to 1 even when the picture associated with the picture header does not have an intra-coded slice.
[0141] d) When a picture has one or more subpictures containing only intra-coded slices and it is anticipated that one or more of the subpictures may be extracted and merged with subpictures which contain one or more inter-coded slices, the value of intra_signalling_present_flag should be set equal to 1.
[0142] 2. A flag in the picture header to specify whether syntax elements that are needed only by inter-coded slices are present in the picture header can be signaled. The flag can be called inter_signalling_present_flag.
[0143] a) When inter_signalling_present_flag is equal to 1, syntax elements needed by inter-coded slices are present in the picture header. Similarly, when inter_signalling_present_flag is equal to 0, syntax elements needed by inter-coded slices are not present in the picture header.
[0144] b) If the picture associated with the picture header has at least one inter-coded slice, the value of inter_signalling_present_flag in the picture header shall be equal to 1 (The value of inter_signalling_present_flag in a picture header shall be equal to 1 on the picture associated with the picture header has at least one inter-coded slice).
[0145] c) Even if the picture associated with the picture header does not have an inter-coded slice, the value of inter_signalling_present_flag in the picture header is 1. (The value of inter_signalling_present_flag in a picture header may be equal to 1 even when the picture associated with the picture header does not have an inter-coded slice.)
[0146] d) When a picture has one or more subpictures containing only inter-coded slices and it is anticipated that one or more of the subpictures may be extracted and merged with subpictures containing one or more intra-coded slices, the value of inter_signalling_present_flag should be set equal to 1.
[0147] 3. The above flags (intra_signalling_present_flag and inter_signalling_present_flag) may be signaled in other parameter sets such as picture parameter sets (PPS) instead of in picture headers.
[0148] 4. An other alternative for signaling the above flags can be as follows:
[0149] a) Two variables, IntraSignallingPresentFlag and InterSignallingPresentFlag, which indicate whether syntax elements needed by intra-coded slices and syntax elements needed by inter-coded slices, respectively, are present in the picture header or not, can be defined.
[0150] b) A flag called mixed_slice_types_present_flag in the picture header can be signaled. When mixed_slice_types_present_flag is equal to 1, the values of IntraSignallingPresentFlag and InterSignallingPresentFlag are set to be equal to 1.
[0151] c) When mixed_slice_types_present_flag is equal to 0, an additional flag called intra_slice_only_flag may be signaled in the picture header and the following applies: If intra_slice_only_flag is equal to 1, the value of IntraSignallingPresentFlag is set equal to 1 and the value of InterSignallingPresentFlag is set equal to 0. Otherwise, the IntraSignallingPresentFlag value is set to 0 and the InterSignallingPresentFlag value is set to 1.
[0152] 5. A fixed length syntax element in the picture header called slice_types_idc, which may be called slice_types_idc, that specifies the following information, can be signaled:
[0153] a) Whether the picture associated with the picture header contains only intra-coded slices. For this type, the value of slice_types_idc may be set equal to 0.
[0154] b) Whether the picture associated with the picture header contains only inter-coded slices (Whether the picture associated with the picture header contain inter-coded slices only), the value of slice_types_idc may be set equal to 1 (The value of slice_types_idc may be set equal to 1).
[0155] c) Whether the picture associated with the picture header contains intra-coded slices and inter-coded slices. The value of slice_types_idc may be set equal to 2.
[0156] Note that when slice_types_idc has a value equal to 2, it is still possible that the picture contains only intra-coded slices or only inter-coded slices.
[0157] d) Other values of slice_types_idc may be reserved for future use.
[0158] 6. For slice_types_idc semantics in a picture header, the following constraints may be further specified:
[0159] a) When the picture associated with the picture header has one or more intra-coded slices, the value of slice_types_idc shall not be equal to 1.
[0160] b) When the picture associated with the picture header has one or more inter-coded slices, the value of slice_types_idc shall not be equal to 0.
[0161] 7. slice_types_idc may be signaled in other parameter sets such as picture parameter sets (PPS) instead of in picture headers.
[0162] As an example, the encoding device and the decoding device can use the following Tables 2 and 3 as the picture header syntax and semantics based on the above methods 1 and 2.
[0163] [Table 2-1]
[0164] [Table 2-2]
[0165] [Table 2-3]
[0166] [Table 3]
[0167] Referring to Tables 2 and 3, if the value of intra_signalling_present_flag is 1, this may indicate that a syntax element used only in intra-coded slices is present in the picture header. If the value of intra_signalling_present_flag is 0, this indicates that a syntax element used only in intra-coded slices is not present in the picture header. Therefore, if a picture associated with the picture header includes one or more slices having a slice type of I-slice, the value of intra_signalling_present_flag is 1. If a picture associated with the picture header does not include a slice having a slice type of I-slice, the value of intra_signalling_present_flag is 0.
[0168] If the value of inter_signalling_present_flag is 1, this may indicate that a syntax element used only in inter-coded slices is present in the picture header. If the value of inter_signalling_present_flag is 0, this indicates that a syntax element used only in inter-coded slices is not present in the picture header. Therefore, if a picture associated with the picture header includes one or more slices having a slice type of P slice and / or B slice, the value of inter_signalling_present_flag is 1. If a picture associated with the picture header does not include a slice having a slice type of P slice and / or B slice, the value of inter_signalling_present_flag is 0.
[0169] Also, for a picture containing one or more sub-pictures containing intra-coded slices that can be merged with one or more sub-pictures containing inter-coded slices, the values of intra_signalling_present_flag and inter_signalling_present_flag are all set to 1.
[0170] For example, if the current picture includes only inter-coded slices (P slices and / or B slices), the encoding device may determine the value of inter_signalling_present_flag to be 1 and the value of intra_signalling_present_flag to be 0.
[0171] As another example, if the current picture includes only intra-coded slices (I-slices), the encoding apparatus may determine the value of inter_signalling_present_flag to be 0 and the value of intra_signalling_present_flag to be 1.
[0172] As another example, if the current picture includes at least one inter-coded slice or at least one intra-coded slice, the encoding device may determine that the values of inter_signalling_present_flag and intra_signalling_present_flag are all 1.
[0173] If the value of intra_signalling_present_flag is determined to be 0, the encoding apparatus may exclude or omit syntax elements required for intra slices and generate video information including only syntax elements required for inter slices in the picture header. If the value of inter_signalling_present_flag is determined to be 0, the encoding apparatus may exclude or omit syntax elements required for inter slices and generate video information including only syntax elements required for intra slices in the picture header.
[0174] When the value of inter_signalling_present_flag obtained from a picture header in video information is 1, the decoding device determines that the corresponding picture includes at least one inter-coded slice and parses syntax elements required for intra prediction from the picture header. When the value of inter_signalling_present_flag is 0, the decoding device determines that the corresponding picture includes only intra-coded slices and parses syntax elements required for intra prediction from the picture header. When the value of intra_signalling_present_flag obtained from a picture header in video information is 1, the decoding device determines that the corresponding picture includes at least one intra-coded slice and parses syntax elements required for intra prediction from the picture header. When the value of intra_signalling_present_flag is 0, the decoding device determines that the corresponding picture includes only inter-coded slices and parses syntax elements required for inter prediction from the picture header.
[0175] As another example, the encoding device and the decoding device can use the following Tables 4 and 5 as the picture header syntax and semantics based on the above methods 5 and 6.
[0176] [Table 4-1]
[0177] [Table 4-2]
[0178] [Table 4-3]
[0179] [Table 4-4]
[0180] [Table 5]
[0181] Referring to Tables 4 and 5, if the value of slice_types_idc is 0, this indicates that the type of all slices in the picture associated with the picture header is I slice. If the value of slice_types_idc is 1, this indicates that the type of all slices in the picture associated with the picture header is P or B slice. If the value of slice_types_idc is 2, this indicates that the slice types of the slices in the picture associated with the picture header are I, P and / or B slices.
[0182] For example, if the current picture includes only intra-coded slices, the encoding apparatus may determine the value of slice_types_idc to be 0 and include only syntax elements required for decoding intra slices in the picture header, i.e., in this case, the picture header does not include syntax elements required for decoding inter slices.
[0183] As another example, if the current picture includes only inter-coded slices, the encoding apparatus may determine the value of slice_types_idc to be 1 and include only syntax elements required for decoding inter slices in the picture header, i.e., in this case, the picture header does not include syntax elements required for intra slices.
[0184] As another example, if the current picture includes at least one inter-coded slice and at least one intra-coded slice, the encoding device may determine the value of slice_types_idc to be 2 and include in the picture header all syntax elements required for decoding inter-slice and intra-slice.
[0185] If the value of slice_types_idc obtained from a picture header in video information is 0, the decoding device determines that the corresponding picture includes only intra-coded slices and parses syntax elements required for decoding the intra-coded slices from the picture header. If the value of slice_types_idc is 1, the decoding device determines that the corresponding picture includes only inter-coded slices and parses syntax elements required for decoding the inter-coded slices from the picture header. If the value of slice_types_idc is 2, the decoding device determines that the corresponding picture includes at least one intra-coded slice and at least one inter-coded slice and parses syntax elements required for decoding the intra-coded slices and inter-coded slices from the picture header.
[0186] In another embodiment, the encoding device and the decoding device may use a single flag to indicate whether a picture includes intra- and inter-coded slices. If the flag is true, i.e., the value of the flag is 1, the picture may include both intra and inter slices. In this case, the following Tables 6 and 7 may be used as the picture header syntax and semantics.
[0187] [Table 6-1]
[0188] [Table 6-2]
[0189] [Table 6-3]
[0190] [Table 6-4]
[0191] [Table 7]
[0192] Referring to Tables 6 and 7, if the value of mixed_slice_signalling_present_flag is 1, this indicates that the picture associated with the corresponding picture header can have one or more slices of different types. If the value of mixed_slice_signalling_present_flag is 0, this indicates that the picture associated with the corresponding picture header contains data related to only a single slice type.
[0193] The variables InterSignallingPresentFlag and IntraSignallingPresentFlag indicate whether the syntax elements required for intra-coded slices and inter-coded slices are present in the corresponding picture header. If the value of mixed_slice_signalling_present_flag is 1, the values of IntraSignallingPresentFlag and InterSignallingPresentFlag are set to 1.
[0194] When the value of intra_slice_only_flag is set to 1, it indicates that the value of IntraSignallingPresentFlag is set to 1 and the value of InterSignallingPresentFlag is set to 0. When the value of intra_slice_only_flag is 0, it indicates that the value of IntraSignallingPresentFlag is set to 0 and the value of InterSignallingPresentFlag is set to 1.
[0195] If the picture associated with the picture header has one or more slices whose slice type is I slice, the value of IntraSignallingPresentFlag is set to 1. If the picture associated with the picture header has one or more slices whose slice type is P or B slice, the value of InterSignallingPresentFlag is set to 1.
[0196] For example, if the current picture includes only intra-coded slices, the encoding device may determine the value of mixed_slice_signalling_present_flag to be 0, the value of intra_slice_only_flag to be 1, the value of IntraSignallingPresentFlag to be 1, and the value of InterSignallingPresentFlag to be 0.
[0197] As another example, if the current picture includes only inter-coded slices, the encoding device may determine the value of mixed_slice_signalling_present_flag to be 0, the value of intra_slice_only_flag to be 0, the value of IntraSignallingPresentFlag to be 0, and the value of InterSignallingPresentFlag to be 1.
[0198] As another example, the encoding apparatus may determine the values of mixed_slice_signalling_present_flag, IntraSignallingPresentFlag, and InterSignallingPresentFlag to be 1 if the current picture includes at least one intra-coded slice and at least one inter-coded slice.
[0199] A decoding apparatus may determine that a corresponding picture includes only intra-coded slices or inter-coded slices if the value of mixed_slice_signalling_present_flag acquired from a picture header in video information is 0. In this case, if the value of intra_slice_only_flag acquired from the picture header is 0, the decoding apparatus may parse only syntax elements required for decoding inter-coded slices from the picture header. If the value of intra_slice_only_flag is 1, the decoding apparatus may parse only syntax elements required for decoding intra-coded slices from the picture header.
[0200] If the value of mixed_slice_signalling_present_flag obtained from the picture header in the video information is 1, the decoding device determines that the corresponding picture includes at least one intra-coded slice and at least one inter-coded slice, and can parse the syntax elements required for decoding the inter-coded slice and the syntax elements required for decoding the intra-coded slice from the picture header.
[0201] 8 and 9 illustrate an example of a video / image encoding method and associated components according to an embodiment of the present document.
[0202] The video / image encoding method disclosed in Figure 8 may be performed by the (video / image) encoding apparatus 200 disclosed in Figures 2 and 9. Specifically, for example, S800 in Figure 8 may be performed by the prediction unit 220 of the encoding apparatus 200, and S810 to S830 may be performed by the entropy encoding unit 240 of the encoding apparatus 200. The video / image encoding method disclosed in Figure 8 may include the embodiments detailed in this document.
[0203] Specifically, referring to FIGS. 8 and 9, the prediction unit 220 of the encoding apparatus may determine a prediction mode of a current block in a current picture (S800). The current picture may include a plurality of slices. The prediction unit 220 of the encoding apparatus may generate prediction samples (predicted blocks) for the current block based on the prediction mode. Here, the prediction mode may include an inter prediction mode and an intra prediction mode. If the prediction mode of the current block is the inter prediction mode, the prediction samples may be generated by an inter prediction unit 221 of the prediction unit 220. If the prediction mode of the current block is the intra prediction mode, the prediction samples may be generated by an intra prediction unit 222 of the prediction unit 220.
[0204] The residual processor 230 of the encoding device may generate residual samples and residual information based on the predicted samples and the original picture (original block, original sample). Here, the residual information is information about the residual samples and may include information about (quantized) transform coefficients for the residual samples.
[0205] The adder (or reconstruction unit) of the encoding device can generate reconstructed samples (reconstructed pictures, reconstruction blocks, reconstructed sample arrays) by adding the residual samples generated by the residual processing unit 230 and the predicted samples generated by the inter prediction unit 221 or the intra prediction unit 222.
[0206] Meanwhile, the entropy encoding unit 240 of the encoding apparatus may generate first information indicating whether information required for an inter prediction operation for a decoding process is present in a picture header associated with the current picture based on the prediction mode (S810). The entropy encoding unit 240 of the encoding apparatus may also generate second information indicating whether information required for an intra prediction operation for the decoding process is present in a picture header associated with the current picture (S820). Here, the first information and the second information are information included in a picture header of the video information and may correspond to the above-mentioned intra_signalling_present_flag, inter_signalling_present_flag, slice_type_idc, mixed_slice_signalling_present_flag, intra_slice_only_flag, IntraSignallingPresentFlag, and / or InterSignallingPresentFlag.
[0207] For example, if the current picture includes an inter-coded slice and thus a picture header associated with the current picture includes information necessary for an inter-prediction operation for a decoding process, the entropy encoding unit 240 of the encoding apparatus may determine the value of the first information to be 1. Also, if the current picture includes an intra-coded slice and thus a picture header associated with the current picture includes information necessary for an intra-prediction operation for a decoding process, the entropy encoding unit 240 of the encoding apparatus may determine the value of the second information to be 1. In this case, the first information may correspond to inter_signalling_present_flag, and the second information may correspond to intra_signalling_present_flag. The first information may be referred to as a first flag, information on whether a syntax element used for an inter slice is present in the picture header, a flag on whether a syntax element used for an inter slice is present in the picture header, information on whether a slice in the current picture is an inter slice, a flag on whether the slice is an inter slice, etc. The second information may be referred to as a second flag, information regarding whether syntax elements used for intra slices are present in the picture header, a flag regarding whether syntax elements used for intra slices are present in the picture header, information regarding whether a slice in the current picture is an intra slice, a flag regarding whether the slice is an intra slice, etc.
[0208] Meanwhile, when the picture includes only intra-coded slices and thus the corresponding picture header includes only information necessary for intra prediction, the entropy encoding unit 240 of the encoding apparatus may determine the value of the first information to be 0 and the value of the second information to be 1. When the picture includes only inter-coded slices and thus the corresponding picture header includes only information necessary for inter prediction, the entropy encoding unit 240 of the encoding apparatus may determine the value of the first information to be 1 and the value of the second information to be 0. Therefore, when the value of the first information is 0, all slices in the current picture may have an I slice type. When the value of the second information is 0, all slices in the current picture may have a P slice type or a B slice type. Here, the information necessary for intra prediction may include syntax elements used for decoding intra slices, and the information necessary for inter prediction may include syntax elements used for decoding inter slices.
[0209] As another example, the entropy encoding unit 240 of the encoding device may determine the value of the information about slice type to be 0 if all slices in the current picture have the I slice type, determine the value of the information about slice type to be 1 if all slices in the current picture have the P slice type or the B slice type, and determine the value of the information about slice type to be 2 if all slices in the current picture have the I slice type, the P slice type, and / or the B slice type (i.e., if the slice types of the slices in the picture are mixed). In this case, the information about slice type may correspond to slice_type_idc.
[0210] As another example, if all slices in a current picture have the same slice type, the entropy encoding unit 240 of the encoding device may determine the value of the information about the slice type to be 0, and if slices in the current picture have different slice types, the entropy encoding unit 240 may determine the value of the information about the slice type to be 1. In this case, the information about the slice type may correspond to mixed_slice_signalling_present_flag.
[0211] If the value of the information about the slice type is determined to be 0, the corresponding picture header may include information about whether the slice includes an intra slice. The information about whether the slice includes an intra slice may correspond to intra_slice_only_flag. If all slices in the picture have an I slice type, the entropy encoding unit 240 of the encoding device may determine the value of the information about whether the slice includes an intra slice to be 1, the value of the information about whether syntax elements used for intra slices are present in the picture header to be 1, and the value of the information about whether syntax elements used for inter slices are present in the picture header to be 0. If the slice types of all slices in the picture are P slice and / or B slice types, the entropy encoding unit 240 of the encoding device may determine the value of the information about whether the slice includes an intra slice to be 0, the value of the information about whether syntax elements used for intra slices are present in the picture header to be 0, and the value of the information about whether syntax elements used for inter slices are present in the picture header to be 1.
[0212] The entropy encoding unit 240 of the encoding apparatus may encode video information including the above-described first information, second information, information on slice type, etc., along with residual information, prediction-related information, etc. (S830). For example, the video information may include partitioning-related information, information on a prediction mode, residual information, in-loop filtering-related information, first information, second information, information on a slice type, etc., and may include various syntax elements related thereto. For example, the video information may include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video information may include various information, such as a picture header syntax, a picture header structure syntax, a slice header syntax, a coding unit syntax, etc. The above-described first information, second information, information on a slice type, information required for the intra prediction operation, and information required for the inter prediction operation may be included in syntax within the picture header.
[0213] The information encoded by the entropy encoding unit 240 of the encoding device can be output in the form of a bitstream, which can be transmitted to a decoding device via a network or a storage medium.
[0214] 10 and 11 show a schematic diagram of an example of a video / image decoding method and associated components according to an embodiment of the present document.
[0215] The video / image decoding method disclosed in Figure 10 may be performed by the (video / image) decoding device 300 disclosed in Figures 3 and 11. Specifically, for example, steps S1000 to S1020 in Figure 10 may be performed by the entropy decoding unit 310 of the decoding device, and step S1030 may be performed by the prediction unit 330 of the decoding device 300. The video / image decoding method disclosed in Figure 10 may include the embodiments detailed in this document.
[0216] 10 and 11, an entropy decoding unit 310 of a decoding device may obtain video information from a bitstream (S1000). The video information may include a picture header associated with a current picture. The current picture may include a plurality of slices.
[0217] Meanwhile, the entropy decoding unit 310 of the decoding device may parse a first flag from the picture header associated with the current picture, the first flag indicating whether information required for an inter prediction operation for the decoding process exists in the picture header associated with the current picture (S1010). The entropy decoding unit 310 of the decoding device may also parse a second flag from the picture header, the second flag indicating whether information required for an intra prediction operation for the decoding process exists in the picture header associated with the current picture (S1020). Here, the first flag and the second flag may correspond to the above-mentioned intra_signalling_present_flag, inter_signalling_present_flag, slice_type_idc, mixed_slice_signalling_present_flag, intra_slice_only_flag, IntraSignallingPresentFlag, and / or InterSignallingPresentFlag. The entropy decoding unit 310 of the decoding device can parse syntax elements included in the picture header of the video information based on any one of the picture header syntaxes of Table 2, Table 4, and Table 6.
[0218] The decoding device may perform at least one of intra prediction or inter prediction on a slice in the current picture based on the first flag, the second flag, information about the slice type, etc., to generate a predicted sample (S1030).
[0219] Specifically, the entropy decoding unit 310 of the decoding device may parse (or acquire) at least one of information necessary for an intra prediction operation or information necessary for an inter prediction operation for the decoding process from a picture header associated with the current picture based on the first flag, the second flag, and / or information about the slice type, etc. The prediction unit 330 of the decoding device may perform intra prediction and / or inter prediction to generate prediction samples based on at least one of the information necessary for the intra prediction operation or the information about the inter prediction. Here, the information necessary for the intra prediction operation may include syntax elements used for decoding intra slices, and the information necessary for the inter prediction operation may include syntax elements used for decoding inter slices.
[0220] For example, if the value of the first flag is 0, the entropy decoding unit 310 of the decoding device may determine that a syntax element used for the inter prediction is not present in the picture header and parse only information required for the intra prediction operation from the picture header. If the value of the first flag is 1, the entropy decoding unit 310 of the decoding device may determine that a syntax element used for the inter prediction is present in the picture header and parse information required for the inter prediction operation from the picture header. In this case, the first flag may correspond to inter_signalling_present_flag.
[0221] Furthermore, if the value of the second flag is 0, the entropy decoding unit 310 of the decoding device may determine that the picture header does not contain a syntax element used for the intra prediction, and parse only information required for the inter prediction operation from the picture header. If the value of the second flag is 1, the entropy decoding unit 310 of the decoding device may determine that the picture header contains a syntax element used for the intra prediction, and parse only information required for the intra prediction operation from the picture header. In this case, the second flag may correspond to intra_signalling_present_flag.
[0222] If the value of the first flag is 0, the decoding device may determine that all slices in the current picture have an I-slice type. If the value of the first flag is 1, the decoding device may determine that zero or more slices in the current picture have a P-slice or B-slice type. That is, if the value of the first flag is 1, the current picture may or may not include a slice having a P-slice or B-slice type.
[0223] Furthermore, if the value of the second flag is 0, the decoding apparatus may determine that all slices in the current picture have a P slice or a B slice type. If the value of the second flag is 1, the decoding apparatus may determine that zero or more slices in the current picture have an I slice type. That is, if the value of the second flag is 1, the current picture may or may not include a slice of an I slice type.
[0224] As another example, if the value of the information about slice type is 0, the entropy decoding unit 310 of the decoding device may determine that all slices in the current picture have an I slice type and parse only the information required for the intra prediction operation from the picture header. If the information about slice type is 1, the entropy decoding unit 310 of the decoding device may determine that all slices in the current picture have a P slice type or a B slice type and parse only the information required for the inter prediction operation from the picture header. If the value of the information about slice type is 2, the entropy decoding unit 310 of the decoding device may determine that the slices in the current picture have a slice type that is a mixture of I slice type, P slice type, and / or B slice type and parse all the information required for the inter prediction operation and the information required for the intra prediction operation from the picture header. In this case, the information about slice type may correspond to slice_type_idc.
[0225] As another example, the entropy decoding unit 310 of the decoding device may determine that all slices in a picture have the same slice type if the value of the information about the slice type is 0, and may determine that the slices in the picture have different slice types if the value of the information about the slice type is 1. In this case, the information about the slice type may correspond to mixed_slice_signalling_present_flag.
[0226] The entropy decoding unit 310 of the decoding device may parse information on whether the slice includes an intra slice from the picture header if the value of the information on slice type is 0. The information on whether the slice includes an intra slice may correspond to the above-mentioned intra_slice_only_flag. If the information on whether the slice includes an intra slice is 1, all slices in the picture may have an I-slice type.
[0227] The entropy decoding unit 310 of the decoding device may parse only information necessary for the intra prediction operation from the picture header when the value of the information on whether the slice includes an intra slice is 1. When the value of the information on whether the slice includes an intra slice is 0, the entropy decoding unit 310 of the decoding device may parse only information necessary for the inter prediction operation from the picture header.
[0228] If the value of the information about the slice type is 1, the entropy decoding unit 310 of the decoding device can parse all of the information necessary for the inter prediction operation and the information necessary for the intra prediction operation from the picture header.
[0229] Meanwhile, the residual processor 320 of the decoding device can generate residual samples based on the residual information acquired by the entropy decoding unit 310 .
[0230] The adder 340 of the decoding device may generate reconstructed samples based on the predicted samples generated by the predictor 330 and the residual samples generated by the residual processor 320. The adder 340 of the decoding device may then generate a reconstructed picture (reconstructed block) based on the reconstructed samples.
[0231] Thereafter, if necessary, in-loop filtering procedures such as deblocking filtering, SAO and / or ALF procedures can be applied to the reconstructed picture to improve the subjective / objective image quality.
[0232] In the above-described embodiments, the methods are described based on flow charts as a series of steps or blocks, but the embodiments are not limited to the order of the steps, and certain steps may occur in a different order or simultaneously with other steps than those described. Furthermore, those skilled in the art will understand that the steps shown in the flow charts are not exclusive, and other steps may be included, or one or more steps in the flow charts may be deleted without affecting the scope of the embodiments herein.
[0233] The methods according to the embodiments of the present document described above can be implemented in software form, and the encoding device and / or decoding device according to the present document can be included in devices that perform video processing, such as TVs, computers, smartphones, set-top boxes, and display devices.
[0234] In this document, when an embodiment is implemented in software, the method described above may be implemented with modules (processes, functions, etc.) that perform the functions described above. The modules may be stored in memory and executed by a processor. The memory may be internal or external to the processor and may be coupled to the processor in various well-known ways. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described herein may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each drawing may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information on instructions) or algorithms for implementation may be stored on a digital storage medium.
[0235] In addition, the decoding device and encoding device to which the embodiments of this document are applied may be included in a multimedia broadcast transmitting / receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video interaction device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a camcorder, a custom video (VoD) service providing device, an over-the-top (OTT) video (over-the-top) device, an internet streaming service providing device, a three-dimensional (3D) video device, a virtual reality (VR) device, an augmented reality (AR) device, an image telephone video device, a vehicle terminal (e.g., a vehicle terminal (including an autonomous vehicle), an airplane terminal, a ship terminal, etc.), a medical video device, etc., and may be used to process video signals or data signals. For example, over-the-top (OTT) video (over-the-top) devices may include a game console, a Blu-ray player, an internet-connected TV, a home theater system, a smartphone, a tablet PC, a digital video recorder (DVR), etc.
[0236] In addition, a processing method to which the embodiments of this document are applied may be produced in the form of a computer-executable program and stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiments of this document may also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. Examples of the computer-readable recording medium include Blu-ray Discs (BDs), Universal Serial Buses (USBs), ROMs, PROMs, EPROMs, EEPROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. The computer-readable recording medium also includes media embodied in the form of carrier waves (e.g., transmission via the Internet). A bitstream generated by the encoding method may be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0237] Furthermore, the embodiments of the present document may be embodied in a computer program product having program code, which may be executed by a computer in accordance with the embodiments of the present document. The program code may be stored on a computer-readable carrier.
[0238] FIG. 12 illustrates an example of a content streaming system in which the embodiments disclosed herein can be applied.
[0239] Referring to FIG. 12, a content streaming system to which the embodiments of this document are applied may largely include an encoding server, a streaming server, a web server, a media repository, a user device, and a multimedia input device.
[0240] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, camcorder, etc. into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, if a multimedia input device such as a smartphone, camera, camcorder, etc. directly generates a bitstream, the encoding server may be omitted.
[0241] The bitstream may be generated by an encoding method or a bitstream generation method applied to an embodiment of this document, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0242] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server, which controls commands and responses between devices in the content streaming system.
[0243] The streaming server can receive content from a media repository and / or an encoding server. For example, if content is received from the encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a certain period of time to provide a smooth streaming service.
[0244] Examples of the user devices include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, and digital signs.
[0245] Each server in the content streaming system can be operated as a distributed server, in which case data received by each server can be processed in a distributed manner.
Claims
1. A video decoding method performed by a decoding device, comprising: receiving a bitstream including video information, the video information including a picture header associated with a current picture, the current picture including a slice; obtaining a first flag from the picture header relating to whether information for inter-slice is present in the picture header; obtaining the information for the inter slice from the picture header based on the first flag; obtaining a second flag from the picture header related to whether information for an intra slice is present in the picture header; obtaining the information for the intra slice from the picture header based on the second flag; performing at least one of intra prediction and inter prediction on blocks in the slice in the current picture based on the first flag and the second flag to generate predicted samples; the information for the inter slice is included in the picture header based on the value of the first flag being equal to 1; the information for the inter slice includes a syntax element indicating a difference between the base 2 logarithm of a minimum size resulting from a quadtree partitioning and the base 2 logarithm of a minimum coding block size in the inter slice in the current picture; the information for the intra slice is included in the picture header based on the value of the second flag being equal to 1; The video decoding method, wherein the information for the intra slice includes a syntax element indicating the difference between the base 2 logarithm of a minimum size resulting from quadtree partitioning and the base 2 logarithm of a minimum coding block size in the intra slice in the current picture.
2. A video encoding method performed by an encoding device, comprising: determining the type of slice in the current picture; generating a first flag related to whether information for inter-slice is present in a picture header associated with the current picture; generating the information for the inter slice based on the first flag; generating a second flag related to whether information for an intra slice is present in the picture header associated with the current picture; generating the information for the intra slice based on the second flag; encoding video information including the first flag, the second flag, the information for the inter slice, and the information for the intra slice; the first flag and the second flag are included in the picture header of the video information; the information for the inter slice is included in the picture header based on the value of the first flag being equal to 1; the information for the inter slice includes a syntax element indicating a difference between the base 2 logarithm of a minimum size resulting from a quadtree partitioning and the base 2 logarithm of a minimum coding block size in the inter slice in the current picture; the information for the intra slice is included in the picture header based on the value of the second flag being equal to 1; The video encoding method, wherein the information for the intra slice includes a syntax element indicating the difference between the base 2 logarithm of a minimum size resulting from quadtree partitioning and the base 2 logarithm of a minimum coding block size in the intra slice in the current picture.
3. 1. A method of transmitting data for video, comprising: obtaining a bitstream for the video, the bitstream comprising: determining the type of slice in the current picture; generating a first flag related to whether information for inter-slice is present in a picture header associated with the current picture; generating the information for the inter slice based on the first flag; generating a second flag related to whether information for an intra slice is present in the picture header associated with the current picture; generating the information for the intra slice based on the second flag; encoding video information including the first flag, the second flag, the information for the inter slice, and the information for the intra slice; transmitting the data including the bitstream; the first flag and the second flag are included in the picture header of the video information; the information for the inter slice is included in the picture header based on the value of the first flag being equal to 1; the information for the inter slice includes a syntax element indicating a difference between the base 2 logarithm of a minimum size resulting from a quadtree partitioning and the base 2 logarithm of a minimum coding block size in the inter slice in the current picture; the information for the intra slice is included in the picture header based on the value of the second flag being equal to 1; The transmission method, wherein the information for the intra slice includes a syntax element indicating the difference between the base 2 logarithm of the smallest size resulting from a quadtree partitioning and the base 2 logarithm of the smallest coding block size in the intra slice in the current picture.