Image / video coding method and device based on slice type
By parsing the picture header flag in the video decoding device and optimizing the signaling transmission of inter-frame prediction and intra-frame prediction, the problem of low efficiency in high-resolution image/video transmission is solved, and more efficient encoding and resource utilization are achieved.
Patent Information
- Application Number
- CN202080084681.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-05
- Filing Date
- 2020-11-05
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2040-11-05
AI Technical Summary
Existing technologies are inefficient in the transmission and storage of high-resolution, high-quality images/videos, and the inter-frame prediction and intra-frame prediction signaling are redundant, increasing costs and resource consumption.
By parsing the flag in the picture header in the video decoding device, the necessity of inter-frame prediction or intra-frame prediction is determined, and only the necessary signaling information is sent to optimize the image/video encoding process.
It improves image/video compression efficiency, reduces unnecessary signaling transmission, and improves coding efficiency and resource utilization.
Smart Images

Figure CN114762350B_ABST
Abstract
Description
Technical Field
[0001] The present technology relates to a method and apparatus for coding an image / video based on information about a slice type. Background Art
[0002] Recently, the demand for high-resolution, high-quality images / videos, such as 4K or 8K or higher ultra-high-definition (UHD) images / videos, has been increasing in various fields. As the image / video resolution or quality becomes higher, a relatively larger amount of information or bits is transmitted compared to conventional image / video data. Therefore, if the image / video data is transmitted via a medium such as an existing wired / wireless broadband line or stored in a traditional storage medium, the cost for transmission and storage is likely to increase.
[0003] In addition, there is growing interest and demand for virtual reality (VR) and artificial reality (AR) content and immersive media such as holograms; and there is also growing broadcasting of images / videos that exhibit image / video characteristics that are different from actual images / videos (e.g., game images / videos).
[0004] Therefore, highly efficient image / video compression technology is required to effectively compress and transmit, store, or play high-resolution, high-quality images / videos exhibiting various characteristics as described above. Summary of the Invention
[0005] Technical issues
[0006] This document provides a method and apparatus for improving image / video coding efficiency.
[0007] This document also provides a method and apparatus for efficiently performing inter-frame prediction and / or intra-frame prediction in image / video coding.
[0008] This document also provides a method and apparatus for efficiently signaling slice type related information when transmitting image / video information.
[0009] This document also provides a method and apparatus for omitting signaling unnecessary for inter-frame prediction and / or intra-frame prediction when transmitting image / video information.
[0010] Technical Solution
[0011] According to an embodiment of the present document, a video decoding method performed by a video decoding device is provided, the method including: obtaining image information from a bitstream, wherein the image information includes a picture header associated with a current picture, and the current picture includes multiple slices; parsing at least one of a first flag or a second flag from the picture header; generating predicted samples by performing at least one of intra-frame prediction or inter-frame prediction on a current block in the current picture based on at least one of the first flag or the second flag; generating reconstructed samples based on the predicted samples; and generating a reconstructed picture based on the reconstructed samples, wherein the first flag indicates whether information necessary for the inter-frame prediction operation is present in the picture header, and wherein the second flag indicates whether information necessary for the inter-frame prediction operation is present in the picture header.
[0012] According to another embodiment of the present document, a video encoding method performed by a video encoding device is provided, the method including: determining a prediction mode of a current block in a current picture, wherein the current picture includes multiple slices; generating prediction samples based on the prediction mode; generating reconstructed samples of the current block based on the prediction samples; generating at least one of first information or second information based on the prediction mode; and encoding image information including at least one of the first information or the second information, wherein the first information and the second information are included in a picture header associated with the current picture, wherein the first information indicates whether information required for an inter-frame prediction operation is present in the picture header, and wherein the second information indicates whether information required for the inter-frame prediction operation is present in the picture header.
[0013] According to another embodiment of the present document, a computer-readable digital storage medium is provided, the computer-readable digital storage medium containing information that causes a decoding device to perform a video decoding method, the decoding method comprising: obtaining image information, wherein the image information includes a picture header associated with a current picture, and the current picture includes multiple slices; parsing at least one of a first flag or a second flag from the picture header; generating predicted samples by performing at least one of intra-frame prediction or inter-frame prediction on a current block in the current picture based on at least one of the first flag or the second flag; generating reconstructed samples based on the predicted samples; and generating a reconstructed picture based on the reconstructed samples, wherein the first flag indicates whether information necessary for the inter-frame prediction operation is present in the picture header, and wherein the second flag indicates whether information necessary for the inter-frame prediction operation is present in the picture header.
[0014] Beneficial effects
[0015] According to the embodiments of this document, the overall image / video compression efficiency can be improved.
[0016] According to the embodiments of this document, inter-frame prediction and / or intra-frame prediction can be efficiently performed when coding an image / video.
[0017] According to an embodiment of this document, when transmitting image / video information, information related to a slice type can be efficiently signaled.
[0018] According to an embodiment of this document, when image / video information is transmitted, signaling of syntax elements unnecessary for inter prediction or intra prediction can be prevented. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 An example of a video / image coding system to which embodiments of the present disclosure are applicable is schematically shown.
[0020] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure is applicable.
[0021] Figure 3 is a diagram schematically illustrating a configuration of a video / image decoding device to which an embodiment of the present disclosure is applicable.
[0022] Figure 4 An exemplary representation of the hierarchical structure for coding images / videos.
[0023] Figure 5 This shows an example of the picture decoding process.
[0024] Figure 6 An example showing the picture encoding process.
[0025] Figure 7 An example of a video / image encoding method based on inter-frame prediction is shown.
[0026] Figure 8 An example of a video / image decoding method based on inter-frame prediction is shown.
[0027] Figure 9 and Figure 10 Schematically represents an example of a video / image encoding method and related components according to an embodiment of this document.
[0028] Figure 11 and Figure 12 Schematically represents an example of a video / image decoding method and related components according to an embodiment of this document.
[0029] Figure 13 represents an example of a content streaming system to which the embodiments disclosed in this document can be applied. DETAILED DESCRIPTION
[0030] This document relates to video / image coding. For example, the methods / embodiments disclosed in this document can be applied to methods disclosed in the Versatile Video Coding (VVC) standard. In addition, the methods / embodiments disclosed in this document can be applied to methods disclosed in the Essential Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the 2nd generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0031] Various embodiments related to video / image coding are presented in this document, and unless otherwise specified, the above embodiments may also be performed in combination with each other.
[0032] The disclosure of the present disclosure can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. The terms used in the present disclosure are only used to describe specific embodiments and are not intended to limit the disclosed methods in the present disclosure. Singular expressions include the expression of "at least one" as long as it is clearly interpreted differently. Terms such as "including" and "having" are intended to indicate the presence of features, quantities, steps, operations, elements, components, or combinations thereof used in the document, and therefore it should be understood that the possibility of the presence or addition of one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0033] In addition, each configuration of the drawings described in this document is an independent diagram for explaining the functions of different features, and does not mean that each configuration is implemented by different hardware or different software. For example, two or more configurations can be combined to form a configuration, and a configuration can also be divided into multiple configurations. Without departing from the gist of the disclosed method of the present disclosure, embodiments of combined and / or separated configurations are included within the scope of the disclosure of the present disclosure.
[0034] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" may mean "A and / or B". Furthermore, "A, B" may mean "A and / or B". Furthermore, "A / B / C" may mean "at least one of A, B, and / or C". Furthermore, "A / B / C" may mean "at least one of A, B, and / or C".
[0035] Furthermore, in this document, the term "or" should be interpreted as meaning "and / or." For example, the expression "A or B" may include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as meaning "additionally or alternatively."
[0036] Furthermore, brackets used in this specification may mean "for example." Specifically, when "prediction (intra-frame prediction)" is expressed, it may indicate that "intra-frame prediction" is presented as an example of "prediction." In other words, the term "prediction" in this specification is not limited to "intra-frame prediction" and may indicate that "intra-frame prediction" is presented as an example of "prediction." Furthermore, even when "prediction (i.e., intra-frame prediction)" is expressed, it may indicate that "intra-frame prediction" is presented as an example of "prediction."
[0037] In this specification, technical features described separately in one drawing may be implemented separately or may be implemented simultaneously.
[0038] Hereinafter, the embodiments of the present document will be described in detail with reference to the accompanying drawings. In addition, in all drawings, the same reference numerals may be used to indicate the same elements, and the same description of the same elements will be omitted.
[0039] Figure 1 An example of a video / image coding system to which embodiments of the present disclosure can be applied is illustrated.
[0040] Reference Figure 1 The video / image coding system may include a first device (source device) and a second device (receiving device). The source device may send the encoded video / image information or data in the form of a file or stream to the receiving device via a digital storage medium or a network.
[0041] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0042] A video source may acquire a video / image through a process of capturing, synthesizing, or generating a video / image. A video source may include a video / image capture device and / or a video / image generation device. For example, a video / image capture device may include one or more cameras, a video / image archive including previously captured videos / images, etc. For example, a video / image generation device may include a computer, a tablet computer, and a smartphone, and may (electronically) generate a video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process of generating relevant data.
[0043] An encoding device can encode input video / images. For compression and coding efficiency, the encoding device may perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output as a bitstream.
[0044] The transmitter can transmit the encoded image / image information or data, output as a bitstream, in the form of a file or stream to a receiver of a receiving device via a digital storage medium or network. Digital storage media may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include components for generating a media file in a predetermined file format and may also include components for transmission via a broadcast / communication network. The receiver may receive / extract the bitstream and transmit the received bitstream to a decoding device.
[0045] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.
[0046] The renderer may render the decoded video / image, and the rendered video / image may be displayed on a display.
[0047] In this document, video may refer to a series of images over time. A picture generally refers to a unit representing an image at a specific time frame, and a slice / tile refers to a unit that constitutes a part of a picture in terms of coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A tile may represent a rectangular area of a CTU row within a tile in a picture. A tile may be divided into multiple tiles, each of which may be composed of one or more CTU rows within a tile. A tile that is not divided into multiple tiles may also be referred to as a tile. Tile scanning may represent a specific sequential ordering of the CTUs of a partitioned picture, where CTUs are continuously ordered within a tile in a CTU raster scan, tiles within a tile are continuously ordered in a raster scan of the tiles of the tile, and tiles in a picture are continuously ordered in a raster scan of the tiles of the picture. A tile is a rectangular area of CTUs within a specific tile column and a specific tile row in a picture. A tile column is a rectangular region of a CTU with a height equal to the height of the picture and a width specified by a syntax element in the picture parameter set. A tile row is a rectangular region of a CTU with a height specified by a syntax element in the picture parameter set and a width equal to the width of the picture. A tile scan is a specific sequential ordering of the CTUs that partition a picture, where the CTUs are ordered consecutively in a CTU raster scan within a tile and the tiles in a picture are ordered consecutively in a raster scan of the tiles of the picture. A slice comprises an integer number of tiles of a picture that can be contained in only a single NAL unit. A slice can consist of multiple complete tiles, or just a sequence of consecutive complete tiles of a tile. In this document, tile group and slice can be used instead of each other. For example, in this document, a tile group / tile group header can be referred to as a slice / slice header.
[0048] A pixel or picture element (pel) may refer to the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component.
[0049] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, a unit may be used interchangeably with terms such as block or area. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients. Alternatively, a sample may refer to a pixel value in the spatial domain, and when such a pixel value is transformed into the frequency domain, it may refer to a transform coefficient in the frequency domain.
[0050] In some cases, a unit may be used interchangeably with terms such as a block or region. Typically, an MxN block may represent a sample or a set of transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or the value of a pixel, and may also represent a pixel / pixel value of only the luma component, or a pixel / pixel value of only the chroma component. A sample may be used as a term corresponding to a pixel or picture element configuring a picture (or image).
[0051] Figure 2 is a diagram schematically illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied. Hereinafter, a device referred to as a video encoding device may include an image encoding device.
[0052] Reference Figure 2 , the encoding device 200 includes an image partitioner 210, a predictor 220, a residual processor 230 and an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250, and the filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0053] The image partitioner 210 may partition an input image (or picture or frame) input to the encoding device 200 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, the coding unit may be recursively partitioned from a coding tree unit (CTU) or a largest coding unit (LCU) based on a quadtree, binary tree, and / or ternary tree (QTBTTT) structure. For example, a coding unit may be partitioned into multiple coding units of greater depth based on a quadtree, binary tree, and / or ternary structure. In this case, for example, the quadtree structure may be applied first, followed by the binary tree and / or ternary structure. Alternatively, the binary tree structure may be applied first. The coding process according to the present disclosure may be performed based on a final coding unit that is no longer partitioned. In this case, the largest coding unit may be used as the final coding unit based on image characteristics, coding efficiency, and the like. Alternatively, if necessary, the coding unit may be recursively partitioned into coding units of greater depth, and the coding unit of the optimal size may be used as the final coding unit. The coding process may include prediction, transformation, and reconstruction processes (described later). As another example, a processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or partitioned from the final coding unit. A prediction unit may be a unit for sample prediction, and a transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0054] The encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown, the unit in the encoder 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as the subtractor 231. The predictor can perform prediction on a processing target block (hereinafter referred to as the current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra prediction or inter prediction in units of the current block or CU. As described later in the description of each prediction mode, the predictor can generate various types of information regarding the prediction (e.g., prediction mode information) and send the generated information to the entropy encoder 240. The information regarding the prediction can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0055] The intra-frame predictor 222 can predict the current block with reference to samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or may be spaced apart. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, the non-directional mode may include a DC mode and a planar mode. For example, depending on the level of detail of the prediction direction, the directional mode may include 33 directional prediction modes or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. The intra-frame predictor 222 may use the prediction mode applied to the neighboring blocks to determine the prediction mode applied to the current block.
[0056] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. It can also include information about the inter-frame prediction direction (e.g., L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture containing the reference block and the reference picture containing the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture containing temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on the neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the neighboring block as the motion information of the current block. In skip mode, unlike merge mode, a residual signal may not be transmitted. In the case of motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling the motion vector difference.
[0057] The predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor 220 can apply intra prediction or inter prediction to predict a block, and can apply intra prediction and inter prediction at the same time. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can be based on an intra block copy (IBC) prediction mode or a palette mode for predicting blocks. The IBC prediction mode or palette mode can be used for image / video coding of content such as games, such as screen content coding (SCC). IBC basically performs prediction in the current picture, but it can be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this document. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample values in the picture can be signaled based on information about the palette table and the palette index.
[0058] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222 ) may be used to generate a reconstructed signal or may be used to generate a residual signal.
[0059] The transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of the following: discrete cosine transform (DCT), discrete sine transform (DST), graph-based transform (GBT), or conditional nonlinear transform (CNT). Here, when the relationship information between pixels is illustrated as a graph, GBT means a transform obtained from the graph. CNT means a transform obtained based on a prediction signal generated by using all previously reconstructed pixels. In addition, the transform process can also be applied to pixel blocks having the same square size, or can also be applied to variable-sized blocks that are not square.
[0060] The quantizer 233 quantizes the transform coefficients and transmits the quantized transform coefficients to the entropy encoder 240, and the entropy encoder 240 encodes the quantized signal (information about the quantized transform coefficients) and outputs the encoded signal as a bitstream. The information about the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange the quantized transform coefficients in the block form in a one-dimensional vector form based on the coefficient scanning order, and also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0061] The entropy encoder 240 can implement various encoding methods, such as Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can also encode information necessary for video / image reconstruction (e.g., syntax element values, etc.), in addition to quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of a network abstraction layer (NAL). The video / image information may also include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Furthermore, the video / image information may include general constraint information. In this document, the video / image information may include information and / or syntax elements signaled / transmitted from the encoding device to the decoding device. The video / image information may be encoded through the aforementioned encoding process and thus included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting a signal output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signal may be configured as an internal / external element of the encoding device 200, or the transmitting unit may also be included in the entropy encoder 240.
[0062] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, dequantization and inverse transform can be applied to the quantized transform coefficients by the dequantizer 234 and the inverse transform unit 235 to reconstruct the residual signal (residual block or residual sample). The adder 250 can add the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). When there is no residual for the processing target block, such as when the skip mode is applied, the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a restorer or a restored block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current picture, and can also be used for inter-frame prediction of the next picture after filtering, as described below.
[0063] Meanwhile, luma mapping and chroma scaling (LMCS) may also be applied during the picture encoding and / or reconstruction process.
[0064] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). For example, the various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various types of information related to filtering and transmit the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0065] The modified reconstructed picture transmitted to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus may be avoided and coding efficiency may be improved.
[0066] The DPB of the memory 270 can store the corrected reconstructed picture for use as a reference picture in the inter-frame predictor 221. The memory 270 can store the motion information of the block from which the motion information in the current picture was derived (or encoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 221 to be used as the motion information of the spatially adjacent block or the motion information of the temporally adjacent block. The memory 270 can store the reconstructed samples of the reconstructed block in the current picture and can transmit the reconstructed samples to the intra-frame predictor 222.
[0067] Figure 3 is a diagram for schematically explaining a configuration of a video / image decoding device to which an embodiment of the present disclosure is applicable.
[0068] Reference Figure 3 , the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. According to an embodiment, the entropy decoding 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be configured by hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0069] When a bit stream including video / image information is input, the decoding apparatus 300 can reconstruct the bit stream corresponding to the bit stream in FIG. Figure 2The image corresponding to the processing of video / image information in the encoding device is processed. For example, the decoding device 300 can derive a unit / block based on the block partition related information obtained from the bitstream. The decoding device 300 can perform decoding using the processing unit applied in the encoding device. Therefore, for example, the processing unit of decoding can be a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to the quadtree structure, binary tree structure and / or ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.
[0070] The decoding device 300 can receive the data in the form of a bit stream from Figure 2 The received signal is output by the encoding device, and the entropy decoder 310 can decode the received signal. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or general constraint information. The signaled / received information and / or syntax elements described later in this document may be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 310 may decode the information within the bitstream based on a coding method such as exponential Golomb coding, context-adaptive variable length coding (CAVLC), or context-adaptive binary arithmetic coding (CABAC), and output syntax elements required for image reconstruction and quantized values of transform coefficients for the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model by using the decoding target syntax element information, the decoding information of the decoding target block, or the information of the symbol / bin decoded in the previous stage, and perform arithmetic decoding on the bin by predicting the probability of the bin appearing according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. The information related to the prediction among the information decoded by the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (i.e., quantized transform coefficient and related parameter information) that has been entropy decoded in the entropy decoder 310 can be input to the residual processor 320.
[0071] The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. At the same time, a receiver (not shown) for receiving a signal output from the encoding device can also be configured as an internal / external element of the decoding device 300, or the receiver can be a component of the entropy decoder 310. At the same time, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of the following: a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0072] The dequantizer 321 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients in a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The dequantizer 321 may dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0073] The inverse transformer 322 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0074] The predictor 330 may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block and determine a specific intra / inter prediction mode based on the information on prediction output from the entropy decoder 310.
[0075] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can apply intra prediction or inter prediction to predict a block, and can apply intra prediction and inter prediction simultaneously. This can be referred to as combined inter and intra prediction (CIIP). In addition, the predictor can predict a block based on an intra block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for image / video coding of content such as games, such as screen content coding (SCC). IBC can basically perform prediction in the current picture, but can be performed similarly to inter prediction so that a reference block is derived within the current picture. That is, IBC can use at least one inter prediction technique described in this document. The palette mode can be considered an example of intra coding or intra prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0076] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or spaced apart from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 may determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0077] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector in a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. Motion information can include a motion vector and a reference picture index. Motion information can also include information regarding the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can construct a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index for the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information regarding the prediction can include information indicating the inter-frame prediction mode used for the current block.
[0078] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If there is no residual for the processing target block, for example, when skip mode is applied, the prediction block can be used as the reconstructed block.
[0079] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, and as described later, may also be output through filtering or may also be used for inter-frame prediction of the next picture.
[0080] In addition, luma mapping with chroma scaling (LMCS) can also be applied to the picture decoding process.
[0081] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in the memory 360, specifically, in the DPB of the memory 360. Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0082] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (decoded) and / or the motion information of the block in the reconstructed picture. The stored motion information can be transmitted to the inter-frame predictor 260 for use as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 360 can store the reconstructed samples of the reconstructed block in the current picture and transmit the reconstructed samples to the intra-frame predictor 331.
[0083] In this document, the embodiments described in the filter 260 , the inter predictor 221 , and the intra predictor 222 of the encoding apparatus 200 may be equally applied to or correspond to the filter 350 , the inter predictor 332 , and the intra predictor 331 .
[0084] The video / image coding method according to this document can be performed based on the following partition structure. Specifically, the prediction, residual processing ((inverse) transform and (de)quantization), syntax element coding and filtering processes to be described later can be performed based on the CTU and CU (and / or TU and PU) derived based on the partition structure. The block partitioning process can be performed by the image partitioner 210 of the above-mentioned encoding device, and the partition-related information can be processed (encoded) by the entropy encoder 240 and can be sent to the decoding device in the form of a bitstream. The entropy decoder 310 of the decoding device can derive the block partition structure of the current picture based on the partition-related information obtained from the bitstream, and based on this, a series of processes for image decoding (e.g., prediction, residual processing, block / picture reconstruction and in-loop filtering) can be performed. The CU size and the TU size can be equal to each other, or multiple TUs can exist in the CU area. At the same time, the CU size can generally represent the luma component (sample) coding block (CB) size. The TU size can generally represent the luma component (sample) transform block (TB) size. The chroma component (sample) CB or TB size can be derived based on the luma component (sample) CB or TB size according to the component ratio according to the color format (chroma format, for example, 4:4:4, 4:2:2, 4:2:0, etc.) of the picture / image. The TU size can be derived based on maxTbSize. For example, if the CU size is larger than maxTbSize, multiple TUs (TBs) of maxTbSize can be derived, and transformation / inverse transformation can be performed in units of TUs (TBs). In addition, for example, when intra prediction is applied, the intra prediction mode / type can be derived in units of CUs (or CBs), and the derivation of neighboring reference samples and the generation of prediction samples can be performed in units of TUs (or TBs). In this case, one or more TUs (or TBs) may exist in one CU (or CB) area, and in this case, multiple TUs (or TBs) may share the same intra prediction mode / type.
[0085] In addition, when encoding a video / image according to this document, the image processing unit may have a hierarchical structure. A picture may be divided into one or more tiles, tiles, slices, and / or tile groups. A slice may include one or more tiles. A tile may include one or more CTU rows in a tile. A slice may include an integer number of tiles in a picture. A tile group may include one or more tiles. A tile is a rectangular area of a CTU within a specific tile column and a specific tile row in a picture. A tile group may include an integer number of tiles according to a tile raster scan in a picture. A slice header may carry information / parameters that can be applied to the corresponding slice (block in a slice). If the encoding / decoding device has a multi-core processor, the encoding / decoding processes for tiles, slices, tiles, and / or tile groups may be processed in parallel. In this document, slices and tile groups may be used interchangeably. That is, the tile group header may be referred to as a slice header. Here, a slice may have one of the slice types including intra (I) slices, predicted (P) slices, and bidirectionally predicted (B) slices. For prediction of blocks in I slices, inter-frame prediction may not be used, but only intra-frame prediction may be used. Even in this case, the original sample values can be coded and signaled without prediction. For blocks in P slices, intra-frame prediction or inter-frame prediction may be used, and when inter-frame prediction is used, only unidirectional prediction may be used. Meanwhile, for blocks in B slices, intra-frame prediction or inter-frame prediction may be used, and when inter-frame prediction is used, maximum reachable bidirectional prediction may be used.
[0086] Depending on the characteristics of the video image (e.g., resolution), or taking into account coding efficiency or parallel processing, the encoder can determine the block / block group, tile, slice, maximum and minimum coding unit sizes, and can include corresponding information or information that can summarize the corresponding information in the bitstream.
[0087] The decoder can obtain information indicating whether a tile / tile group, tile, slice, or CTU in the current picture has been partitioned into multiple coding units. Efficiency can be improved by obtaining (sending) this information only under specific conditions.
[0088] Figure 4 The hierarchical structure for coding images / videos is shown as an example.
[0089] refer to Figure 4 The coded image / video is divided into the VCL (Video Coding Layer) that handles the image / video decoding process and itself, the subsystem that sends and stores coding information, and the Network Abstraction Layer (NAL) that exists between the VCL and the subsystem and is responsible for network adaptation functions.
[0090] VCL can generate VCL data including compressed image data (slice data), or generate parameter sets including picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS), etc., or supplementary enhancement information (SEI) messages additionally required for the decoding process of images.
[0091] In NAL, a NAL unit can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in the VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header may include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0092] As shown in the figure, NAL units can be divided into VCL NAL units and non-VCL NAL units according to the RBSP generated in the VCL. A VCL NAL unit may refer to a NAL unit including information about an image (slice data), while a non-VCL NAL unit may refer to a NAL unit containing information necessary for decoding an image (parameter set or SEI message).
[0093] By appending header information according to the data standard of the subsystem, the VCL NAL units and non-VCL NAL units can be transmitted over the network. For example, the NAL units can be converted into a predetermined standard data format such as the H.266 / VVC file format, the Real-time Transport Protocol (RTP), or the Transport Stream (TS) and transmitted over various networks.
[0094] As described above, in a NAL unit, a NAL unit type may be specified according to an RBSP data structure included in a corresponding NAL unit, and information about this NAL unit type may be stored and signaled in a NAL unit header.
[0095] For example, NAL units can be roughly classified into VCL NAL unit types and non-VCL NAL unit types depending on whether the NAL unit includes information about an image (slice data). VCL NAL unit types can be classified according to the characteristics and type of the picture included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the type of parameter set.
[0096] The following is an example of a NAL unit type specified according to a parameter set type included in a non-VCL NAL unit type.
[0097] -APS (Adaptation Parameter Set) NAL unit: type used for NAL units containing APS
[0098] -DPS (Decoding Parameter Set) NAL unit: type used for NAL units containing DPS
[0099] - VPS (Video Parameter Set) NAL unit: type used for NAL units containing VPS
[0100] - SPS (Sequence Parameter Set) NAL unit: type used for NAL units containing SPS
[0101] -PPS (Picture Parameter Set) NAL unit: type used for NAL units containing PPS
[0102] - PH (Picture Header) NAL unit: type used for NAL units including PH
[0103] The above-mentioned NAL unit type has syntax information for the NAL unit type, and the syntax information can be stored and signaled in the NAL unit header. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.
[0104] At the same time, as described above, a picture may include multiple slices, and a slice may include a slice header and slice data. In this case, a picture header may be further added for multiple slices in a picture (a set of slice headers and slice data). The picture header (picture header syntax) may include information / parameters that can be commonly applied to the picture. The slice header (slice header syntax) may include information / parameters that can be commonly applied to the slices. The adaptation parameter set (APS) or the picture parameter set (PPS) may include information / parameters that can be commonly applied to one or more pictures. The sequence parameter set (SPS) may include information / parameters that can be commonly applied to one or more sequences. The video parameter set (VPS) may include information / parameters that can be commonly applied to multiple layers. The decoding parameter set (DPS) may include information / parameters that can be commonly applied to the entire video. The DPS may include information / parameters related to the concatenation of the coded video sequence (CVS).
[0105] In this document, the high-level syntax may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0106] Furthermore, for example, information on partitioning and configuration of tiles / tile groups / tiles / slices may be configured by the encoding end through a higher-level syntax and may be transmitted to the decoding device in the form of a bitstream.
[0107] In this article, the image / video information encoded in the form of a bitstream from the encoding device and notified to the decoding device using a signal may include not only information related to intra-picture partitions, intra-frame / inter-frame prediction information, residual information, in-loop filtering information, etc., but also information included in a slice header, information included in a picture header, information included in an APS, information included in a PPS, information included in an SPS, information included in a VPS, and / or information included in a DPS. In addition, the image / video information may also include information in a NAL unit header.
[0108] In this document, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or may still be referred to as a transform coefficient for consistency of expression.
[0109] In this document, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and information about the transform coefficients may be signaled using residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficients), and scaled transform coefficients may be derived by inverse transforming (scaling) the transform coefficients. Residual samples may be derived based on the inverse transform (transform) of the scaled transform coefficients. This may also be applied / expressed in other parts of this document.
[0110] As described above, the encoding device can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). In addition, the decoding device can decode the information in the bitstream based on a coding method such as exponential Golomb, CAVLC, or CABAC, and can output the values of syntax elements necessary for image reconstruction and the quantized values of the transform coefficients for the residual. For example, the above-mentioned coding method can be performed as described later.
[0111] In this document, intra prediction may refer to the generation of prediction samples for the current block based on reference samples in the picture to which the current block belongs (hereinafter, the current picture). When intra prediction is applied to the current block, neighboring reference samples to be used for intra prediction of the current block may be derived. The neighboring reference samples of the current block may include samples adjacent to the left boundary of the current block (nWxnH) and a total of 2xnH samples adjacent to the left bottom, samples adjacent to the top boundary of the current block and a total of 2xnW samples adjacent to the right top, and one sample adjacent to the left top of the current block. In addition, the neighboring reference samples of the current block may include multiple columns of top neighboring samples and multiple rows of left neighboring samples. In addition, the neighboring reference samples of the current block may include a total of nH samples adjacent to the right boundary of the current block (nWxnH), a total of nW samples adjacent to the bottom boundary of the current block, and one sample adjacent to the right bottom of the current block.
[0112] However, some neighboring reference samples of the current block may not have been decoded or enabled. In this case, the decoding device can configure the neighboring reference samples to be used for prediction by replacing unenabled samples with enabled samples. In addition, the neighboring reference samples to be used for prediction can be configured by interpolating enabled samples.
[0113] If neighboring reference samples are derived, prediction samples can be (i) summarized based on the average or interpolation of neighboring reference samples of the current block, and (ii) summarized based on reference samples located in a specific (prediction) direction of the prediction samples among the neighboring reference samples of the current block. Case (i) is referred to as non-directional mode or non-angular mode, and case (ii) is referred to as directional mode or angular mode. Furthermore, prediction samples can be generated by interpolating first and second neighboring samples located in a direction opposite to the prediction direction of the intra prediction mode for the current block based on the prediction samples of the current block among neighboring reference samples. This is referred to as linear interpolation intra prediction (LIP). Furthermore, chroma prediction samples can be generated based on luma samples using a linear model. This is referred to as LM mode. Furthermore, temporary prediction samples for the current block can be derived based on filtered neighboring reference samples, and the prediction samples for the current block can be derived by calculating a weighted sum of the temporary prediction samples and at least one reference sample derived according to the intra prediction mode (i.e., an unfiltered neighboring reference sample) among the existing neighboring reference samples. This situation can be referred to as directionally dependent intra prediction (PDPC). Furthermore, prediction samples can be derived by using reference samples located in the prediction direction on a reference sample line with the highest prediction accuracy among multiple reference sample lines adjacent to the current block through corresponding line selection. In this case, intra prediction coding can be performed in a method for indicating (signaling) the reference sample line to be used to the decoding device. This situation can be referred to as multiple reference line (MRL) intra prediction or MRL-based intra prediction. Furthermore, intra prediction can be performed based on the same intra prediction mode by dividing the current block into vertical or horizontal subpartitions, and adjacent reference samples can be derived and used on a subpartition basis. That is, in this case, since the intra prediction mode for the current block is applied equally to the subpartitions, and adjacent reference samples are derived and used on a subpartition basis, intra prediction performance can be improved in some cases. This prediction method can be referred to as intra subpartition (ISP) or ISP-based intra prediction. The above intra prediction methods can be referred to as intra prediction types to distinguish them from intra prediction modes. Intra-frame prediction types may be referred to by various terms such as intra-frame prediction techniques or additional intra-frame prediction modes. For example, an intra-frame prediction type (or additional intra-frame prediction mode) may include at least one of the aforementioned LIP, PDPC, MRL, or ISP. A general intra-frame prediction method that excludes specific intra-frame prediction types such as LIP, PDPC, MRL, or ISP may be referred to as a normal intra-frame prediction type. In the absence of a specific intra-frame prediction type, a normal intra-frame prediction type may generally be applied, and prediction may be performed based on the aforementioned intra-frame prediction modes. Furthermore, post-filtering may be performed on the derived prediction samples as needed.
[0114] Specifically, the intra prediction process may include the steps of determining an intra prediction mode / type, deriving neighboring reference samples, and deriving prediction samples based on the intra prediction mode / type. In addition, a post-filtering step may be performed on the derived prediction samples as needed.
[0115] In addition to the aforementioned prediction types, affine linear weighted intra prediction (ALWIP) can also be used. ALWIP can be referred to as linear weighted intra prediction (LWIP), matrix weighted intra prediction (MIP), or matrix-based intra prediction. When MIP is applied to the current block, i) neighboring reference samples that have been averaged are used, ii) a matrix-vector multiplication process can be performed, and iii) prediction samples for the current block can be derived by further performing horizontal / vertical interpolation as needed. The intra prediction mode used for MIP can be configured differently from the aforementioned LIP, PDPC, MRL, or ISP intra prediction, or the intra prediction mode can be used for normal intra prediction. The intra prediction mode used for MIP can be referred to as MIP intra prediction mode, MIP prediction mode, or MIP mode. For example, depending on the intra prediction mode used for MIP, the matrix and offset used for matrix-vector multiplication can be configured differently. Here, the matrix can be referred to as a (MIP) weighting matrix, and the offset can be referred to as a (MIP) offset vector or a (MIP) bias vector.
[0116] In image / video coding, the pictures constituting the image / video can be encoded / decoded according to the decoding order. The picture order corresponding to the output order of the decoded pictures can be set differently from the decoding order, based on which backward prediction and forward prediction can also be performed in inter-frame prediction.
[0117] Figure 5 This shows an example of the picture decoding process.
[0118] Figure 5 An example of a schematic image decoding process to which embodiments of this document can be applied is shown. Figure 5 In the above Figure 3 S500 may be performed in the entropy decoder 310 of the decoding device described in
[15] ; S510 may be performed in the predictor 330; S520 may be performed in the residual processor 320; S530 may be performed in the adder 340; and S540 may be performed in the filter 350. S500 may include the information decoding process described herein; S510 may include the inter / intra prediction process described in this document; S520 may include the residual processing process described in this document; S530 may include the block / picture reconstruction process described in this document; and S540 may include the in-loop filtering process described in this document.
[0119] refer to Figure 5 , such as about Figure 3 As described in this specification, the picture decoding process may illustratively include a process (S500) for obtaining image / video information from a bitstream (via decoding), picture reconstruction processes (S510 to S530), and an in-loop filtering process (S540) for reconstructing the picture. The picture reconstruction process may be performed based on residual samples and prediction samples obtained through the inter / intra prediction (S510) and residual processing (S520) (dequantization and inverse transformation of quantized transform coefficients) processes described herein. By performing an in-loop filtering process on the reconstructed picture generated by the picture reconstruction process, a modified reconstructed picture may be generated, which may be output as a decoded picture and may also be stored in a decoded picture buffer or memory 360 of a decoding device and used as a reference picture in subsequent inter-frame prediction processes for picture decoding. Depending on the situation, the in-loop filtering process may be skipped, in which case the reconstructed picture may be output as a decoded picture and may also be stored in a decoded picture buffer or memory 360 of a decoding device and used as a reference picture in subsequent inter-frame prediction processes for picture decoding. The in-loop filtering process S540 may include the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and / or the bilateral filtering process as described above, and all or some of them may be skipped. In addition, one or some of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and the bilateral filtering process may be applied sequentially, or all of them may be applied sequentially. For example, after the deblocking filtering process is applied to the reconstructed picture, the SAO process may be performed thereon. Alternatively, for example, after the deblocking filtering process is applied to the reconstructed picture, the ALF process may be performed thereon. This may also be performed in the encoding device.
[0120] Figure 6 An example showing the picture encoding process.
[0121] Figure 6 An example of a schematic diagram encoding process to which embodiments of this document can be applied is shown. Figure 6 In the above Figure 2 S600 may be performed in the predictor 220 of the encoding device described in
[15] ; S610 may be performed in the residual processor 230; and S620 may be performed in the entropy encoder 240. S600 may include the inter / intra prediction process described in this document; S610 may include the residual processing process described in this document; and S620 may include the information encoding process described in this document.
[0122] refer to Figure 6 , as in about Figure 2As shown in the description, the picture encoding process can schematically include a process of generating a reconstructed picture of the current picture and a process of applying in-loop filtering to the reconstructed picture (optional), as well as a process of encoding information used for picture reconstruction (e.g., prediction information, residual information, partition information, etc.) and outputting it in the form of a bitstream. The encoding device can derive (modified) residual samples from the quantized transform coefficients through the dequantizer 234 and the inverse transformer 235, and can generate a reconstructed picture based on the (modified) residual samples and the prediction samples as the output of S600. The reconstructed picture generated in this manner can be the same as the above-mentioned reconstructed picture generated in the decoding device. By performing an in-loop filtering process on the reconstructed picture, similar to the case of the decoding device, a modified reconstructed picture can be generated, which can be stored in the decoded picture buffer or memory 270 and used as a reference picture in the inter-frame prediction process of the subsequent picture encoding. As described above, all or part of the in-loop filtering process can be skipped depending on the situation. In the case of performing an in-loop filtering process, (in-loop) filtering-related information (parameters) can be encoded in the entropy encoder 240 and output in the form of a bit stream, and the decoding device can perform the in-loop filtering process in the same manner as the encoding device based on the filtering-related information.
[0123] This in-loop filtering process can reduce noise generated during image / video coding, such as blocking artifacts and ringing artifacts, and improve subjective / objective visual quality. In addition, when the in-loop filtering process is performed in both the encoding device and the decoding device, the encoding device and the decoding device can derive the same prediction results, improve the reliability of picture coding, and reduce the amount of data to be transmitted for picture coding.
[0124] As described above, the picture reconstruction process can be performed in the encoding device and in the decoding device. Based on the intra prediction / inter prediction of each block unit, a reconstructed block can be generated, and a reconstructed picture including the reconstructed block can be generated. In the case where the current picture / slice / patchwork group is an I picture / slice / patchwork group, the blocks included in the current picture / slice / patchwork group can be reconstructed based only on intra prediction. At the same time, in the case where the current picture / slice / patchwork group is a P or B picture / slice / patchwork group, the blocks included in the current picture / slice / patchwork group can be reconstructed based on intra prediction or inter prediction. In this case, inter prediction can be applied to some blocks in the current picture / slice / patchwork group, and intra prediction can be applied to some of the remaining blocks. The color components of the picture may include a luminance component and a chrominance component, and unless explicitly limited by this document, the methods and embodiments proposed in this document may be applied to the luminance component and the chrominance component.
[0125] Meanwhile, the video / image encoding process based on inter-frame prediction may schematically include, for example, the following.
[0126] Figure 7 An example of a video / image encoding method based on inter-frame prediction is shown.
[0127] refer to Figure 7 , the encoding device performs inter-frame prediction on the current block (S700). The encoding device can derive the inter-frame prediction mode and motion information of the current block and generate prediction samples for the current block. Here, the inter-frame prediction mode determination, motion information derivation and prediction sample generation processes can be performed simultaneously or one after another. For example, the inter-frame predictor of the encoding device may include a prediction mode determiner, a motion information deriver and a prediction sample deriver. The prediction mode determiner can determine the prediction mode for the current block, the motion information deriver can derive the motion information of the current block, and the prediction sample deriver can derive the prediction samples for the current block. For example, the inter-frame predictor of the encoding device can search for blocks similar to the current block in a certain area (search area) of the reference picture through motion estimation, and derive a reference block whose difference with the current block is the smallest or less than or equal to a certain level. Based on this, a reference picture index indicating the reference picture above which the reference block is located can be derived, and based on the position difference between the reference block and the current block, a motion vector can be derived. The encoding device can determine a mode from among the various prediction modes applied to the current block. The encoding apparatus may compare rate-distortion (RD) costs of various prediction modes and determine an optimal prediction mode for the current block.
[0128] For example, when skip mode or merge mode is applied to the current block, the encoding device may construct a merge candidate list and derive a reference block whose difference with the current block is the smallest or less than or equal to a certain level from the reference blocks indicated by the merge candidates included in the merge candidate list. In this case, a merge candidate associated with the derived reference block may be selected, and merge index information indicating the selected merge candidate may be generated and signaled to the decoding device. The motion information of the selected merge candidate may be used to derive the motion information of the current block.
[0129] As another example, when the (A)MVP mode is applied to the current block, the encoding device may construct an (A)MVP candidate list and use the motion vector of an MVP candidate selected from among the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. In this case, for example, a motion vector indicating a reference block derived by the above-mentioned motion estimation may be used as the motion vector of the current block, and among the MVP candidates, the MVP candidate having the motion vector with the smallest difference from the motion vector of the current block may be the selected MVP candidate. An MVD (motion vector difference) may be derived, which is the difference obtained by subtracting the MVP from the motion vector of the current block. In this case, information about the MVD may be signaled to the decoding device. Additionally, when the (A)MVP mode is applied, the value of the reference picture index may be configured as reference picture index information and separately signaled to the decoding device.
[0130] The encoding apparatus may derive residual samples based on the prediction samples ( S710 ).The encoding apparatus may derive residual samples by comparing original samples and prediction samples of the current block.
[0131] The encoding device encodes the image information including prediction information and residual information (S720). The encoding device can output the encoded image information in the form of a bitstream. The prediction information may include prediction mode information (e.g., skip flag, merge flag, mode index, etc.) and information about motion information as information about the prediction process. The information about the motion information may include candidate selection information (e.g., merge index, mvp flag or mvp index), which is information for deriving a motion vector. In addition, the information about the motion information may include information about the above-mentioned MVD and / or reference picture index information. In addition, the information about the motion information may include information indicating whether L0 prediction, L1 prediction or bidirectional prediction is applied. The residual information is information about the residual sample. The residual information may include information about the quantized transform coefficients used for the residual sample.
[0132] The output bitstream may be stored in a (digital) storage medium and transmitted to the decoding device, or may be transmitted to the decoding device over a network.
[0133] At the same time, as described above, the encoding device can generate a reconstructed picture (including reconstructed samples and reconstructed blocks) based on the reference samples and the residual samples. This is to derive the same prediction result as the prediction result performed in the decoding device in the encoding device, and the reason is that coding efficiency can be improved through this. Therefore, the encoding device can store the reconstructed picture (or reconstructed sample, reconstructed block) in a memory and use it as a reference picture for inter-frame prediction. As described above, the reconstructed picture can be further applied with an in-loop filtering process or the like.
[0134] The video / image decoding process based on inter-frame prediction may schematically include, for example, the following.
[0135] Figure 8 An example of a video / image decoding method based on inter-frame prediction is shown.
[0136] The decoding device may perform operations corresponding to those already performed in the encoding device.The decoding device may perform prediction on the current block and derive a prediction sample based on the received prediction information.
[0137] Specifically, refer to Figure 8 The decoding device may determine a prediction mode for the current block based on the prediction information received from the bitstream (S800). The decoding device may determine which inter-frame prediction mode to apply to the current block based on the prediction mode information in the prediction information.
[0138] For example, it is possible to determine whether to apply the merge mode to the current block, or to determine the (A)MVP mode based on the merge flag. Alternatively, an inter-frame prediction mode can be selected from various inter-frame prediction mode candidates based on the merge index. The inter-frame prediction mode candidates may include various inter-frame prediction modes, such as skip mode, merge mode, and / or (A)MVP mode.
[0139] The decoding device derives the motion information of the current block based on the determined inter-frame prediction mode (S810). For example, when skip mode or merge mode is applied to the current block, the decoding device may construct a merge candidate list to be described later and select one of the merge candidates included in the merge candidate list. The selection may be performed based on the above-mentioned selection information (merge index). The motion information of the selected merge candidate may be used to derive the motion information of the current block. The motion information of the selected merge candidate may be used as the motion information of the current block.
[0140] As another example, when the (A)MVP mode is applied to the current block, the decoding device may construct an (A)MVP candidate list and use the motion vector of an MVP candidate selected from among the MVP (motion vector predictor) candidates included in the (A)MVP candidate list as the MVP of the current block. The selection may be performed based on the above-mentioned selection information (MVP flag or MVP index). In this case, the MVD of the current block may be derived based on the information about the MVD, and the motion vector of the current block may be derived based on the MVD and MVP of the current block. In addition, the reference picture index of the current block may be derived based on the reference picture index information. The picture in the reference picture list related to the current block indicated by the reference picture index may be derived as the reference picture referenced for inter-frame prediction of the current block.
[0141] Meanwhile, the motion information of the current block may be derived without constructing a candidate list, and in this case, the construction of the candidate list as described above may be omitted.
[0142] The decoding device may generate prediction samples for the current block based on the motion information of the current block (S820). In this case, a reference picture may be derived based on a reference picture index of the current block, and the prediction samples for the current block may be derived using samples of a reference block on the reference picture indicated by the motion vector of the current block. In this case, a prediction sample filtering process may be further performed on all or some of the prediction samples for the current block, as described later, depending on the situation.
[0143] For example, the inter-frame predictor of the encoding device may include a prediction mode determiner, a motion information exporter, and a prediction sample exporter. It can determine the prediction mode for the current block based on the prediction mode information received at the prediction mode determiner, can export the motion information (motion vector and / or reference picture index, etc.) of the current block based on the information about the motion information received at the motion information exporter, and can export the prediction sample of the current block at the prediction sample exporter.
[0144] The decoding device generates residual samples of the current block based on the received residual information (S830). The decoding device can generate reconstructed samples of the current block based on the residual samples and the predicted samples, and generate a reconstructed picture based on these reconstructed samples (S840). Hereinafter, the in-loop filtering process can be applied to the reconstructed picture as described above.
[0145] At the same time, as described above, high-level syntax (HLS) can be compiled / signaled for video / image coding. A coded picture can be composed of one or more slices. Parameters describing the coded picture are signaled in the picture header, and parameters describing the slice are signaled in the slice header. The picture header is carried in the form of the NAL unit itself. The slice header is present at the beginning of the NAL unit that includes the payload of the slice (i.e., the slice data).
[0146] Each picture is associated with a picture header. A picture can be composed of different types of slices: intra-coded slices (i.e., I slices) and inter-coded slices (i.e., P slices and B slices). Therefore, the picture header can include syntax elements necessary for intra slices and inter slices of the picture. For example, the syntax of the picture header can be as shown in Table 1 below.
[0147] [Table 1]
[0148]
[0149]
[0150]
[0151]
[0152]
[0153] Among the syntax elements of Table 1, the syntax element including "intra_slice" in its title (e.g., pic_log2_diff_min_qt_min_cb_intra_slice_luma) is a syntax element being used in the I slice of the corresponding picture, and the syntax elements related to the syntax elements including "inter_slice" in its title (e.g., pic_log2_diff_min_qt_min_cb_inter_slice, mvp, mvd, mmvd, and merge) (e.g., pic_temporal_mvp_enabled_flag) are syntax elements being used in the P slice and / or B slice of the corresponding picture.
[0154] That is, the picture header includes all syntax elements required for intra-coded slices and syntax elements required for inter-coded slices for each single picture. However, this is only useful for pictures that include mixed-type slices (pictures that include all intra-coded slices and inter-coded slices). Generally speaking, since a picture does not include mixed-type slices (i.e., a general picture includes only intra-coded slices or only inter-coded slices), it is not necessary to perform signaling of all data (syntax elements used in intra-coded slices and syntax elements used in inter-coded slices).
[0155] The following drawings have been prepared to illustrate detailed examples of this document. Since the names of detailed devices or the names of detailed signals / information are presented exemplarily, the technical features of this document are not limited to the detailed names used in the following drawings.
[0156] This document provides the following methods to solve the above problems. The items of each method can be applied individually or in combination.
[0157] 1. A flag in the picture header that can be used to signal whether the syntax elements required to specify intra-coded slices are present in the picture header. This flag can be called intra_signaling_present_flag.
[0158] a) When intra_signaling_present_flag is equal to 1, the syntax elements required for intra-coded slices are present in the picture header. Similarly, when intra_signaling_present_flag is equal to 0, the syntax elements required for intra-coded slices are not present in the picture header.
[0159] b) When the picture associated with the picture header has at least one intra-coded slice, the value of intra_signaling_present_flag in the picture header shall be equal to 1.
[0160] c) Even when a picture associated with a picture header does not have intra-coded slices, the value of intra_signaling_present_flag in the picture header may be equal to 1.
[0161] d) When a picture has one or more sub-pictures containing only intra-coded slices and it is expected that the one or more sub-pictures can be extracted and merged with a sub-picture containing one or more inter-coded slices, the value of intra_signaling_present_flag should be set equal to 1.
[0162] 2. A flag in the picture header that can be used to signal whether the syntax elements required to specify inter-frame-only coded slices are present in the picture header. This flag can be called inter_signaling_present_flag.
[0163] a) When inter_signaling_present_flag is equal to 1, the syntax elements required for inter-coded slices are present in the picture header. Similarly, when inter_signaling_present_flag is equal to 0, the syntax elements required for inter-coded slices are not present in the picture header.
[0164] b) The value of inter_signaling_present_flag in the picture header shall be equal to 1 when the picture associated with the picture header has at least one inter-coded slice.
[0165] c) Even when a picture associated with a picture header does not have inter-coded slices, the value of inter_signaling_present_flag in the picture header may be equal to 1.
[0166] d) When a picture has one or more sub-pictures containing only inter-coded slices and it is expected that the one or more sub-pictures can be extracted and merged with a sub-picture containing one or more intra-coded slices, the value of inter_signaling_present_flag should be set equal to 1.
[0167] 3. The above flags (intra_signaling_present_flag and inter_signaling_present_flag) may be signaled in other parameter sets such as picture parameter set (PPS) instead of in the picture header.
[0168] 4. Another alternative for signaling the above flags could be as follows.
[0169] a) Two variables, IntraSignalingPresentFlag and InterSignalingPresentFlag, may be defined to specify whether syntax elements required for intra-coded slices and syntax elements required for inter-coded slices are respectively present in the picture header.
[0170] b) A flag called mixed_slice_types_present_flag in the picture header may be signaled. When mixed_slice_types_present_flag is equal to 1, the values of IntraSignalingPresentFlag and InterSignalingPresentFlag are set equal to 1.
[0171] c) When mixed_slice_types_present_flag is equal to 0, an additional flag called intra_slice_only_flag may be signaled in the picture header and the following applies. If intra_slice_only_flag is equal to 1, the value of IntraSignalingPresentFlag is set to 1 and the value of InterSignalingPresentFlag is set to 0. Otherwise, the value of IntraSignalingPresentFlag is set to 0 and the value of InterSignalingPresentFlag is set to 1.
[0172] 5. A fixed-length syntax element in the picture header may be signaled, which may be called slice_types_idc, specifying the following information.
[0173] a) Whether the picture associated with the picture header contains only intra-coded slices. For this type, the value of slice_types_idc can be set to 0.
[0174] b) Whether the picture associated with the picture header contains only inter-coded slices. The value of slice_types_idc may be set to 1.
[0175] c) Whether a picture associated with a picture header can contain intra-coded slices and inter-coded slices. The value of slice_types_idc can be set to 2.
[0176] Note that when slice_types_idc has a value equal to 2, it is still possible that a picture contains only intra-coded slices or only inter-coded slices.
[0177] d) Other values of slice_types_idc may be reserved for future use.
[0178] 6. For the slice_types_idc semantics in the picture header, the following constraints may be further specified.
[0179] a) When the picture associated with the picture header has one or more intra-coded slices, the value of slice_types_idc shall not be equal to 1.
[0180] b) When the picture associated with the picture header has one or more inter-coded slices, the value of slice_types_idc shall not be equal to 0.
[0181] 7. slice_types_idc may be signaled in other parameter sets such as picture parameter set (PPS) instead of in the picture header.
[0182] As an embodiment, the encoding device and the decoding device may use the following Table 2 and Table 3 as the syntax and semantics of the picture header based on Methods 1 and 2 described above.
[0183] [Table 2]
[0184]
[0185]
[0186] [Table 3]
[0187]
[0188]
[0189] Referring to Tables 2 and 3, if the value of intra_signaling_present_flag is 1, this may indicate that syntax elements used only in intra-coded slices are present in the picture header. If the value of intra_signaling_present_flag is 0, this may indicate that syntax elements used only in intra-coded slices are not present in the picture header. Therefore, if the picture associated with the picture header includes one or more slices with an I-slice slice type, the value of intra_signaling_present_flag becomes 1. In addition, if the picture associated with the picture header does not include a slice with an I-slice slice type, the value of intra_signaling_present_flag becomes 0.
[0190] If the value of inter_signaling_present_flag is 1, this may indicate that syntax elements used only in inter-coded slices are present in the picture header. If the value of inter_signaling_present_flag is 0, this may indicate that syntax elements used only in inter-coded slices are not present in the picture header. Therefore, if the picture associated with the picture header includes one or more slices having a slice type of P slice and / or B slice, the value of intra_signaling_present_flag becomes 1. In addition, if the picture associated with the picture header does not include a slice having a slice type of P slice and / or B slice, the value of intra_signaling_present_flag becomes 0.
[0191] Furthermore, in a case where a picture includes one or more sub-pictures including intra-coded slices that can be merged with one or more sub-pictures including inter-coded slices, the value of intra_signaling_present_flag and the value of inter_signaling_present_flag are both set to 1.
[0192] For example, in the case where only inter-coded slices (P slices and / or B slices) are included in the current picture, the encoding apparatus may determine the value of inter_signaling_present_flag to 1 and the value of intra_signaling_present_flag to 0.
[0193] As another example, in the case where only intra-coded slices (I slices) are included in the current picture, the encoding apparatus may determine the value of inter_signaling_present_flag to be 0 and the value of intra_signaling_present_flag to be 1.
[0194] As yet another example, in a case where at least one inter-coded slice or at least one intra-coded slice is included in the current picture, the encoding apparatus may determine the value of inter_signaling_present_flag and the value of intra_signaling_present_flag to be 1 in total.
[0195] In the case where the value of intra_signaling_present_flag is determined to be 0, the encoding device may generate image information in which syntax elements necessary for intra slices are excluded or omitted and only syntax elements necessary for inter slices are included in the picture header. If the value of inter_signaling_present_flag is determined to be 0, the encoding device may generate image information in which syntax elements necessary for inter slices are excluded or omitted and only syntax elements necessary for intra slices are included in the picture header.
[0196] If the value of inter_signaling_present_flag obtained from the picture header in the image information is 1, the decoding device can determine that at least one inter-coded slice is included in the corresponding picture, and can parse the syntax elements necessary for intra-frame prediction from the picture header. If the value of inter_signaling_present_flag is 0, the decoding device can determine that only intra-coded slices are included in the corresponding picture, and can parse the syntax elements necessary for intra-frame prediction from the picture header. If the value of intra_signaling_present_flag obtained from the picture header in the image information is 1, the decoding device can determine that at least one intra-coded slice is included in the corresponding picture, and can parse the syntax elements necessary for intra-frame prediction from the picture header. If the value of intra_signaling_present_flag is 0, the decoding device can determine that only inter-coded slices are included in the corresponding picture, and can parse the syntax elements necessary for inter-frame prediction from the picture header.
[0197] As another embodiment, the encoding device and the decoding device may use the following Table 4 and Table 5 as the syntax and semantics of the picture header based on the above methods 5 and 6.
[0198] [Table 4]
[0199]
[0200]
[0201] [Table 5]
[0202]
[0203]
[0204] Referring to Tables 4 and 5, if the value of slice_types_idc is 0, this indicates that the type of all slices in the picture associated with the picture header is I slice. If the value of slice_types_idc is 1, this indicates that the type of all slices in the picture associated with the picture header is P or B slice. If the value of slice_types_idc is 2, this indicates that the slice type of the slices in the picture associated with the picture header is I, P, and / or B slice.
[0205] For example, if only intra-coded slices are included in the current picture, the encoding device may determine the value of slice_types_idc to be 0 and may include syntax elements required for decoding only intra slices in the picture header. That is, in this case, the syntax elements required for inter slices are not included in the picture header.
[0206] As another example, if only inter-coded slices are included in the current picture, the encoding device may determine the value of slice_types_idc to be 1 and may include syntax elements required for decoding only inter-frame slices in the picture header. That is, in this case, the syntax elements required for intra-frame slices are not included in the picture header.
[0207] As another example, if at least one inter-coded slice and at least one intra-coded slice are included in the current picture, the encoding device may determine the value of slice_types_idc to be 2, and may include all of the syntax elements required for decoding of the inter-frame slices and the syntax elements required for decoding of the intra-frame slices in the picture header.
[0208] If the value of slice_types_idc obtained from the picture header in the image information is 0, the decoding device can determine that only intra-coded slices are included in the corresponding picture, and can parse the syntax elements necessary for decoding the intra-coded slices from the picture header. If the value of slice_types_idc is 1, the decoding device can determine that only inter-coded slices are included in the corresponding picture, and can parse the syntax elements necessary for decoding the inter-coded slices from the picture header. If the value of slice_types_idc is 2, the decoding device can determine that at least one intra-coded slice and at least one inter-coded slice are included in the corresponding picture, and can parse the syntax elements necessary for decoding the intra-coded slices and the syntax elements necessary for decoding the inter-coded slices from the picture header.
[0209] As another embodiment, the encoding device and the decoding device may use a flag indicating whether a picture includes intra-coded slices and inter-coded slices. If the flag is true, that is, if the value of the flag is 1, all intra-frame slices and inter-frame slices may be included in the corresponding picture. In this case, the following Tables 6 and 7 may be used as the syntax and semantics of the picture header.
[0210] [Table 6]
[0211]
[0212]
[0213]
[0214] [Table 7]
[0215]
[0216] Referring to Tables 6 and 7, if the value of mixed_slice_signaling_present_flag is 1, this may indicate that a picture associated with a corresponding picture header has one or more slices of different types. If the value of mixed_slice_signaling_present_flag is 0, this may mean that a picture associated with a corresponding picture header includes data associated with only a single slice type.
[0217] The variables InterSignalingPresentFlag and IntraSignalingPresentFlag indicate whether the syntax elements required for intra-coded slices and inter-coded slices are present in the corresponding picture header, respectively. If the value of mixed_slice_signaling_present_flag is 1, the values of IntraSignalingPresentFlag and InterSignalingPresentFlag are set to 1.
[0218] If the value of intra_slice_only_flag is set to 1, this means that the value of IntraSignalingPresentFlag is set to 1, and the value of InterSignalingPresentFlag is set to 0. If the value of intra_slice_only_flag is set to 0, this means that the value of IntraSignalingPresentFlag is set to 0, and the value of InterSignalingPresentFlag is set to 1.
[0219] If the picture associated with the picture header has one or more slices of the slice type I slice, the value of IntraSignalingPresentFlag is set to 1. If the picture associated with the picture header has one or more slices of the slice type P or B slice, the value of InterSignalingPresentFlag is set to 1.
[0220] For example, if only intra-coded slices are included in the current picture, the encoding device may determine the value of mixed_slice_signaling_present_flag to 0, may determine the value of intra_slice_only_flag to 1, may determine the value of IntraSignalingPresentFlag to 1, and may determine the value of InterSignalingPresentFlag to 0.
[0221] As another example, if only inter-coded slices are included in the current picture, the encoding device may determine the value of mixed_slice_signaling_present_flag to 0, may determine the value of intra_slice_only_flag to 0, may determine the value of IntraSignalingPresentFlag to 0, and may determine the value of InterSignalingPresentFlag to 1.
[0222] As yet another example, if at least one intra-coded slice and at least one inter-coded slice are included in the current picture, the encoding apparatus may determine the values of mixed_slice_signaling_present_flag, IntraSignalingPresentFlag, and InterSignalingPresentFlag to be 1, respectively.
[0223] If the value of mixed_slice_signaling_present_flag obtained from the picture header in the image information is 0, the decoding device can determine whether the corresponding picture includes only intra-frame coded slices or inter-frame coded slices. In this case, if the value of intra_slice_only_flag obtained from the picture header is 0, the decoding device can parse the syntax elements necessary for decoding only inter-frame coded slices from the picture header. If the value of intra_slice_only_flag is 1, the decoding device can parse the syntax elements necessary for decoding only intra-frame coded slices from the picture header.
[0224] If the value of mixed_slice_signaling_present_flag obtained from the picture header in the image information is 1, the decoding device can determine that at least one intra-coded slice and at least one inter-coded slice are included in the corresponding picture, and can parse the syntax elements required for decoding the inter-coded slice and the syntax elements required for decoding the intra-coded slice from the picture header.
[0225] Figure 9 and Figure 10 Schematically represents an example of a video / image encoding method and related components according to an embodiment of this document.
[0226] Figure 9 The video / image encoding method disclosed in Figure 2 and Figure 10 Specifically, for example, Figure 9 S900 and S910 may be performed by the predictor 220 of the encoding apparatus 200 ; S920 may be performed by the adder 250 of the encoding apparatus 200 ; and S930 and S940 may be performed by the entropy encoder 240 of the encoding apparatus 200 . Figure 9 The video / image encoding method disclosed in may include the embodiments described above in this document.
[0227] Specifically, refer to Figure 9 and Figure 10, the predictor 220 of the encoding device may determine a prediction mode for a current block in a current picture (S900). The current picture may include multiple slices. The predictor 220 of the encoding device may generate a prediction sample (prediction block) for the current block based on the prediction mode (S910). Here, the prediction mode may include an inter-frame prediction mode and an intra-frame prediction mode. When the prediction mode of the current block is an inter-frame prediction mode, the prediction sample may be generated by the inter-frame predictor 221 of the predictor 220. When the prediction mode of the current block is an intra-frame prediction mode, the prediction sample may be generated by the intra-frame predictor 222 of the predictor 220.
[0228] The residual processor 230 of the encoding device can generate residual samples and residual information based on the predicted samples and the original picture (original block, original sample). Here, the residual information is information about the residual samples and may include information about (quantized) transform coefficients for the residual samples.
[0229] The adder (or reconstructor) of the encoding device can generate reconstructed samples (reconstructed pictures, reconstructed blocks, reconstructed sample arrays) by adding the residual samples generated by the residual processor 230 and the prediction samples generated by the inter-frame predictor 221 or the intra-frame predictor 222 (S920).
[0230] At the same time, the entropy encoder 240 of the encoding device may generate at least one of first information indicating whether information necessary for an inter-frame prediction operation in a decoding process is present in a picture header associated with the current picture, or second information indicating whether information necessary for an intra-frame prediction operation in a decoding process is present in a picture header associated with the current picture, based on the prediction mode (S930). Here, the first information and the second information are information included in the picture header of the image information, and may correspond to the aforementioned intra_signalling_present_flag, inter_signalling_present_flag, slice_type_idc, mixed_slice_signalling_present_flag, intra_slice_only_flag, IntraSignallingPresentFlag, and / or InterSignallingFlag.
[0231] For example, if the picture header associated with the current picture includes information necessary for an inter-frame prediction operation during decoding due to the inclusion of an inter-frame coded slice in the current picture, the entropy encoder 240 of the encoding device may determine the value of the first information to be 1. Furthermore, if the picture header associated with the current picture includes information necessary for an intra-frame prediction operation during decoding due to the inclusion of an intra-frame coded slice in the current picture, the entropy encoder 240 of the encoding device may determine the value of the second information to be 1. In this case, the first information may correspond to inter_signaling_present_flag, and the second information may correspond to intra_signaling_present_flag. The first information may be referred to as a first flag, information regarding the presence of syntax elements for inter-frame slices in the picture header, a flag indicating whether syntax elements for inter-frame slices are present in the picture header, information regarding whether a slice in the current picture is an inter-frame slice, or a flag indicating whether a slice is an inter-frame slice. The second information may be referred to as a second flag, information about whether a syntax element for intra-frame slicing exists in a picture header, a flag for whether a syntax element for intra-frame slicing exists in a picture header, information about whether a slice in a current picture is an intra-frame slice, or a flag for whether a slice is an intra-frame slice.
[0232] Meanwhile, if a picture includes only intra-coded slices and the corresponding picture header includes information necessary only for intra-frame prediction operations, the entropy encoder 240 of the encoding device may determine the value of the first information to be 0 and the value of the second information to be 1. Furthermore, if a picture includes only inter-coded slices and the corresponding picture header includes information necessary only for inter-frame prediction operations, the value of the first information may be 1 and the value of the second information may be 0. Accordingly, if the value of the first information is 0, all slices in the current picture may have an I-slice type. If the value of the second information is 0, all slices in the current picture may have a P-slice type or a B-slice type. Here, the information necessary for intra-frame prediction operations may include syntax elements for decoding intra-frame slices, while the information necessary for inter-frame prediction operations may include syntax elements for decoding inter-frame slices.
[0233] As another example, if all slices in the current picture have an I slice type, the entropy encoder 240 of the encoding device may determine the value of the information on the slice type to be 0, and if all slices in the current picture have a P slice type or a B slice type, the entropy encoder 240 of the encoding device may determine the value of the information on the slice type to be 1. If all slices in the current picture have an I slice type, a P slice type, and / or a B slice type (i.e., the slice types of the slices in the picture are mixed), the entropy encoder 240 of the encoding device may determine the value of the information on the slice type to be 2. In this case, the information on the slice type may correspond to slice_type_idc.
[0234] As yet another example, if all slices in the current picture have the same slice type, the entropy encoder 240 of the encoding device may determine the value of the information about the slice type to be 0, and if the slices in the current picture have different slice types, the entropy encoder 240 of the encoding device may determine the value of the information about the slice type to be 1. In this case, the information about the slice type may correspond to mixed_slice_signaling_present_flag.
[0235] If the value of the information on the slice type is determined to be 0, information on whether an intra slice is included in the slice may be included in the corresponding picture header. The information on whether an intra slice is included in the slice may correspond to intra_slice_only_flag. If all slices in the picture have an I slice type, the entropy encoder 240 of the encoding device may determine the value of the information on whether an intra slice is included in the slice to be 1, determine the value of the information on whether a syntax element for an intra slice is present in the picture header to be 1, and determine the value of the information on whether a syntax element for an inter slice is present in the picture header to be 0. If the slice type of all slices in the picture is a P slice type and / or a B slice type, the entropy encoder 240 of the encoding device may determine the value of the information on whether an intra slice is included in the slice to be 0, determine the value of the information on whether a syntax element for an intra slice is present in the picture header to be 0, and determine the value of the information on whether a syntax element for an inter slice is present in the picture header to be 1.
[0236] The entropy encoder 240 of the encoding device can encode the image information including the above-mentioned first information, second information and slice type information, as well as residual information and prediction-related information (S940). For example, the image information may include partition-related information, information about prediction mode, residual information, in-loop filtering-related information, first information, second information, slice type information, etc., and include various syntax elements related to them. In an example, the image information may include information about various parameter sets such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), etc. In addition, the image information may include various information such as picture header syntax, picture header structure syntax, slice header syntax, coding unit syntax, etc. The above-mentioned first information, second information, information about slice type, information required for intra-frame prediction operation, and information required for inter-frame prediction operation may be included in the syntax in the picture header.
[0237] The information encoded by the entropy encoder 240 of the encoding device may be output in the form of a bitstream. The bitstream may be transmitted to the decoding device through a network or a storage medium.
[0238] Figure 11 and Figure 12 Schematically represents an example of a video / image decoding method and related components according to an embodiment of this document.
[0239] Figure 11 The video / image decoding method disclosed in Figure 3 and Figure 12 Specifically, for example, the entropy decoder 310 of the decoding device may be used to perform the decoding. Figure 11 S1100 and S1120 of the decoding apparatus 300 may be performed, and S1130 may be performed in the predictor 330 of the decoding apparatus 300. S1130 and S1140 may be performed by the adder 340 of the decoding apparatus 300. Figure 11 The video / image decoding method disclosed in the disclosure may include the embodiments described above in this document.
[0240] refer to Figure 11 and Figure 12 The entropy decoder 310 of the decoding device may obtain image information from the bitstream (S1100). The image information may include a picture header associated with the current picture. The current picture may include multiple slices.
[0241] At the same time, the entropy decoder 310 of the decoding device can parse at least one of a first flag indicating whether information required for an inter-frame prediction operation of the decoding process is present in the picture header associated with the current picture, or a second flag indicating whether information required for an intra-frame prediction operation of the decoding process is present in the picture header associated with the current picture (S1110). Here, the first flag and the second flag may correspond to the aforementioned intra_signalling_present_flag, inter_signalling_present_flag, slice_type_idc, mixed_slice_signalling_present_flag, intra_slice_only_flag, IntraSignallingPresentFlag, and / or InterSignallingPresentFlag. The entropy decoder 310 of the decoding device can parse the syntax elements included in the picture header of the image information based on the picture header syntax of any one of Tables 2, 4, and 6 above.
[0242] The decoding apparatus may generate a prediction sample by performing at least one of intra prediction and inter prediction on a current block in a current picture based on the first flag, the second flag, and the information about the slice type ( S1120 ).
[0243] Specifically, the entropy decoder 310 of the decoding device can parse (or obtain) at least one of the information required for the intra-frame prediction operation and / or the information required for the inter-frame prediction operation for the decoding process from the picture header related to the current picture based on the first flag, the second flag and / or the information about the slice type. The predictor 330 of the decoding device can generate a prediction sample by performing intra-frame prediction and / or inter-frame prediction based on at least one of the information required for the intra-frame prediction operation or the information for inter-frame prediction. Here, the information required for the intra-frame prediction operation may include syntax elements for decoding of intra-frame slices, and the information required for the inter-frame prediction operation may include syntax elements for decoding of inter-frame slices.
[0244] As an example, if the value of the first flag is 0, the entropy decoder 310 of the decoding device can determine (or decide) that the syntax element for inter prediction is not present in the picture header, and can parse only the information necessary for the intra prediction operation from the picture header. If the value of the first flag is 1, the entropy decoder 310 of the decoding device can determine (or decide) that the syntax element for inter prediction is present in the picture header, and can parse the information necessary for the inter prediction operation from the picture header. In this case, the first flag may correspond to inter_signaling_present_flag.
[0245] In addition, if the value of the second flag is 0, the entropy decoder 310 of the decoding device can determine (or decide) that the syntax elements for intra prediction are not present in the picture header, and can parse only the information necessary for inter prediction operation from the picture header. If the value of the second flag is 1, the entropy decoder 310 of the decoding device can determine (or decide) that the syntax elements for intra prediction are present in the picture header, and can parse only the information necessary for intra prediction operation from the picture header. In this case, the second flag may correspond to intra_signaling_present_flag.
[0246] If the value of the first flag is 0, the decoding device may determine that all slices in the current picture are of the I-slice type. If the value of the first flag is 1, the decoding device may determine that 0 or more slices in the current picture are of the P-slice type or the B-slice type. In other words, if the value of the first flag is 1, a slice of the P-slice type or the B-slice type may be included in the current picture, or a slice of the P-slice type or the B-slice type may not be included in the current picture.
[0247] Furthermore, if the value of the second flag is 0, the decoding device may determine that all slices in the current picture are of the P slice or B slice type. If the value of the second flag is 1, the decoding device may determine that 0 or more slices in the current picture are of the I slice type. In other words, if the value of the second flag is 1, a slice of the I slice type may be included in the current picture, or a slice of the I slice type may not be included in the current picture.
[0248] As another example, if the value of the information on the slice type is 0, the entropy decoder 310 of the decoding device can determine that all slices in the current picture have an I slice type and can parse only the information required for intra-frame prediction operation. If the information on the slice type is 1, the entropy decoder 310 of the decoding device can determine that all slices in the corresponding picture have a P slice type or a B slice type and can parse only the information required for inter-frame prediction operation from the picture header. If the value of the information on the slice type is 2, the entropy decoder 310 of the decoding device can determine that the slice in the corresponding picture has a slice type in which I slice type, P slice type and / or B slice type are mixed, and can parse all of the information required for inter-frame prediction operation and the information required for intra-frame prediction operation from the picture header. In this case, the information on the slice type can correspond to slice_type_idc.
[0249] As yet another example, the entropy decoder 310 of the decoding apparatus may determine that all slices in the current picture have the same slice type when the value of the information about the slice type is determined to be 0, and may determine that the slices in the current picture have different slice types when the value of the information about the slice type is determined to be 1. In this case, the information about the slice type may correspond to mixed_slice_signalling_present_flag.
[0250] If the value of the information on the slice type is determined to be 0, the entropy decoder 310 of the decoding device can parse information on whether an intra slice is included in the slice from the picture header. The information on whether an intra slice is included in the slice may correspond to the intra_slice_only_flag described above. If the information on whether an intra slice is included in the slice is 1, all slices in the picture may have an I slice type.
[0251] If the value of the information on whether an intra slice is included in the slice is 1, the entropy decoder 310 of the encoding device can parse only information necessary for intra prediction operation from the picture header. If the value of the information on whether an intra slice is included in the slice is 0, the entropy decoder 310 of the decoding device can parse only information necessary for inter prediction operation from the picture header.
[0252] If the value of the information about the slice type is 1, the entropy decoder 310 of the decoding apparatus may parse all of information required for an inter prediction operation and information necessary for an intra prediction operation from a picture header.
[0253] Meanwhile, the residual processor 320 of the decoding apparatus may generate residual samples based on the residual information obtained by the entropy decoder 310 .
[0254] The adder 340 of the decoding apparatus may generate a reconstructed sample (S1130) based on the predicted sample generated by the predictor 330 and the residual sample generated by the residual processor 320. In addition, the adder 340 of the decoding apparatus may generate a reconstructed picture (reconstructed block) based on the reconstructed sample (S1140).
[0255] After that, an in-loop filtering process such as an ALF process, SAO and / or deblocking filtering may be applied to the reconstructed picture as needed in order to improve the subjective / objective video quality.
[0256] Although the methods have been described in the above embodiments based on flowcharts in which steps or blocks are listed in sequence, the steps of the present disclosure are not limited to a specific order, and a step may be performed in a different step, in a different order, or simultaneously relative to the above order. In addition, it should be understood by those skilled in the art that the steps in the flowchart are not exclusive, and another step may be included therein or one or more steps in the flowchart may be deleted without affecting the scope of the present disclosure.
[0257] The above-mentioned method according to the present disclosure may be in the form of software, and the encoding device and / or decoding device according to the present disclosure may be included in an apparatus for performing image processing (e.g., TV, computer, smart phone, set-top box, display device, etc.).
[0258] When the embodiments of the present disclosure are implemented with software, the above-mentioned methods can be implemented with modules (processing or functions) that perform the above-mentioned functions. The modules can be stored in a memory and executed by a processor. The memory can be installed inside or outside the processor and can be connected to the processor via various well-known devices. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits and / or data processing devices. The memory may include a read-only memory (ROM), a random access memory (RAM), a flash memory, a memory card, a storage medium and / or other storage devices. In other words, according to the embodiments of the present disclosure, it can be implemented and executed on a processor, a microprocessor, a controller or a chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, a processor, a microprocessor, a controller or a chip. In this case, information about the implementation (for example, information about instructions) or the algorithm can be stored in a digital storage medium.
[0259] In addition, the decoding device and encoding device of the embodiment of the present disclosure can be included in a multimedia broadcast transceiver, a mobile communication terminal, a home theater video device, a digital theater video device, a surveillance camera, a video chat device, a real-time communication device such as video communication, a mobile streaming device, a storage medium, a portable camera, a video on demand (VoD) service provider, an over-the-top (OTT) video device, an Internet streaming service provider, a 3D video device, a virtual reality (VR) device, an augmented reality (AR) device, an image phone video device, a vehicle terminal (e.g., a vehicle (including an autonomous vehicle) terminal, an aircraft terminal, or a ship terminal) and a medical video device; and can be used to process image signals or data. For example, OTT video devices may include game consoles, Blueray players, networked TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0260] In addition, the processing method of the embodiment of the present disclosure can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the embodiment of the present disclosure can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all kinds of storage devices and distributed storage devices storing computer-readable data. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. The computer-readable recording medium also includes a medium implemented in the form of a carrier wave (for example, transmission on the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium, or can be transmitted through a wired or wireless communication network.
[0261] In addition, the embodiments of the present disclosure can be implemented as a computer program product based on a program code, and the program code can be executed on a computer according to the embodiments of this document.The program code can be stored on a computer readable carrier.
[0262] Figure 13 1 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0263] refer to Figure 13 A content streaming system to which the embodiments of the present disclosure are applied may generally include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0264] The encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generate a bitstream, and send it to the streaming server. As another example, if the multimedia input device such as smartphones, cameras, and camcorders directly generates the bitstream, the encoding server can be omitted.
[0265] The bitstream may be generated by the encoding method or the bitstream generation method to which the embodiment of the present disclosure is applied, and the streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0266] The streaming server sends multimedia data to the user device via a network server based on the user's request. The network server serves as a tool for notifying the user of available services. When the user requests a desired service, the network server transfers the request to the streaming server, which then transmits the multimedia data to the user. In this regard, the content streaming system may include a separate control server, in which case the control server is used to control commands and responses between the various devices in the content streaming system.
[0267] The streaming server may receive content from a media storage device and / or an encoding server. For example, when receiving content from an encoding server, the content may be received in real time. In this case, the streaming server may store the bitstream for a predetermined period of time to smoothly provide a streaming service.
[0268] For example, user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation, tablet PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smart watches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0269] Each server in the content streaming system may be operated as a distributed server, and in such case, data received by each server may be processed in a distributed manner.
Claims
1. A video decoding method performed by a video decoding device, the method comprising: Receiving a bitstream including image information, wherein the image information includes a picture header associated with a current picture, and a plurality of slices are included in the current picture; Parsing from the picture header a syntax element specifying whether all slices in the current picture are I slices, or whether slices in the current picture are P slices or B slices; determining, based on a value of the syntax element, whether the picture header includes a syntax element related to the I slice, or includes a syntax element related to the P slice or the B slice; deriving prediction samples by performing at least one of intra prediction or inter prediction on a block in the slice in the current picture based on the syntax elements; and Based on the derived prediction samples, reconstructed samples are generated, wherein, based on the syntax element specifying that the slice in the current picture is the P slice or the B slice, the picture header includes syntax elements related to the P slice or the B slice, wherein the syntax elements related to the P slice or the B slice include syntax elements indicating a maximum hierarchical depth of coding units generated by multi-type tree splitting in an inter slice in the current picture, and Particularly, based on the syntax element specifying that all the slices in the current picture are I slices, the picture header includes syntax elements related to the I slices, wherein the syntax elements related to the I slices include syntax elements indicating a maximum hierarchical depth of coding units generated by multiple types of tree splitting in intra slices in the current picture.
2. A video encoding method performed by a video encoding device, the method comprising: determining a type of a slice in a current picture, wherein the slice is included in the current picture; generating a syntax element that specifies whether all slices in the current picture are I slices, or whether slices in the current picture are P slices or B slices; determining, based on a value of the syntax element, whether the picture header includes a syntax element related to the I slice, or includes a syntax element related to the P slice or the B slice; Based on the determination, generating a syntax element associated with the I slice, or a syntax element associated with the P slice or the B slice; and encoding picture information, wherein the picture information includes at least one of the syntax element specifying whether all the slices in the current picture are I slices, or whether the slices in the current picture are P slices or B slices, a syntax element related to the I slice, or a syntax element related to the P slice or the B slice; wherein, based on the syntax element specifying that the slice in the current picture is the P slice or the B slice, the picture header includes syntax elements related to the P slice or the B slice, wherein the syntax elements related to the P slice or the B slice include syntax elements indicating a maximum hierarchical depth of coding units generated by multi-type tree splitting in an inter slice in the current picture, and Particularly, based on the syntax element specifying that all the slices in the current picture are I slices, the picture header includes syntax elements related to the I slices, wherein the syntax elements related to the I slices include syntax elements indicating a maximum hierarchical depth of coding units generated by multiple types of tree splitting in intra slices in the current picture.
3. A method for transmitting video data, the method comprising: Obtaining a bitstream of the video, wherein the bitstream is generated based on the following steps: determining a type of slice in a current picture, wherein the slice is included in the current picture, generating a syntax element specifying whether all slices in the current picture are I slices, or whether slices in the current picture are P slices or B slices, determining based on a value of the syntax element whether the picture header includes a syntax element related to the I slice, or includes a syntax element related to the P slice or the B slice, generating a syntax element related to the I slice, or a syntax element related to the P slice or the B slice based on the determination, and encoding image information, wherein the image information includes at least one of the syntax element specifying whether all slices in the current picture are I slices, or whether slices in the current picture are P slices or B slices, the syntax element related to the I slice, or the syntax element related to the P slice or the B slice; and sending said data comprising said bitstream, wherein, based on the syntax element specifying that the slice in the current picture is the P slice or the B slice, the picture header includes syntax elements related to the P slice or the B slice, wherein the syntax elements related to the P slice or the B slice include syntax elements indicating a maximum hierarchical depth of coding units generated by multi-type tree splitting in an inter slice in the current picture, and Particularly, based on the syntax element specifying that all the slices in the current picture are I slices, the picture header includes syntax elements related to the I slices, wherein the syntax elements related to the I slices include syntax elements indicating a maximum hierarchical depth of coding units generated by multiple types of tree splitting in intra slices in the current picture.
Citation Information
Patent Citations
Image-decoding method and apparatus including a method for configuring a reference picture list
WO2012033327A2