Image decoding method and apparatus for encoding image information including picture header
By obtaining the logo of the picture head (PH) network abstraction layer (NAL) unit in the image decoding device, adaptively control the bit rate of the NAL unit, solving the problem of high-resolution and high-quality images with high transmission and storage costs, and improving image encoding efficiency is achieved.
Patent Information
- Application Number
- CN202080097941.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-02
- Filing Date
- 2020-12-29
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2040-12-29
AI Technical Summary
The transmission and storage cost of high-resolution high-quality images is high, and image encoding efficiency needs to be improved to reduce bit volume.
By obtaining the flag of the picture head (PH) network abstraction layer (NAL) unit in the image decoding device, the bit rate of the NAL unit is adaptively controlled to optimize the encoding efficiency, including determining the presence of the PH NAL unit in the encoding device and generating corresponding flags for encoding.
By adaptively controlling the bit rate of the NAL unit, the overall efficiency of image encoding is improved and transmission and storage costs are reduced.
Smart Images

Figure CN115211122B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to image coding technology, and more particularly, to a video decoding method and device for adaptively encoding PH NAL units in an image coding system. Background Art
[0002] Recently, demand for high-resolution, high-quality images, such as HD (High Definition) and UHD (Ultra High Definition), has been growing in various fields. Because image data has high resolution and high quality, the amount of information or bits to be transmitted has increased compared to conventional image data. Consequently, when image data is transmitted using media such as conventional wired / wireless broadband lines or stored using existing storage media, transmission and storage costs increase.
[0003] Therefore, there is a need for efficient image compression technology for effectively transmitting, storing, and reproducing information of high-resolution and high-quality images. Summary of the Invention
[0004] Technical issues
[0005] The technical purpose of the present disclosure is to provide a method and device for improving image coding efficiency.
[0006] Another technical objective of the present disclosure is to provide a method and apparatus for encoding a flag indicating whether a PH NAL unit exists.
[0007] Technical Solution
[0008] According to an embodiment of the present disclosure, an image decoding method performed by a decoding device is provided. The method includes the following steps: obtaining a flag indicating whether a picture header (PH) network abstraction layer (NAL) unit exists; obtaining the PH based on the flag; and decoding a current picture related to the PH based on the PH.
[0009] According to another embodiment of the present disclosure, a decoding device for performing image decoding is provided. The decoding device includes: an entropy decoder configured to obtain a flag indicating whether a picture header (PH) network abstraction layer (NAL) unit exists, and obtain the PH based on the flag; and a predictor configured to decode a current picture related to the PH based on the PH.
[0010] According to another embodiment of the present disclosure, a video encoding method performed by an encoding device is provided. The method includes the following steps: determining whether a picture header (PH) network abstraction layer (NAL) unit including a PH header related to a current picture exists; generating a flag indicating whether the PH NAL unit exists based on the determination result; and encoding image information including the flag.
[0011] According to another embodiment of the present disclosure, a video encoding device is provided. The encoding device includes an entropy encoder configured to determine whether a picture header (PH) network abstraction layer (NAL) unit including a PH related to a current picture exists, generate a flag indicating whether the PH NAL unit exists based on a result of the determination, and encode image information including the flag.
[0012] According to another embodiment of the present disclosure, a computer-readable digital storage medium is provided that stores a bitstream including image information and enables execution of an image decoding method. The computer-readable digital storage medium includes the following steps: obtaining a flag indicating whether a picture header (PH) network abstraction layer (NAL) unit exists; obtaining the PH based on the flag; and decoding a current picture related to the PH based on the PH.
[0013] Beneficial effects
[0014] According to the present disclosure, a flag indicating whether a PH NAL unit exists may be signaled, the NAL unit may be controlled adaptively to the bit rate of a bitstream based on the flag, and overall coding efficiency may be improved.
[0015] According to the present disclosure, based on a flag indicating whether a PH NAL unit exists, constraints on the number of slices in the current picture and on the presence of a PH NAL unit can be set for the relevant picture to control the NAL unit adaptively according to the bit rate of the bitstream, thereby improving overall coding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 An example of a video / image encoding device to which the embodiments of the present disclosure can be applied is briefly illustrated.
[0017] Figure 2 is a schematic diagram illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied.
[0018] Figure 3 FIG. 1 is a schematic diagram illustrating a configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.
[0019] Figure 4 The hierarchical structure of the encoded image information is schematically shown.
[0020] Figure 5 The encoding process according to an embodiment of the present disclosure is schematically illustrated.
[0021] Figure 6 The decoding process according to an embodiment of the present disclosure is schematically illustrated.
[0022] Figure 7 The picture header configuration in the NAL unit according to whether the PH NAL unit exists is schematically shown.
[0023] Figure 8 The image encoding method of the encoding device according to this document is schematically shown.
[0024] Figure 9 A coding device for performing the image coding method according to the present document is schematically shown.
[0025] Figure 10 The image decoding method of the decoding device according to this document is schematically shown.
[0026] Figure 11 A decoding device for performing the image decoding method according to the present document is schematically shown.
[0027] Figure 12 The structure diagram of the content streaming system to which the present disclosure is applied is illustrated. DETAILED DESCRIPTION
[0028] The present disclosure can be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, the embodiments are not intended to limit the present disclosure. The terms used in the following description are only used to describe specific embodiments and are not intended to limit the present disclosure. As long as it is clearly understood in different ways, singular expressions include plural expressions. Terms such as "including" and "having" are intended to indicate the presence of features, quantities, steps, operations, elements, components, or combinations thereof used in the following description, so it should be understood that the possibility of the presence or addition of one or more different features, quantities, steps, operations, elements, components, or combinations thereof is not excluded.
[0029] Furthermore, the elements in the drawings described in this disclosure are drawn independently for the purpose of explaining different specific functions, and do not imply that these elements are specifically implemented by independent hardware or independent software. For example, two or more elements in the drawings may be combined to form a single element, or an element may be divided into multiple elements. Implementations of combining and / or dividing elements belong to this disclosure and do not depart from the concepts of this disclosure.
[0030] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. In addition, throughout the drawings, like reference numerals are used to indicate like elements, and the same description of like elements will be omitted.
[0031] Figure 1 An example of a video / image encoding device to which the embodiments of the present disclosure can be applied is briefly illustrated.
[0032] Reference Figure 1The video / image coding system may include a first device (source device) and a second device (receiver device). The source device may transmit coded video / image information or data to the receive device in the form of a file or stream via a digital storage medium or a network.
[0033] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0034] The video source can obtain the video / image by capturing, synthesizing, or generating a video / image. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured videos / images, etc. The video / image generation device may include, for example, a computer, a tablet computer, and a smartphone, and may (electronically) generate the video / image. For example, a virtual video / image may be generated by a computer, etc. In this case, the video / image capture process may be replaced by a process that generates relevant data.
[0035] The encoding device can encode the input video / image. The encoding device can perform a series of processes such as prediction, transformation, and quantization to achieve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0036] The transmitter can transmit encoded video / image information or data in the form of a bitstream to a receiver of a receiving device via a digital storage medium or network in the form of a file or stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include components for generating media files in a predetermined file format and can also include components for transmitting via a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received bitstream to a decoding device.
[0037] The decoding device may decode a video / image by performing a series of processes such as dequantization, inverse transformation, and prediction corresponding to the operations of the encoding device.
[0038] The renderer may render the decoded video / image, and the rendered video / image may be displayed on a display.
[0039] The present disclosure relates to video / image coding. For example, the methods / implementations disclosed in the present disclosure can be applied to methods disclosed in Versatile Video Coding (VVC), EVC (Essential Video Coding) standards, AOMedia Video 1 (AV1) standards, the second generation Audio Video Coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0040] The present disclosure presents various embodiments of video / image encoding, and unless otherwise mentioned, the embodiments may be performed in combination with each other.
[0041] In the present disclosure, video may refer to a series of images over time. Generally, a picture refers to a unit representing an image in a specific time region, and a sub-picture / slice / tile is a unit that constitutes a part of a picture in encoding. A sub-picture / slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more sub-pictures / slices / tiles. A picture may be composed of one or more tile groups. A tile group may include one or more tiles. A tile may represent a rectangular area of a CTU row within a tile in a picture. A tile may be partitioned into multiple tiles, each tile consisting of one or more CTU rows within the tile. A tile that is not partitioned into multiple tiles may also be referred to as a tile. Tile scanning is a specific sequential ordering of CTUs that partition a picture, wherein CTUs are continuously ordered in a tile by a CTU raster scan, tiles within a tile are continuously ordered by a raster scan of the tiles of the tile, and tiles in the picture are continuously ordered by a raster scan of the tiles of the picture. In addition, a sub-picture can represent a rectangular area of one or more slices within a picture. That is, a sub-picture contains one or more slices that together cover a rectangular area of the picture. A tile is a rectangular area of a CTU within a specific tile column and a specific tile row in a picture. A tile column is a rectangular area of a CTU whose height is equal to the height of the picture and whose width is specified by a syntax element in the picture parameter set. A tile row is a rectangular area of a CTU whose height is specified by a syntax element in the picture parameter set and whose width is equal to the width of the picture. Tile scan is a specific sequential ordering of CTUs that partition a picture, where CTUs can be ordered continuously within a tile according to a CTU raster scan, while tiles in a picture can be ordered continuously according to a raster scan of the tiles of the picture. A slice includes an integer number of tiles of a picture that can be exclusively contained in a single NAL unit. A slice can consist of multiple complete tiles or only a continuous sequence of complete tiles of a tile. In this disclosure, tile group and slice can be used interchangeably. For example, in this disclosure, a tile / tile header may be referred to as a slice / slice header.
[0042] A pixel or picture element (pel) may represent the smallest unit constituting a picture (or image). In addition, "sample" may be used as a term corresponding to a pixel. A sample may generally represent a pixel or a pixel value, and may represent only a pixel / pixel value of a luminance component or only a pixel / pixel value of a chrominance component.
[0043] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to the area. A unit may include a luminance block and two chrominance (e.g., CB, CR) blocks. In some cases, a unit may be used interchangeably with terms such as block or area. In general, an M×N block may include M columns and N rows of samples (or sample arrays) or a set (or array) of transform coefficients.
[0044] In this specification, "A or B" may mean "only A," "only B," or "A and B." In other words, in this specification, "A or B" may be interpreted as "A and / or B." For example, "A, B, or C" herein means "only A," "only B," "only C," or "any one and any combination of A, B, and C."
[0045] As used herein, a slash ( / ) or a comma may mean "and / or." For example, "A / B" may mean "A and / or B." Thus, "A / B" may mean "only A," "only B," or "A and B." For example, "A,B,C" may mean "A, B, or C."
[0046] In this specification, “at least one of A and B” may mean “only A”, “only B”, or “both A and B”. In addition, in this specification, the expression “at least one of A or B” or “at least one of A and / or B” may be interpreted as the same as “at least one of A and B”.
[0047] In addition, in this specification, "at least one of A, B, and C" means "only A", "only B", "only C", or "any combination of A, B, and C". In addition, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0048] Furthermore, brackets used in this specification may refer to "for example." Specifically, when "prediction (intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-frame prediction," and "intra-frame prediction" may be provided as an example of "prediction." Furthermore, even when "prediction (i.e., intra-frame prediction)" is indicated, "intra-frame prediction" may be provided as an example of "prediction."
[0049] In this specification, technical features described separately in one drawing may be implemented separately or simultaneously.
[0050] The following figures are created to explain specific examples of this specification. Since the names of specific devices or the names of specific signals / messages / fields described in the figures are presented by way of example, the technical features of this specification are not limited to the specific names used in the following figures.
[0051] Figure 2 1 is a schematic diagram illustrating a configuration of a video / image encoding device to which an embodiment of the present disclosure can be applied. Hereinafter, a video encoding device may include an image encoding device.
[0052] Reference Figure 2 , the encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230 and an entropy encoder 240, an adder 250, a filter 260 and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234 and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstructed block generator. According to an embodiment, the image segmenter 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 250 and the filter 260 may be composed of at least one hardware component (e.g., an encoder chipset or processor). In addition, the memory 270 may include a decoded picture buffer (DPB) or may be composed of a digital storage medium. The hardware components may also include the memory 270 as an internal / external component.
[0053] The image splitter 210 can split the input image (or picture or frame) input to the encoding device 200 into one or more processors. For example, the processor can be referred to as a coding unit (CU). In this case, the coding unit can be recursively split from the coding tree unit (CTU) or the largest coding unit (LCU) according to the quadtree binary tree ternary tree (QTBTTT) structure. For example, a coding unit can be split into multiple coding units with a deeper depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quadtree structure can be applied first, and then the binary tree structure and / or the ternary structure can be applied. Alternatively, the binary tree structure can be applied first. The encoding process according to the present disclosure can be performed based on the final coding unit that is no longer split. In this case, the maximum coding unit can be used as the final coding unit based on coding efficiency according to image characteristics, or if necessary, the coding unit can be recursively split into coding units with a deeper depth and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processor may further include a prediction unit or a transform unit. In this case, the prediction unit and the transform unit may be separated or partitioned from the final coding unit. The prediction unit may be a unit for sample prediction, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving a residual signal from the transform coefficients.
[0054] In some cases, the term "unit" can be used interchangeably with terms such as "block" or "region." In general, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or pixel value, and can represent only the pixel / pixel value of the luma component or only the pixel / pixel value of the chroma component. A sample can be used as a term corresponding to a picture (or image) of pixels or picture elements.
[0055] In the encoding device 200, the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 is subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the transformer 232. In this case, as shown in the figure, the unit in the encoding device 200 for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as a subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction on a current block or CU basis. As described later in the description of each prediction mode, the predictor can generate various information related to the prediction (such as prediction mode information) and send the generated information to the entropy encoder 240. The information about the prediction can be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0056] The intra-frame predictor 222 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples can be located near the current block or can be far away from the current block. In intra-frame prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The non-directional mode can include, for example, a DC mode and a planar mode. Depending on the level of detail of the prediction direction, the directional mode can include, for example, 33 directional prediction modes or 65 directional prediction modes. However, this is only an example, and more or fewer directional prediction modes can be used depending on the settings. The intra-frame predictor 222 can determine the prediction mode applied to the current block by using the prediction mode applied to the adjacent block.
[0057] The inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. To reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatially neighboring blocks in the current picture and temporally neighboring blocks in the reference picture. The reference picture including the reference block and the reference picture including the temporally neighboring block can be the same or different. Temporally neighboring blocks can be referred to as collocated reference blocks, collocated CUs (colCUs), etc., and the reference picture including temporally neighboring blocks can be referred to as collocated pictures (colPics). For example, the inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference picture index for the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in the case of skip mode and merge mode, the inter-frame predictor 221 can use the motion information of the adjacent block as the motion information of the current block. In skip mode, unlike merge mode, it may not be possible to send a residual signal. In the case of motion vector prediction (MVP) mode, the motion vector of the adjacent block can be used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0058] The predictor 220 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply both intra prediction and inter prediction at the same time. This can be called inter-frame intra-frame combined prediction (CIIP). In addition, the predictor can predict the block based on the intra block copy (IBC) prediction mode or palette mode. The IBC prediction mode or palette mode can be used for content image / video encoding of games, etc., for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction because the reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample value within the picture can be signaled based on information about the palette table and the palette index.
[0059] The prediction signal generated by the predictor (including the inter-frame predictor 221 and / or the intra-frame predictor 222) can be used to generate a reconstruction signal or to generate a residual signal. The transformer 232 can generate a transform coefficient by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a Karhunen-Loève transform (KLT), a graph-based transform (GBT), or a conditional nonlinear transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to a transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process can be applied to square pixel blocks of the same size, or can be applied to blocks of variable size other than square.
[0060] The quantizer 233 can quantize the transform coefficients and send them to the entropy encoder 240. The entropy encoder 240 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be called residual information. The quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scanning order, and generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Information about the transform coefficients can be generated. The entropy encoder 240 can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 240 can encode information required for video / image reconstruction (e.g., syntax element values, etc.) in addition to the quantized transform coefficients, together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (Network Abstraction Layer). The video / image information may also include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. In the present disclosure, information and / or syntax elements sent / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded through the above-mentioned encoding process and included in the bitstream. The bitstream may be transmitted over a network or may be stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 and / or a storage unit (not shown) that stores the signal may be included as internal / external elements of the encoding device 200. Alternatively, the transmitter may be included in the entropy encoder 240.
[0061] The quantized transform coefficients output from the quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transformation to the quantized transform coefficients using the dequantizer 234 and the inverse transformer 235. The adder 250 adds the reconstructed residual signal to the prediction signal output from the inter-frame predictor 221 or the intra-frame predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array). If the block to be processed has no residual (such as when skip mode is applied), the prediction block can be used as a reconstructed block. The adder 250 can be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current picture, and can be used for inter-frame prediction of the next picture through filtering as described below.
[0062] Furthermore, during picture coding and / or reconstruction, luma mapping and chroma scaling (LMCS) may be applied.
[0063] The filter 260 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 260 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 270 (specifically, the DPB of the memory 270). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filter 260 can generate various information related to filtering and send the generated information to the entropy encoder 240, as described later in the description of various filtering methods. The information related to filtering can be encoded by the entropy encoder 240 and output in the form of a bit stream.
[0064] The modified reconstructed picture sent to the memory 270 may be used as a reference picture in the inter predictor 221. When inter prediction is applied by the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus may be avoided, and encoding efficiency may be improved.
[0065] The DPB of the memory 270 can store a modified reconstructed picture used as a reference picture in the inter-frame predictor 221. The memory 270 can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of a reconstructed block in the picture. The stored motion information can be sent to the inter-frame predictor 221 and used as motion information of spatially neighboring blocks or motion information of temporally neighboring blocks. The memory 270 can store reconstructed samples of the reconstructed blocks in the current picture and can transmit the reconstructed samples to the intra-frame predictor 222.
[0066] Figure 3 FIG. 1 is a schematic diagram illustrating a configuration of a video / image decoding device to which an embodiment of the present disclosure can be applied.
[0067] Reference Figure 3 , the decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to an embodiment, the entropy decoder 310, the residual processor 320, the predictor 330, the adder 340, and the filter 350 may be composed of hardware components (e.g., a decoder chipset or processor). In addition, the memory 360 may include a decoded picture buffer (DPB) or may be composed of a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0068] When a bit stream including video / image information is input, the decoding device 300 can be used with Figure 2 The image is reconstructed accordingly to the processing of the video / image information in the encoding device. For example, the decoding device 300 can derive the unit / block based on the block segmentation related information obtained from the bit stream. The decoding device 300 can perform decoding using a processor applied in the encoding device. Therefore, the decoding processor can be, for example, a coding unit, and the coding unit can be divided from the coding tree unit or the maximum coding unit according to the quadtree structure, the binary tree structure and / or the ternary tree structure. One or more transform units can be derived from the coding unit. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.
[0069] The decoding device 300 can receive the data in the form of a bit stream from Figure 2The signal output by the encoding device of the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information required for image reconstruction (or picture reconstruction) (for example, video / image information). The video / image information may also include information about various parameter sets such as an adaptive parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may also include general constraint information. The decoding device may also decode the picture based on the information about the parameter set and / or the general constraint information. The signaled / received information and / or syntax elements described later in this disclosure can be decoded through a decoding process and obtained from the bitstream. For example, the entropy decoder 310 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and outputs the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residual. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, use the decoding target syntax element information, the decoding information of the decoding target block, or the information of the symbol / bin decoded in the previous stage to determine the context model, and arithmetically decode the bin by predicting the probability of occurrence of the bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the decoded symbol / bin information for the context model of the next symbol / bin. The information related to prediction among the information decoded by the entropy decoder 310 can be provided to the predictor (inter-frame predictor 332 and intra-frame predictor 331), and the residual value (that is, quantized transform coefficient and related parameter information) on which entropy decoding is performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive a residual signal (residual block, residual sample, residual sample array). In addition, the information about filtering among the information decoded by the entropy decoder 310 can be provided to the filter 350. In addition, a receiver (not shown) for receiving a signal output from the encoding device may be further configured as an internal / external element of the decoding device 300, or the receiver may be a component of the entropy decoder 310. In addition, the decoding device according to the present disclosure may be referred to as a video / image / picture decoding device, and the decoding device may be classified into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0070] The dequantizer 321 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 may rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding device. The dequantizer 321 may dequantize the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0071] The inverse transformer 322 performs an inverse transform on the transform coefficients to obtain a residual signal (residual block, residual sample array).
[0072] The predictor may perform prediction on the current block and generate a prediction block including prediction samples of the current block. The predictor may determine whether to apply intra prediction or inter prediction to the current block based on the information on prediction output from the entropy decoder 310, and may determine a specific intra / inter prediction mode.
[0073] The predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra prediction or inter prediction to predict a block, but also apply intra prediction and inter prediction at the same time. This can be called inter-frame intra-frame combined prediction (CIIP). In addition, the predictor can predict the block based on the intra block copy (IBC) prediction mode or palette mode. The IBC prediction mode or palette mode can be used for content image / video encoding of games, etc., for example, screen content coding (SCC). IBC basically performs prediction in the current picture, but can be performed similarly to inter prediction because the reference block is derived in the current picture. That is, IBC can use at least one of the inter prediction techniques described in this disclosure. The palette mode can be regarded as an example of intra coding or intra prediction. When the palette mode is applied, the sample value within the picture can be signaled based on information about the palette table and the palette index.
[0074] The intra-frame predictor 331 can predict the current block by referencing samples in the current picture. Depending on the prediction mode, the referenced samples may be located near the current block or may be located far away from the current block. In intra-frame prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. The intra-frame predictor 331 may determine the prediction mode to be applied to the current block by using the prediction modes applied to the neighboring blocks.
[0075] The inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted in units of blocks, sub-blocks, or samples based on the correlation of motion information between neighboring blocks and the current block. The motion information can include a motion vector and a reference picture index. The motion information can also include information about the inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter-frame prediction, neighboring blocks can include spatial neighboring blocks in the current picture and temporal neighboring blocks in the reference picture. For example, the inter-frame predictor 332 can configure a motion information candidate list based on the neighboring blocks and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the prediction information can include information indicating the inter-frame prediction mode for the current block.
[0076] The adder 340 can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including the inter-frame predictor 332 and / or the intra-frame predictor 331). If the block to be processed has no residual (for example, when the skip mode is applied), the prediction block can be used as the reconstructed block.
[0077] The adder 340 may be referred to as a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra-frame prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter-frame prediction of the next picture.
[0078] In addition, luma mapping and chroma scaling (LMCS) can be applied during picture decoding.
[0079] The filter 350 can improve the subjective / objective image quality by applying filtering to the reconstructed signal. For example, the filter 350 can generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture, and store the modified reconstructed picture in the memory 360 (specifically, the DPB of the memory 360). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0080] The (modified) reconstructed picture stored in the DPB of the memory 360 can be used as a reference picture in the inter-frame predictor 332. The memory 360 can store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the reconstructed block in the picture. The stored motion information can be sent to the inter-frame predictor 332 to be used as the motion information of the spatially adjacent blocks or the motion information of the temporally adjacent blocks. The memory 360 can store the reconstructed samples of the reconstructed blocks in the current picture and can transmit the reconstructed samples to the intra-frame predictor 331.
[0081] In the present disclosure, the embodiments described in the filter 260, the inter-frame predictor 221, and the intra-frame predictor 222 of the encoding device 200 may be the same as the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300 or may be applied to correspond to the filter 350, the inter-frame predictor 332, and the intra-frame predictor 331 of the decoding device 300, respectively. The same contents can also be applied to the inter-frame predictor 332 and the intra-frame predictor 331.
[0082] In the present disclosure, at least one of quantization / inverse quantization and / or transform / inverse transform may be omitted. When quantization / inverse quantization is omitted, the quantized transform coefficient may be referred to as a transform coefficient. When transform / inverse transform is omitted, the transform coefficient may be referred to as a coefficient or a residual coefficient, or, for uniformity of expression, may still be referred to as a transform coefficient.
[0083] In the present disclosure, quantized transform coefficients and transform coefficients may be referred to as transform coefficients and scaled transform coefficients, respectively. In this case, residual information may include information about the transform coefficients, and information about the transform coefficients may be signaled via residual coding syntax. Transform coefficients may be derived based on the residual information (or information about the transform coefficients), and scaled transform coefficients may be derived by inversely transforming (scaling) the transform coefficients. Residual samples may be derived based on inversely transforming (transforming) the scaled transform coefficients. This may also be applied / expressed in other parts of the present disclosure.
[0084] Figure 4 Schematically shows the hierarchical structure of the encoded image information.
[0085] Figure 4 The video / image encoded according to the coding layer and structure of the present disclosure can be schematically shown. Figure 4 , the coded video / image can be divided into a video coding layer (VCL) that processes video / image and video / image decoding processing, a subsystem that sends and stores coded information, and a network abstraction layer (NAL) that exists between the VCL and the subsystem and is responsible for functions.
[0086] For example, in the VCL, VCL data including compressed image data (slice data) can be generated, or a picture parameter set (PPS), a sequence parameter set (SPS), a video parameter set (VPS), or a parameter set including supplemental enhancement information (SEI) messages additionally required for image decoding processing can be generated.
[0087] For example, in NAL, a NAL unit can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in the VCL. In this case, RBSP can refer to slice data, parameter sets, SEI messages, etc. generated in the VCL. The NAL unit header can include NAL unit type information specified according to the RBSP data included in the corresponding NAL unit.
[0088] For example, Figure 4 As shown, NAL units can be classified into VCL NAL units and non-VCL NAL units according to the RBSP generated in the VCL. A VCL NAL unit may refer to a NAL unit including information about an image (slice data), and a non-VCL NAL unit may refer to a NAL unit including information required for image decoding (parameter set or SEI message).
[0089] The header information may be attached to the VCL NAL unit and the non-VCL NAL unit according to the data standard of the subsystem, and the VCL NAL unit and the non-VCL NAL unit including the header information may be transmitted over the network. For example, the NAL unit may be converted into a data format of a predetermined standard such as the H.266 / VVC file format, the Real-time Transport Protocol (RTP), or the Transport Stream (TS) and transmitted via various networks.
[0090] In addition, as described above, the type of the NAL unit may be specified according to the RBSP data structure included in the NAL unit, and information on the NAL unit type may be stored in the NAL unit header and signaled.
[0091] For example, NAL units can be classified into VCL NAL unit types and non-VCL NAL unit types according to whether they include information about images (slice data). In addition, VCL NAL unit types can be classified according to the characteristics and types of pictures included in the VCL NAL unit, and non-VCL NAL unit types can be classified according to the types of parameter sets.
[0092] The following may be examples of NAL unit types specified according to the type of parameter sets included in a non-VCL NAL unit type.
[0093] - Adaptation Parameter Set (APS) NAL unit: includes the NAL unit type of the APS
[0094] -Decoding parameter set (DPS) NAL unit: includes the NAL unit type of the DPS
[0095] - Video Parameter Set (VPS) NAL unit: includes the NAL unit type of the VPS
[0096] - Sequence Parameter Set (SPS) NAL unit: NAL unit type that includes SPS
[0097] - Picture parameter set (PPS) NAL unit: NAL unit type including PPS
[0098] - Picture header (PH) NAL unit: includes the NAL unit type of the PH
[0099] The above-mentioned NAL unit type may have syntax information about the NAL unit type, which may be stored in the NAL unit header and signaled. For example, the syntax information may be nal_unit_type, and the NAL unit type may be specified as a nal_unit_type value.
[0100] In addition, as described above, a picture may include multiple slices, and a slice may include a slice header and slice data. In this case, a picture header may be added (embedded) for multiple slices (a collection of slice headers and slice data). The picture header (picture header syntax) may include information / parameters that are commonly applicable to the picture. The slice header (slice header syntax) may include information / parameters that are commonly applicable to the slices. The APS (APS syntax) or PPS (PPS syntax) may include information / parameters that are commonly applicable to one or more slices or pictures. The SPS (SPS syntax) may include information / parameters that are commonly applicable to one or more sequences. The VPS (VPS syntax) may include information / parameters that are commonly applicable to multiple layers. The DPS (DPS syntax) may include information / parameters that are commonly applicable to the entire image. The DPS may include information / parameters related to the concatenation of coded video sequences (CVS). In the present disclosure, the high-level syntax (HLS) may include at least one of the APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, picture header syntax, and slice header syntax.
[0101] In addition, as described above, generally, one NAL unit type can be set for one picture, and as described above, the NAL unit type can be signaled by nal_unit_type in the NAL unit header of the NAL unit including the slice. The following table shows examples of NAL unit type codes and NAL unit type classes.
[0102] [Table 1]
[0103]
[0104]
[0105] Furthermore, as described above, a picture can be composed of one or more slices. Furthermore, parameters describing a picture can be signaled via a picture header (PH), and parameters describing a slice can be signaled via a slice header (SH). The PH can be transmitted in its own NAL unit type. Furthermore, the SH can be present at the beginning of a NAL unit containing the slice's payload (i.e., slice data).
[0106] For example, the syntax elements of the signaled PH may be as follows.
[0107] [Table 2]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114] Furthermore, adopting PH may mean that each coded picture must have at least two NAL units. For example, one of the two units may be a NAL unit for PH, and the other may be a NAL unit for a coded slice including a slice header (SH) and slice data. This may be a problem for bitstreams with low bitrates, as the additional NAL units per picture may significantly affect the bitrate. Therefore, it may be desirable for PH to have a mode that does not consume new NAL units.
[0115] Therefore, the present disclosure proposes embodiments for solving the above problems. The proposed embodiments can be applied individually or in combination.
[0116] As an example, a method is proposed for signaling a flag indicating whether a PH NAL unit is present in a coded layer video sequence (CLVS) in an advanced parameter set. That is, the flag may indicate whether a picture header is present in a NAL unit (i.e., a PH NAL unit) or a slice header. Here, for example, CLVS may mean a sequence of picture units (PUs) having the same value of nuh_layer_id. A picture unit may be a set of NAL units for a coded picture. Furthermore, for example, an advanced parameter set may be a sequence parameter set (SPS), a picture parameter set (PPS), or a slice header. The flag may be referred to as ph_nal_present_flag. Alternatively, the flag may be referred to as a PH NAL present flag.
[0117] In addition, regarding the PH NAL presence flag, the present disclosure proposes an implementation method in which the value of ph_nal_present_flag is constrained to be the same for all SPSs referenced by pictures of the same CVS. This constraint may mean that the value of ph_nal_present_flag must be the same for a coded video sequence in a multi-layer bitstream.
[0118] Furthermore, as an example, when the value of ph_nal_present_flag is equal to 1, there is one PH NAL unit, and the PH NAL unit is associated with a video coding layer (VCL) NAL unit of the picture.
[0119] Furthermore, as an example, when the value of ph_nal_present_flag is equal to 0 (ie, when the PH NAL unit does not exist for each picture), a method of applying the following constraints is proposed.
[0120] For example, the above constraints can be as follows.
[0121] First, all pictures of a CLVS may include only one slice.
[0122] Secondly, the PH NAL unit may not exist. The PH syntax table may be present in the slice layer RBSP together with the slice header (SH) and slice data. That is, the PH syntax table may be present in the slice header.
[0123] Third, the PH syntax table and the SH syntax table can start at byte-aligned positions. To achieve this, a byte alignment bit can be added between PH and SH.
[0124] Fourth, the value of picture_header_extension_present_flag may be 0 in all PPSs referencing an SPS.
[0125] Fifth, all syntax elements that can be present in either PH or SH can be present in PH but not SH.
[0126] In addition, as an example, a method for updating access unit detection may be proposed. That is, instead of checking the PH, each new VCL NAL unit may mean a new access unit (AU). That is, when the value of ph_nal_present_flag indicates that there is no PH NAL unit, the VCL NAL unit including the ph_nal_present_flag is not the VCL NAL unit of the previous AU (i.e., the picture of the previous AU), and may mean parsing the VCL NAL unit of the new AU (i.e., the picture of the new AU). Therefore, when the value of ph_nal_present_flag indicates that there is no PH NAL unit, the VCL NAL unit including the ph_nal_present_flag may be the first VCL NAL unit of the picture of the new AU (e.g., the current picture to be decoded). Here, AU may mean a set of picture units (PUs) belonging to different layers and including coded pictures related to the same time of the output of the decoded picture buffer (DPB). In addition, PU may mean a set of NAL units including one coded picture that are associated and have a continuous decoding order. That is, PU may mean a set of NAL units for one coded picture that are associated and have a continuous decoding order. In addition, when the bitstream is a single-layer bitstream rather than a multi-layer bitstream, AU may be the same as PU.
[0127] The embodiments proposed in the present disclosure can be implemented as follows.
[0128] For example, the SPS syntax for signaling ph_nal_present_flag proposed in the embodiments of the present disclosure may be as follows.
[0129] [Table 3]
[0130]
[0131] Referring to Table 3, the SPS may include ph_nal_present_flag.
[0132] For example, the semantics of the syntax element ph_nal_present_flag may be as shown in the following table.
[0133] [Table 4]
[0134]
[0135] For example, referring to Table 4, the syntax element ph_nal_present_flag may indicate whether a NAL unit having the same nal_unit_type as PH_NUT exists for each coded picture of the CLVS referencing the SPS. For example, ph_nal_present_flag being equal to 1 may indicate that a NAL unit having the same nal_unit_type as PH_NUT exists for each coded picture of the CLVS referencing the SPS. In addition, for example, ph_nal_present_flag being equal to 0 may indicate that a NAL unit having the same nal_unit_type as PH_NUH does not exist for each coded picture of the CLVS referencing the SPS.
[0136] Furthermore, for example, when ph_nal_present_flag is 1, the following may apply.
[0137] - A NAL unit whose nal_unit_type is PH_NUT (ie, a PH NAL unit) may not be present in a CLVS referencing an SPS.
[0138] - Each picture of the CLVS referring to the SPS may include one slice.
[0139] -PH can exist in the slice layer RBSP.
[0140] In addition, although Tables 3 and 4 propose a method of signaling ph_nal_present_flag through SPS, the methods shown in Tables 3 and 4 are implementations proposed in the present disclosure, and an implementation may also be proposed of signaling ph_nal_present_flag through PPS or a slice header instead of SPS.
[0141] Furthermore, for example, according to an embodiment proposed in the present disclosure, the picture header syntax table and the picture header RBSP may be separately signaled as shown in the following table.
[0142] [Table 5]
[0143]
[0144] Furthermore, the signaled picture header syntax table may be as follows.
[0145] [Table 6]
[0146]
[0147]
[0148]
[0149]
[0150]
[0151]
[0152] Furthermore, according to an embodiment proposed in the present disclosure, for example, a slice-layer RBSP may be signaled as follows.
[0153] [Table 7]
[0154]
[0155] Furthermore, for example, according to the embodiments proposed in the present disclosure, one or more constraints shown in the following table may be applied.
[0156] [Table 8]
[0157]
[0158] For example, referring to Table 8, if ph_nal_present_flag is 0, the value of picture_header_extension_present_flag may be 0.
[0159] Additionally, for example, bitstream applicability may require that the value of pic_rpl_present_flag must be equal to 1 when the following two conditions are true.
[0160] - ph_nal_present_flag is 0, and the picture associated with PH is not an IDR picture.
[0161] - ph_nal_present_flag is 0, the picture associated with PH is an IDR picture, and sps_id_rpl_present_flag is equal to 1.
[0162] Here, rpl may mean reference picture list.
[0163] In addition, for example, when the value of ph_nal_present_flag is 0, the bitstream applicability may require that the value of pic_sao_enabled_present_flag must be equal to 1.
[0164] In addition, for example, when the value of ph_nal_present_flag is 0, the bitstream applicability may require that the value of pic_alf_enabled_present_flag must be equal to 1.
[0165] In addition, for example, when the value of ph_nal_present_flag is 0, the bitstream applicability may require that the value of pic_deblocking_filter_override_present_flag must be equal to 1.
[0166] Furthermore, for example, the embodiment can be applied according to the following procedure.
[0167] Figure 5 The encoding process according to an embodiment of the present disclosure is schematically illustrated.
[0168] Reference Figure 5 , the encoding apparatus may generate a NAL unit including information about a picture (S500). The encoding apparatus may determine whether a NAL unit for a picture header exists (S510) and may decide whether a NAL unit for a picture header exists (S520).
[0169] For example, when the NAL unit for the picture header exists, the encoding apparatus may generate a bitstream including a VCL NAL unit including a slice header and a PH NAL unit including a picture header ( S530 ).
[0170] In addition, for example, when there is no NAL unit for the picture header, the encoding apparatus may generate a bitstream including a VCL NAL unit including a slice header and a picture header (S540). That is, the picture header syntax structure may exist in the slice header.
[0171] Figure 6 The decoding process according to an embodiment of the present disclosure is schematically illustrated.
[0172] Reference Figure 6 , the decoding apparatus may receive a bitstream including a NAL unit (S600). Thereafter, the decoding apparatus may determine whether a NAL unit for a picture header exists (S610).
[0173] For example, when a NAL unit for a picture header exists, the decoding apparatus may decode / reconstruct a picture / slice / block / sample based on a slice header in a VCL NAL unit and a picture header in a PH NAL unit ( S620 ).
[0174] Also, for example, when the NAL unit for the picture header does not exist, the decoding apparatus may decode / reconstruct the picture / slice / block / sample based on the slice header and the picture header in the VCL NAL unit ( S630 ).
[0175] Here, the (coded) bitstream may include one or more NAL units for decoding a picture. In addition, the NAL unit may be a VCL NAL unit or a non-VCL NAL unit. For example, a VCL NAL unit may include information about a coded slice, and a VAL NAL unit may have a NAL unit type of the NAL unit type class "VCL" shown in Table 1 above.
[0176] In addition, according to the embodiments proposed in the present disclosure, the bitstream may include a PH NAL unit (NAL unit for a picture header), or the bitstream may not include a PH NAL unit for the current picture. Information indicating whether a PH NAL unit is present (e.g., ph_nal_present_flag) may be signaled through HLS (e.g., VPS, DPS, SPS, slice header, etc.).
[0177] Figure 7 The picture header configuration in the NAL unit according to whether the PH NAL unit exists is schematically shown. For example, Figure 7 (a) shows the case where there is a PH NAL unit for the current picture, Figure 7 (b) shows a case where there is no PH NAL unit for the current picture, but the picture header is included in the VCL NAL unit.
[0178] For example, when a PH NAL unit is present, the picture header may be included in the PH NAL unit. On the other hand, when the PH NAL unit is not present, the picture header may still be configured, but may be included in another type of NAL unit. For example, the picture header may be included in a VCL NAL unit. The VCL NAL unit may include information about coded slices. The VCL unit may include a slice header for a coded slice. For example, when a specific slice header includes information indicating that the coded / associated slice is the first slice in a picture or sub-picture, the picture header may be included in a specific VAL NAL unit including the specific slice header. Alternatively, for example, when the PH NAL unit is not present, the picture header may be included in a non-VCL NAL unit such as a PPS NAL unit, an APS NAL unit, or the like.
[0179] Figure 8 The image encoding method of the encoding device according to this document is schematically shown. Figure 8 The method disclosed in Figure 2 The encoding device shown is executed. Specifically, for example, Figure 8 S800 to S820 of the encoding apparatus may be performed by an entropy encoder. In addition, although not shown, the process of decoding the current picture may be performed by a predictor and a residual processor of the encoding apparatus.
[0180] The encoding device determines whether a picture header (PH) network abstraction layer (NAL) unit including a PH related to the current picture exists (S800). The encoding device may generate a NAL unit for the current picture. For example, the NAL unit for the current picture may include: a PH NAL unit including the PH related to the current picture and / or a video coding layer (VCL) NAL unit including information about slices in the current picture (e.g., slice header and slice data). The encoding device may determine whether a PH NAL unit exists. For example, when a PH NAL unit exists, the encoding device may generate a PH NAL unit including the PH related to the current picture and / or a video coding layer (VCL) NAL unit including information about slices in the current picture (e.g., slice header and slice data). Alternatively, for example, when the PH NAL unit does not exist, the encoding device may generate a video coding layer (VCL) NAL unit including the PH related to the current picture and information about one slice of the current picture (e.g., slice header and slice data). Furthermore, for example, when the flag indicates that the PH NAL unit does not exist, the current picture may include one slice. Here, for example, the PH may include syntax elements representing parameters of the current picture.
[0181] The encoding device generates a flag for whether a PH NAL unit exists based on the determination result (S810). For example, the encoding device may generate a flag for whether a PH NAL unit exists based on the determination result. For example, the flag may indicate whether a PH NAL unit exists. For example, when the value of the flag is 1, the flag may indicate the presence of a PH NAL unit, and when the value of the flag is 0, the flag may indicate the absence of a PH NAL unit. Alternatively, for example, when the value of the flag is 0, the flag may indicate the presence of a PH NAL unit, and when the value of the flag is 1, the flag may indicate the absence of a PH NAL unit. The syntax element of the flag may be the above-mentioned ph_nal_present_flag.
[0182] The encoding device encodes the image information including the flag (S820). The encoding device may encode the image information including the flag. The image information may include the flag. In addition, for example, the image information may include a high-level syntax, and the flag may be included in the high-level syntax. For example, the high-level syntax may be a sequence parameter set (SPS). Or, for example, the high-level syntax may be a slice header (SH). That is, for example, the flag may be included in a slice header.
[0183] Furthermore, for example, when the flag indicates the presence of a PH NAL unit, the PH may be included in the PH NAL unit, and when the flag indicates the absence of a PH NAL unit, the PH may be included in a slice header associated with the current picture. For example, when the flag indicates the presence of a PH NAL unit, the PH may be included in the PH NAL unit, and when the flag indicates the absence of a PH NAL unit, the PH may be included in a VCL NAL unit that includes a slice header. That is, for example, when the flag indicates the presence of a PH NAL unit, the PH may be included in the PH NAL unit, and when the flag indicates the absence of a PH NAL unit, the PH may be included in a slice header. For example, when the flag indicates the presence of a PH NAL unit, the image information may include a PH NAL unit that includes the PH and at least one VCL NAL unit that includes a slice header associated with the current picture, and when the flag indicates the absence of a PH NAL unit, the image information may include a VCL NAL unit that includes the PH and a slice header. Furthermore, for example, when the flag indicates the absence of a PH NAL unit, the image information may not include the PH NAL unit.
[0184] Furthermore, for example, when the flag indicates that the PH NAL unit does not exist, the PH NAL unit may not exist for all pictures in the coding layer video sequence (CLVS) including the current picture. That is, for example, the flag indicating whether the PH NAL unit exists for all pictures in the coding layer video sequence (CLVS) may have the same value. Furthermore, for example, when the flag indicates that the PH NAL unit does not exist, the picture headers of all pictures in the CLVS may be included in the slice headers of all pictures.
[0185] Furthermore, for example, AU detection can be modified from existing methods. For example, a new VCL NAL unit can mean a new AU. That is, for example, when the flag indicates that there is no PH NAL unit, the VCL NAL unit including the slice header can be the first VCL NAL unit of the current picture (for a new AU (i.e., an AU for the current picture)). For example, the flag can be included in the slice header of the VCL NAL unit. Or, for example, when the flag indicates that there is a PH NAL unit, it can be the first VCL NAL unit of the current picture after the PH NAL unit (i.e., signaled after the PH NAL unit).
[0186] In addition, the encoding device can decode the current picture. For example, the encoding device can decode the current picture based on the syntax elements of the PH. For example, the syntax elements of the PH can be the syntax elements shown in Table 6. The PH can include syntax elements representing parameters of the current picture, and the encoding device can decode the current picture based on these syntax elements. In addition, for example, the VCL NAL unit including the slice header can include slice data of the slice in the current picture, and the encoding device can decode the slice in the current picture based on the slice data. For example, the decoding device can derive the predicted samples and residual samples of the current picture, and generate reconstructed samples / reconstructed pictures based on the predicted samples and the residual samples.
[0187] Furthermore, for example, the encoding device may generate and encode prediction information for a block in the current picture. In this case, various prediction methods disclosed in the present disclosure (e.g., inter-frame prediction or intra-frame prediction) may be applied. For example, the encoding device may determine whether to perform inter-frame prediction or intra-frame prediction on a block, and may determine a specific inter-frame prediction mode or a specific intra-frame prediction mode based on the RD cost. Based on the determined mode, the encoding device may derive prediction samples for the block. The prediction information may include prediction mode information for the block. The image information may include prediction information.
[0188] In addition, for example, the encoding apparatus may encode residual information of a block of the current picture.
[0189] For example, the encoding device may derive residual samples by subtracting predicted samples from original samples of the block.
[0190] Then, for example, the encoding device may quantize the residual samples to derive quantized residual samples, may derive transform coefficients based on the quantized residual samples, and generate and encode residual information based on the transform coefficients. Alternatively, for example, the encoding device may quantize the residual samples to derive quantized residual samples, transform the quantized residual samples to derive transform coefficients, and generate and encode residual information based on the transform coefficients. The image information may include residual information. In addition, for example, the encoding device may encode the image information and output the encoded image information in the form of a bitstream.
[0191] The encoding device can generate a reconstructed sample and / or a reconstructed picture by adding the predicted sample and the residual sample. As described above, an in-loop filtering process such as deblocking filtering, SAO and / or ALF process can be applied to the reconstructed sample to improve subjective / objective picture quality.
[0192] In addition, the bit stream including the image information can be sent to the decoding device via a network or a (digital) storage medium. Here, the network can include a broadcast network and / or a communication network, and the digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD and SSD.
[0193] Figure 9 A coding device for performing the image coding method according to the present document is schematically shown. Figure 8 The method shown can be used by Figure 9 The encoding device shown is executed. Specifically, for example, Figure 9 The entropy encoder of the encoding device may perform S800 to S820. Although not shown, the process of decoding the current picture may be performed by a predictor and a residual processor of the encoding device.
[0194] Figure 10 The image decoding method of the decoding device according to this document is schematically shown. Figure 10 The method shown can be used by Figure 3 The decoding device shown is executed. Specifically, for example, Figure 10 S1000 to S1010 may be performed by an entropy decoder of a decoding device, Figure 10 S1020 may be performed by a predictor and a residual processor of the decoding device.
[0195] The decoding device obtains a flag for whether a picture header (PH) network abstraction layer (NAL) unit exists (S1000). The decoding device can obtain the flag for whether a picture header (PH) network abstraction layer (NAL) unit exists through the bitstream. For example, the decoding device can obtain image information through the bitstream, and the image information can include the flag. In addition, for example, the image information can include a high-level syntax, and the flag can be included in the high-level syntax. For example, the high-level syntax can be a sequence parameter set (SPS). Or, for example, the high-level syntax can be a slice header (SH). That is, for example, the flag can be included in the slice header.
[0196] For example, this flag may indicate whether a PH NAL unit exists. For example, when the value of this flag is 1, the flag may indicate the presence of a PH NAL unit, and when the value of this flag is 0, the flag may indicate the absence of a PH NAL unit. Alternatively, for example, when the value of this flag is 0, the flag may indicate the presence of a PH NAL unit, and when the value of this flag is 1, the flag may indicate the absence of a PH NAL unit. The syntax element of this flag may be the above-mentioned ph_nal_present_flag.
[0197] The decoding device obtains PH based on the flag (S1010). The decoding device can obtain PH from the PH NAL unit or the VCL NAL unit including the slice header based on the flag. That is, for example, the decoding device can obtain PH from the PH NAL unit or the slice header based on the flag.
[0198] For example, when the flag indicates the presence of a PH NAL unit, the PH can be obtained from the PH NAL unit, and when the flag indicates the absence of a PH NAL unit, the PH can be obtained from the slice header. For example, when the flag indicates the presence of a PH NAL unit, the PH can be included in the PH NAL unit, and when the flag indicates the absence of a PH NAL unit, the PH can be included in a VCL NAL unit including a slice header. That is, for example, when the flag indicates the presence of a PH NAL unit, the PH can be included in the PH NAL unit, and when the flag indicates the absence of a PH NAL unit, the PH can be included in the slice header. For example, when the flag indicates the presence of a PH NAL unit, the image information can include a PH NAL unit including the PH and a VCL NAL unit including the slice header, and when the flag indicates the absence of a PH NAL unit, the image information can include a VCL NAL unit including the PH and the slice header. Furthermore, for example, when the flag indicates the absence of a PH NAL unit, the image information may not include the PH NAL unit.
[0199] In addition, for example, when the flag indicates that the PH NAL unit does not exist, the current picture of the PH may include one slice. That is, for example, when the flag indicates that the PH NAL unit does not exist, the image information may include a VCL NAL unit containing a slice header of one slice in the current picture.
[0200] In addition, for example, when the flag indicates the presence of the PH NAL unit, the PH NAL unit for the current picture and at least one VCL NAL unit including a slice header of the current picture can be obtained through the bitstream. That is, for example, when the flag indicates the presence of the PH NAL unit, the image information can include the PH NAL unit for the current picture and a VCL NAL unit including a slice header of at least one slice in the current picture.
[0201] Furthermore, for example, when the flag indicates that the PH NAL unit does not exist, the PH NAL unit may not exist for all pictures in the coding layer video sequence (CLVS) including the current picture. That is, for example, the flag indicating whether the PH NAL unit exists for all pictures in the coding layer video sequence (CLVS) may have the same value. Furthermore, for example, when the flag indicates that the PH NAL unit does not exist, the picture headers of all pictures in the CLVS may be included in the slice headers of all pictures.
[0202] Furthermore, for example, AU detection can be modified from existing methods. For example, a new VCL NAL unit can mean a new AU. That is, for example, when the flag indicates that there is no PH NAL unit, the VCL NAL unit including the slice header can be the first VCL NAL unit of the current picture (for a new AU (i.e., an AU for the current picture)). For example, the flag can be included in the slice header of the VCL NAL unit. Or, for example, when the flag indicates that there is a PH NAL unit, it can be the first VCL NAL unit of the current picture after the PH NAL unit (i.e., signaled after the PH NAL unit).
[0203] The decoding device decodes the current picture related to the PH based on the PH (S1020). The decoding device can decode the current picture based on the syntax elements of the PH. For example, the syntax elements of the PH can be the syntax elements shown in Table 6. The PH can include syntax elements representing parameters of the current picture, and the decoding device can decode the current picture based on these syntax elements. In addition, for example, the VCL NAL unit including the slice header can include slice data of the slice in the current picture, and the decoding device can decode the slice in the current picture based on the slice data. For example, the decoding device can derive the predicted samples and residual samples of the current picture, and generate the reconstructed samples / reconstructed pictures of the current picture based on the predicted samples and the residual samples.
[0204] As described above, in-loop filtering processes such as deblocking filtering, SAO and / or ALF processes may be applied to the reconstructed samples as needed in order to improve subjective / objective picture quality.
[0205] Figure 11 A decoding device for executing the image decoding method according to the present document is schematically shown. Figure 10 The method shown can be used by Figure 11 The decoding device shown is executed. Specifically, for example, Figure 11 The entropy decoder of the decoding device can perform Figure 10 S1000 to S1010, Figure 11 The predictor and residual processor of the decoding device can perform Figure 10 S1020.
[0206] According to the above disclosure, a flag indicating whether a PH NAL unit exists can be signaled, the NAL unit can be adaptively adjusted according to the bit rate of the bitstream based on the flag, and the overall coding efficiency can be improved.
[0207] In addition, according to the present disclosure, constraints on the number of slices in the current picture and on whether the PH NAL unit exists can be set for the relevant picture based on the flag indicating whether the PH NAL unit exists and adaptive to the bit rate control NAL unit, thereby improving overall encoding efficiency.
[0208] In the above-mentioned embodiment, method is described based on the flow chart with a series of steps or square frames. The present disclosure is not limited to the order of the above steps or square frames. Some steps or square frames can be performed in an order different from the above-mentioned other steps or square frames or performed simultaneously. In addition, it will be understood by those skilled in the art that the steps shown in the flow chart are not exclusive and may also include other steps, or may delete one or more steps in the flow chart without affecting the scope of the present disclosure.
[0209] The embodiments described in this specification can be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information (e.g., information about instructions) or algorithms used for implementation can be stored in a digital storage medium.
[0210] In addition, the decoding device and encoding device to which the present disclosure is applied may be included in the following devices: multimedia broadcast transmission / reception devices, mobile communication terminals, home theater video devices, digital theater video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, portable cameras, VoD service providing devices, over-the-top (OTT) video devices, Internet streaming service providing devices, three-dimensional (3D) video devices, teleconferencing video devices, transportation user devices (e.g., vehicle user devices, aircraft user devices, and ship user devices), and medical video devices; and the decoding device and encoding device to which the present disclosure is applied may be used to process video signals or data signals. For example, over-the-top (OTT) video devices may include game consoles, Blu-ray players, Internet access televisions, home theater systems, smart phones, tablet computers, digital video recorders (DVRs), etc.
[0211] In addition, the processing method of the present invention can be generated in the form of a program executed by a computer and can be stored in a computer-readable recording medium. The multimedia data with a data structure according to the present invention can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a BD, a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (for example, via transmission over the Internet). In addition, the bit stream generated by the encoding method can be stored in a computer-readable recording medium or transmitted via a wired / wireless communication network.
[0212] In addition, the embodiments of the present disclosure may be implemented using a computer program product according to a program code, and the program code may be executed in a computer through the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0213] Figure 12 The structure diagram of the content streaming system to which the present disclosure is applied is illustrated.
[0214] A content streaming system to which embodiments of the present disclosure are applied may mainly include an encoding server, a streaming server, a network server, a media storage, a user device, and a multimedia input device.
[0215] The encoding server compresses content input from a multimedia input device such as a smartphone, camera, or camcorder into digital data to generate a bitstream and transmits the bitstream to the streaming server. As another example, when the multimedia input device such as a smartphone, camera, or camcorder directly generates the bitstream, the encoding server can be omitted.
[0216] A bitstream may be generated by an encoding method or a bitstream generating method to which an embodiment of the present disclosure is applied, and a streaming server may temporarily store the bitstream in the process of transmitting or receiving the bitstream.
[0217] The streaming server transmits multimedia data to user devices via a network server based on user requests, and the network server serves as an intermediary for notifying users of services. When a user requests a desired service from the network server, the network server transmits the request to the streaming server, which then transmits the multimedia data to the user. In this case, the content streaming system may include a separate control server. In this case, the control server is used to control commands and responses between devices within the content streaming system.
[0218] The streaming server can receive content from a media storage and / or encoding server. For example, when receiving content from an encoding server, the content can be received in real time. In this case, in order to provide a stable streaming service, the streaming server can store the bitstream for a predetermined time.
[0219] Examples of user devices may include mobile phones, smart phones, laptop computers, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touch-screen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smart watches, smart glasses, and head-mounted displays), digital TVs, desktop computers, and digital signage, etc. Each server within the content streaming system may operate as a distributed server, in which case data received from each server may be distributed.
[0220] The claims described in this disclosure can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined to implement a device, and the technical features of the device claims of this disclosure can be combined to implement a method. Furthermore, the technical features of the method claims of this disclosure and the technical features of the device claims of this disclosure can be combined to implement a device, and the technical features of the method claims of this disclosure and the technical features of the device claims of this disclosure can be combined to implement a method.
Claims
1. An image decoding method performed by a decoding device, the image decoding method comprising the following steps: Get the flag of whether the picture header PH network abstraction layer NAL unit exists; Based on the flag indicating that the PH NAL unit exists, obtaining a picture header for a current picture from the PH NAL unit, where the picture header includes parameters for the current picture; obtaining, based on the flag indicating that the PH NAL unit is absent, the picture header for the current picture from a slice layer raw byte sequence payload RBSP of the current picture, the slice layer RBSP including a slice header, the slice header including parameters of the slice for the current picture; as well as Decoding the current picture based on the PH, Wherein, based on the flag indicating that the PH NAL unit does not exist, the current picture includes only one slice, wherein, based on the flag indicating that the PH NAL unit does not exist, the slice layer RBSP includes all of the parameters of the picture header and the parameters of the slice header, and Wherein, based on the flag indicating the presence of the PH NAL unit, the slice layer RBSP only includes the parameters of the slice header.
2. An image encoding method performed by an encoding device, the image encoding method comprising the following steps: Determining whether there is a picture header (PH) network abstraction layer (NAL) unit including a picture header related to the current picture; Based on determining that the PH NAL unit exists, generating the PH NAL unit including the picture header, the picture header including parameters for the current picture; generating, based on determining that the PH NAL unit does not exist, a slice layer raw byte sequence payload RBSP for the current picture, the slice layer RBSP including the picture header, the slice layer RBSP including a slice header, the slice header including parameters for the slice of the current picture; generating a flag indicating whether the PH NAL unit exists based on a result of the determination; as well as encoding the image information including the logo, Wherein, based on the flag indicating that the PH NAL unit does not exist, the current picture includes only one slice, wherein, based on the flag indicating that the PH NAL unit does not exist, the slice layer RBSP includes all of the parameters of the picture header and the parameters of the slice header, and Wherein, based on the flag indicating the presence of the PH NAL unit, the slice layer RBSP only includes the parameters of the slice header.
3. A method for transmitting image data, the method comprising the following steps: Obtaining a bitstream of image information, wherein the image information includes a flag indicating whether a picture header (PH) network abstraction layer (NAL) unit exists; as well as sending the data of the bit stream including the image information, the image information including the flag, wherein, based on the flag indicating the presence of the PH NAL unit, a picture header for the current picture is obtained from the PH NAL unit, the picture header including parameters for the current picture; wherein, based on the flag indicating that the PH NAL unit does not exist, obtaining the picture header for the current picture from a slice layer raw byte sequence payload RBSP of the current picture, the slice layer RBSP including a slice header, the slice header including parameters of the slice for the current picture, Wherein, based on the flag indicating that the PH NAL unit does not exist, the current picture includes only one slice, wherein, based on the flag indicating that the PH NAL unit does not exist, the slice layer RBSP includes all of the parameters of the picture header and the parameters of the slice header, and Wherein, based on the flag indicating the presence of the PH NAL unit, the slice layer RBSP only includes the parameters of the slice header.