Video decoding device, video encoding device, and device for transmitting image data
By utilizing the sps_video_parameter_set_id syntax element in the video decoding and encoding device, the number of output layer sets is inferred to be 1, enabling inter-frame or intra-frame prediction. This solves the problem of efficient compression of high-resolution images/videos, and improves coding efficiency and single-layer bitstream processing capabilities.
Patent Information
- Application Number
- CN202511225402.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-12
- Filing Date
- 2021-05-11
- Publication Date
- 2025-11-14
AI Technical Summary
Existing image/video coding technologies have high transmission and storage costs when processing high-resolution, high-quality images/videos, and are difficult to efficiently compress and process parameter sets in single-layer bitstreams.
By using the video decoding and encoding devices, and taking advantage of the fact that the value of the sps_video_parameter_set_id syntax element is equal to 0, it is inferred that the total number of output layer sets OLS and the number of layers are 1, thus realizing inter-frame or intra-frame prediction, and generating prediction samples in a single-layer bitstream to reconstruct the current block.
It improves the encoding efficiency of images/videos, enhances the processing efficiency of single-layer bitstreams, effectively derives the layer identifier of the parameter set and the information of the output layer set, and reduces transmission and storage costs.
Smart Images

Figure CN120956916A_ABST
Abstract
Description
[0001] This application is a divisional application of the original patent application No. 202180034840.8 (International Application No.: PCT / KR2021 / 005896, Application Date: May 11, 2021, Invention Title: Method and Apparatus for Processing References to Parameter Sets within a Single-Layer Bitstream in an Image / Video Coding System). Technical Field
[0002] This disclosure relates to a method and apparatus for processing a set of parameters within a single-layer bitstream when encoding / decoding image / video information in an image / video coding system. Background Technology
[0003] Recently, there has been an increasing demand for high-resolution, high-quality images / videos (such as 4K, 8K, or even Ultra High Definition (UHD) images / videos) across various fields. As image / video resolution or quality increases, more information or bits are transmitted compared to traditional image / video data. Therefore, transmitting and storing image / video data via media such as existing wired / wireless broadband lines or stored in traditional storage media increases transmission and storage costs.
[0004] Furthermore, there is growing interest and demand for virtual reality (VR) and artificial reality (AR) content, as well as immersive media such as holograms; and the broadcasting of images / videos that present characteristics different from those of actual images / videos, such as game images / videos, is also on the rise.
[0005] Therefore, efficient image / video compression technology is needed to effectively compress and send, store, or play high-resolution, high-quality images / videos exhibiting the various characteristics described above. Summary of the Invention
[0006] Technical issues
[0007] The technical objective of this disclosure is to provide methods and apparatus for improving the coding efficiency of images / videos.
[0008] Another technical objective of this disclosure is to provide a method and apparatus for efficiently processing single-layer bitstreams.
[0009] Another technical objective of this disclosure is to provide a method and apparatus for deriving a layer identifier of a parameter set referenced by a VCL NAL cell when the bitstream is a single-layer bitstream.
[0010] Another technical objective of this disclosure is to provide a method and apparatus for deriving information about the output layer set when the bit stream is a single-layer bit stream.
[0011] Technical solution
[0012] According to embodiments of this specification, a video decoding method performed by a video decoding device is provided. The method may include the following steps: obtaining image information including Video Coding Layer (VCL) Network Abstraction Layer (NAL) units from a bitstream; generating a prediction sample for the current block by performing inter-frame prediction or intra-frame prediction on the current block within the current image based on the image information; and reconstructing the current block based on the prediction sample, wherein the image information may include an `sps_video_parameter_set_id` syntax element indicating the value of a video parameter set; and wherein, based on the value of the `sps_video_parameter_set_id` syntax element being equal to 0, the total number of output layer sets (OLS) specified by the video parameter set and the number of layers within the OLS can be inferred to be equal to 1.
[0013] According to another embodiment of this specification, a video coding method performed by a video coding apparatus is provided. The method may include the following steps: performing inter-frame prediction or intra-frame prediction on a current block within a current image; generating prediction information for the current block based on the inter-frame prediction or the intra-frame prediction; and encoding image information including the prediction information, wherein the image information may include video coding layer (VCL) network abstraction layer (NAL) units and an `sps_video_parameter_set_id` syntax element indicating the value of a video parameter set; and wherein, based on the value of the `sps_video_parameter_set_id` syntax element being equal to 0, the total number of output layer sets (OLS) specified by the video parameter set and the number of layers within the OLS can be inferred to be equal to 1.
[0014] According to another embodiment of this specification, a computer-readable digital recording medium is provided that stores information that causes a video decoding apparatus to perform a video decoding method, wherein the video decoding method may include the following steps: obtaining image information including video coding layer (VCL) network abstraction layer (NAL) units; generating a prediction sample of the current block by performing inter-frame prediction or intra-frame prediction on the current block within the current image based on the image information; and reconstructing the current block based on the prediction sample, wherein the image information may include an sps_video_parameter_set_id syntax element indicating an identifier value of a video parameter set; and wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, the value of the total number of output layer sets (OLS) specified by the video parameter set, and the value of the number of layers within the OLS, can be inferred to be equal to 1.
[0015] Technical effect
[0016] According to embodiments of this disclosure, the overall compression efficiency of images / videos can be enhanced.
[0017] According to the embodiments of this disclosure, single-layer bitstreams can be processed efficiently.
[0018] According to embodiments of this disclosure, when the bitstream is a single-layer bitstream, the layer identifier of the parameter set referenced by the VCL NAL unit can be derived.
[0019] According to embodiments of this disclosure, constraints corresponding to a single-layer bitstream can be provided.
[0020] According to embodiments of this disclosure, even when the bitstream is a single-layer bitstream that does not include a video parameter set, information about the output layer set can be obtained. Attached Figure Description
[0021] Figure 1 An example of a video / image coding system to which the embodiments described herein can be applied is illustrated schematically.
[0022] Figure 2 This is a diagram schematically illustrating the configuration of a video / image encoding apparatus to which the embodiments described herein may be applied.
[0023] Figure 3 This is a diagram illustrating the configuration of a multi-layer-based video / image coding apparatus to which embodiments of the present disclosure may be applied.
[0024] Figure 4 This is a diagram illustrating the configuration of a video / image decoding apparatus to which embodiments of the present disclosure may be applied.
[0025] Figure 5 This is a diagram illustrating the configuration of a multi-layer-based video / image decoding apparatus to which embodiments of the present disclosure may be applied.
[0026] Figure 6 A schematic example of an image decoding process to which embodiments of the present disclosure may be applied is shown.
[0027] Figure 7 A schematic example of an image encoding process to which embodiments of the present disclosure can be applied is shown.
[0028] Figure 8 An example of a layered structure for encoded images / videos is shown.
[0029] Figure 9 and Figure 10 General examples of video / image encoding methods and related components according to embodiments of the present disclosure are shown respectively.
[0030] Figure 11 and Figure 12 General examples of video / image decoding methods and related components according to embodiments of the present disclosure are shown respectively.
[0031] Figure 13 An example of a content streaming system to which embodiments of the present disclosure can be applied is shown. Detailed Implementation
[0032] The disclosure of this document may be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. The terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the methods disclosed herein. Expressions of a single number include the expression "at least one," provided that it is clearly read differently. Terms such as "comprising" and "having" are intended to indicate the presence of features, numbers, steps, operations, elements, components, or combinations thereof used in the disclosure, and therefore should be understood that the possibility of having or adding one or more different features, numbers, steps, operations, elements, components, or combinations thereof is not excluded.
[0033] This disclosure relates to video / image coding. For example, the methods / implementations disclosed in this disclosure can be applied to methods disclosed in the Universal Video Coding (VVC) standard. Additionally, the methods / implementations disclosed in this disclosure can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AO Media Video 1 (AV1) standard, the second-generation Audio Video Coding Standard (AVS2), or next-generation video / image coding standards (e.g., H.267, H.268, etc.).
[0034] Various implementations relating to video / image encoding are presented in this disclosure, and unless otherwise stated, the implementations may be combined with each other.
[0035] Furthermore, the various configurations in the accompanying drawings described in this disclosure are independent illustrative diagrams used to explain the functions that are different features from each other, and do not imply that the various configurations are implemented by different hardware or different software. For example, two or more configurations in a configuration may be combined to form a single configuration, and a single configuration may also be divided into multiple configurations. Embodiments that combine and / or separate configurations are included within the scope of this disclosure without departing from the spirit of the methods disclosed herein.
[0036] In this disclosure, the terms “ / ” and “,” should be interpreted as indicating “and / or”. For example, the expression “A / B” can mean “A and / or B”. Furthermore, “A, B” can mean “A and / or B”. Additionally, “A / B / C” can mean “at least one of A, B and / or C”. Furthermore, “A / B / C” can mean “at least one of A, B and / or C”.
[0037] Furthermore, in this disclosure, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can include: 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this disclosure should be interpreted as indicating "alternatively or alternatively".
[0038] Furthermore, the parentheses used in this disclosure may mean "for example". Specifically, when referring to "prediction (intra-frame prediction)", it may indicate that "intra-frame prediction" is proposed as an example of "prediction". In other words, the word "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" is proposed as an example of "prediction". Moreover, even when referring to "prediction (i.e., intra-frame prediction)", it may indicate that "intra-frame prediction" is proposed as an example of "prediction".
[0039] In this disclosure, technical features may be implemented individually or simultaneously and explained individually in one of the accompanying drawings.
[0040] In the following, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Furthermore, throughout the drawings, the same reference numerals are used to indicate the same elements, and the same descriptions of the same elements may be omitted.
[0041] Figure 1 Examples of video / image coding systems to which embodiments of the present disclosure can be applied are illustrated.
[0042] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiver). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0043] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0044] Video sources can acquire video / images through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, video / image archives including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc. In this case, the video / image capture process can be replaced by a process that generates related data.
[0045] An encoding device can encode input video / images. It can perform a series of processes such as prediction, transformation, and quantization to achieve compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0046] The transmitter can send encoded images / image information or data, output as a bitstream, to the receiver of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received bitstream to a decoding device.
[0047] Decoding devices can decode video / images by performing a series of processes, such as dequantization, inverse transform, and prediction, that correspond to the operations of encoding devices.
[0048] The renderer can render decoded video / images. The rendered video / images can then be displayed on a monitor.
[0049] In this disclosure, video can refer to a series of images over time. An image generally refers to a unit representing an image at a specific time frame, and a slice / tile refers to a unit that constitutes part of an image in terms of coding. A slice / tile may include one or more coding tree units (CTUs). An image may consist of one or more slices / tiles. An image may consist of one or more groups of tiles. A group of tiles may include one or more tiles. A tile may represent a rectangular area of a row of CTUs within a tile in an image. A tile may be divided into multiple tiles, each of which consists of one or more rows of CTUs within the tile. A tile that is not divided into multiple tiles may also be referred to as a tile. A tile scan is a specific ordering of the CTUs that divide the image, wherein the CTUs are ordered consecutively in a CTU raster scan within a tile, tiles within a tile are ordered consecutively in a raster scan of tiles of a tile, and tiles in an image are ordered consecutively in a raster scan of tiles of an image. A tile is a rectangular area of a CTU within a specific tile column and a specific tile row in an image. A tile column is a rectangular region of multiple CTUs having a height equal to the height of the image and a width specified by a syntax element in the image parameter set. A tile scan is a specific ordering of the CTUs that segment the image, where the CTUs are ordered consecutively in a CTU raster scan within a tile, and the tiles in the image are ordered consecutively in a tile raster scan within the image. A slice comprises an integer number of tiles of an image that can be proprietaryly included in a single NAL unit. A slice can consist of a consecutive sequence of multiple complete tiles or only complete tiles of a single tile. In this disclosure, tile groups and slices are used interchangeably. For example, in this disclosure, a tile group / tile group header can be referred to as a slice / slice header.
[0050] A pixel, or image unit, can refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0051] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the terms "unit" and "block" or "region" may be used interchangeably. Typically, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients in M columns and N rows. Alternatively, samples may represent pixel values in the spatial domain, and when such pixel values are transformed to the frequency domain, they may represent transform coefficients in the frequency domain.
[0052] In some cases, a unit can be used interchangeably with terms such as block or region. Typically, an M×N block can represent a sample consisting of M columns and N rows or a set of transform coefficients. A sample can typically represent a pixel or pixel value, and can also represent only the pixel / pixel value of the luminance component, and only the pixel / pixel value of the chrominance component. A sample can be used as an item corresponding to the pixels or picometers of a picture (or image).
[0053] Figure 2 This is a schematic diagram illustrating the configuration of a video / image encoding apparatus to which embodiments of the present disclosure may be applied. Hereinafter, the video encoding apparatus may include an image encoding apparatus.
[0054] Reference Figure 2 The encoding apparatus 200 may include and be configured with an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 described above may be constituted by at least one hardware component (e.g., an encoder chipset or a processor). Additionally, the memory 270 may include a decoded image buffer (DPB) or may be constituted by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0055] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, coding units can be recursively segmented from coding tree units (CTUs) or maximum coding units (LCUs) according to a quadtree-binary-tritree (QTBTTT) structure. For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary tree structure. Alternatively, a binary tree structure can be applied first. The encoding process according to this disclosure can be performed based on the final coding unit that is no longer segmented. In this case, the maximum coding unit can be used as the final coding unit based on encoding efficiency according to image characteristics, or, if necessary, the coding unit can be recursively segmented into deeper coding units, and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be separated or divided from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving the residual signal from the transform coefficients.
[0056] In the encoding apparatus 200, a prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 can be subtracted from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to the converter 232. In this case, as shown, the unit in the encoding apparatus 200 used to subtract the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called subtractor 231. Predictor 220 can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction on a unit of the current block or CU. As described later in the description of each prediction mode, predictor 220 can generate various information related to the prediction, such as prediction mode information, and send the generated information to entropy encoder 240. The information about the prediction can be encoded in entropy encoder 240 and output in the form of a bitstream.
[0057] Intra-predictor 222 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples can be located near or far from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, the directional modes can include, for example, 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 222 can also determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.
[0058] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be called a juxtaposed reference block, a co-located CU (colCU), etc., and the reference image including the temporally neighboring block may be called a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, the motion vector of the current block can be indicated by using motion vectors from neighboring blocks as motion vector predictors and signaling the motion vector difference.
[0059] Predictor 220 can generate prediction signals based on various prediction methods described below. For example, predictor 220 can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as Inter-Frame Intra-Frame Combined Prediction (CIIP). Alternatively, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. IBC prediction mode or palette mode can be used for image / video coding of content such as games, for example, Screen Content Coding (SCC). IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction because the reference block is derived in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure. Palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying palette mode, sample values within the frame can be signaled based on information about the palette table and palette index.
[0060] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate the reconstructed signal or the residual signal.
[0061] Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to the transform obtained from a graph when the relationship information between pixels is represented by a graph. CNT refers to the transform generated based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or it can be applied to blocks of variable size that are not square.
[0062] Quantizer 233 quantizes the transform coefficients and sends them to entropy encoder 240, which encodes the quantized signal (information about the quantized transform coefficients) and outputs a bitstream. This information about the quantized transform coefficients can be called residual information. Quantizer 233 can rearrange the block-form quantized transform coefficients into a one-dimensional vector based on the coefficient scan order, and generate information about the quantized transform coefficients based on this one-dimensional vector form.
[0063] The entropy encoder 240 can perform various encoding methods, such as, for example, Golomb, Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), etc. The entropy encoder 240 can encode information required for video / image reconstruction (e.g., values of syntax elements, etc.) together or separately, except for quantization transform coefficients. It can transmit or store encoded information (e.g., encoded video / image information) in the form of a bitstream at NAL (Network Abstraction Layer) units. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may also include general constraint information. In this disclosure, information and / or syntax elements that are transmitted / signaled from the encoding device to the decoding device may be included in the video / image information. The video / image information can be encoded by the above-described encoding process and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include a broadcast network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitting unit (not shown) for transmitting signals output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signals may be included as internal / external components of the encoding device 200, and alternatively, the transmitter may be included in the entropy encoder 240.
[0064] The quantization transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantization transform coefficients using dequantizer 234 and inverse transformer 235. Adder 250 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If the block to be processed has no residual (such as when a skip mode is applied), the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstruction unit or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be used for inter-frame prediction of the next image by filtering as described below.
[0065] In addition, Luminance Mapping and Chromatography Scaling (LMCS) can be applied during the image encoding and / or reconstruction process.
[0066] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 270 (specifically, the DPB of memory 270). Various filtering methods may include, for example, deblocking filtering, sample adaptive offsetting, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various filtering-related information and send the generated information to entropy encoder 240, as described later in the description of the various filtering methods. The filtering-related information can be encoded by entropy encoder 240 and output as a bitstream.
[0067] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. When inter-frame prediction is applied through the encoding device, prediction mismatch between encoding device 200 and decoding device can be avoided, and encoding efficiency can be improved.
[0068] The DPB of memory 270 can store a modified reconstructed image used as a reference image in inter-frame predictor 221. Memory 270 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of reconstructed blocks in the image. The stored motion information can be sent to inter-frame predictor 221 and used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and can transmit these reconstructed samples to intra-frame predictor 222.
[0069] Furthermore, image / video coding according to this disclosure can include multi-layer image / video coding. Multi-layer image / video coding can include scalable coding. Multi-layer coding or scalable coding can process input signals from individual layers. The input signals (input images / pictures) can vary depending on the layers in at least one of resolution, frame rate, bit depth, color format, aspect ratio, and view. In this case, redundant transmission / processing of information can be reduced and compression efficiency improved by performing prediction between layers using the differences between layers (i.e., based on scalability).
[0070] Figure 3 This is a block diagram of a multi-layer-based encoding apparatus for executing video / image signals according to embodiments of the present disclosure.
[0071] Figure 3 The encoding device may include Figure 2 The encoding device. In Figure 3In this specification, the image segmenter and adder are omitted. However, the encoding device may include both an image segmenter and an adder. In this case, the image segmenter and adder may be included on a layer-by-layer basis. This disclosure primarily describes multi-layer-based prediction. Further descriptions may be found in the references to [reference needed]. Figure 2 The given description.
[0072] For ease of description, Figure 3 The example assumes a multi-layer structure consisting of two layers. However, embodiments of this disclosure are not limited to the specific example, and it should be noted that multi-layer structures to which embodiments of this disclosure are applied may include two or more layers.
[0073] Reference Figure 3 The encoding device 200 includes an encoder 200-1 for layer 1 and an encoder 200-0 for layer 0.
[0074] Layer 0 can be a base layer, a reference layer, or a lower layer; Layer 1 can be an enhancement layer, the current layer, or a higher layer.
[0075] The encoder 200-1 of layer 1 includes a predictor 220-1, a residual processor 230-1, a filter 260-1, a memory 270-1, an entropy encoder 240-1, and a multiplexer (MUX) 270. The MUX may be included as an external component.
[0076] The encoder 200-0 of layer 0 includes a predictor 220-0, a residual processor 230-0, a filter 260-0, a memory 270-0, and an entropy encoder 240-0.
[0077] Predictors 220-0 and 220-1 can perform predictions on the input image based on various prediction techniques described above. For example, predictors 220-0 and 220-1 can perform inter-frame prediction and intra-frame prediction. Predictors 220-0 and 220-1 can perform predictions in predetermined processing units. The prediction unit can be a coding unit (CU) or a transform unit (TU). Prediction blocks (including prediction samples) can be generated based on the prediction results, and the residual processor can derive residual blocks (including residual samples) based on the prediction blocks.
[0078] Inter-frame prediction generates prediction blocks by performing predictions based on information from at least one of the previous and / or subsequent images of the current image. Intra-frame prediction generates prediction blocks by performing predictions based on neighboring samples within the current image.
[0079] The various prediction mode methods described above can be used for inter-frame prediction modes or methods. Inter-frame prediction can select a reference image relative to the current block to be predicted and a reference block within the reference image related to the current block. Predictors 220-0 and 220-1 can generate prediction blocks based on the reference blocks.
[0080] Furthermore, predictor 220-1 can use information from layer 0 to perform predictions for layer 1. In this disclosure, for ease of description, the method of using information from another layer to predict information from the current layer is referred to as inter-layer prediction.
[0081] Information of the current layer predicted based on information from another layer (i.e., predicted through inter-layer prediction) includes at least one of texture, motion information, cell information, and predetermined parameters (e.g., filtering parameters).
[0082] Furthermore, the information of another layer used for prediction of the current layer (i.e., for inter-layer prediction) may include at least one of texture, motion information, cell information, and predetermined parameters (e.g., filtering parameters).
[0083] In inter-layer prediction, the current block can be a block within the current picture of the current layer (e.g., layer 1) and can be the target block to be encoded. The reference block can be a block within a picture (reference picture) of the layer (reference layer, e.g., layer 0) referenced for the prediction of the current block, belonging to the same access unit (AU) as the picture to which the current block belongs (the current picture), and can be a block corresponding to the current block. Here, an access unit can be a set of picture units (PUs) including encoded pictures associated with the same temporal outputs from different layers and DPBs. Picture units can be a set of NAL units that are correlated with each other according to a specific classification rule, are consecutive in decoding order, and contain only one encoded picture. The encoded video sequence (CVS) can be a set of AUs.
[0084] An example of inter-layer prediction is inter-layer motion prediction, which uses motion information from a reference layer to predict the motion information of the current layer. Based on inter-layer motion prediction, the motion information of the current block can be predicted based on the motion information of the reference block. In other words, when deriving motion information based on the inter-frame prediction mode (described later), motion information candidates can be derived using motion information from an inter-layer reference block rather than temporally neighboring blocks.
[0085] When applying interlayer motion prediction, predictor 220-1 can scale and use motion information from the reference block of the reference layer (i.e., the interlayer reference block).
[0086] In another example of inter-layer prediction, inter-layer texture prediction can use the texture of the reconstructed reference block as the predicted value for the current block. In this case, predictor 220-1 can scale the texture of the reference block by upsampling. Inter-layer texture prediction can be referred to as inter-layer (reconstructed) sample prediction or simply as inter-layer prediction.
[0087] In interlayer parameter prediction (another example of interlayer prediction), parameters derived from the reference layer can be reused in the current layer, or the parameters of the current layer can be derived based on the parameters used in the reference layer.
[0088] In interlayer residual prediction (another example of interlayer prediction), residual information from another layer can be used to predict the residual of the current layer, and prediction of the current block can be performed based on the predicted residual.
[0089] In inter-layer differential prediction (another example of inter-layer prediction), prediction of the current block can be performed using the difference between the images obtained by upsampling or downsampling the reconstructed image of the current layer and the reconstructed image of the reference layer.
[0090] In inter-layer syntax prediction (another example of inter-layer prediction), the syntax information of a reference layer can be used to predict or generate the texture of the current block. In this case, the syntax information of the referenced layer can include information about intra-frame prediction modes and motion information.
[0091] When predicting a specific block, multiple prediction methods using inter-layer prediction can be used across multiple layers.
[0092] Here, as examples of interlayer prediction, interlayer texture prediction, interlayer motion prediction, interlayer cell information prediction, interlayer parameter prediction, interlayer residual prediction, interlayer difference prediction, and interlayer syntax prediction have been described; however, the interlayer prediction applicable to this disclosure is not limited to the examples above.
[0093] For example, inter-layer prediction can be applied as an extension of inter-frame prediction for the current layer. In other words, inter-frame prediction for the current block can be performed by including a reference image derived from a reference layer in a reference image that can be referenced for inter-frame prediction of the current block.
[0094] In this case, inter-layer reference images can be included in the reference image list for the current block. Using the inter-layer reference images, predictor 220-1 can perform inter-frame prediction for the current block.
[0095] Here, the inter-layer reference image can be a reference image constructed by sampling the reconstructed image of a reference layer to correspond to the current layer. Therefore, when the reconstructed image of the reference layer corresponds to the image of the current layer, the reconstructed image of the reference layer can be used as the inter-layer reference image without resampling. For example, when the width and height of the samples in the reconstructed image of the reference layer are the same as the width and height of the samples in the reconstructed image of the current layer; and when the offsets between the top-left, top-right, bottom-left, and bottom-right of the reference layer image and the top-left, top-right, bottom-left, and bottom-right of the current layer image are 0, the reconstructed image of the reference layer can be used as the inter-layer reference image of the current layer without resampling.
[0096] In addition, the reconstructed image of the reference layer used to derive the interlayer reference image can be an image belonging to the same AU as the current image to be encoded.
[0097] When performing inter-frame prediction for the current block by including inter-layer reference images in the reference image list, the positions of the inter-layer reference images within reference image lists L0 and L1 can differ. For example, in the case of reference image list L0, the inter-layer reference image can be located after a short-term reference image preceding the current image, while in the case of reference image list L1, the inter-layer reference image can be located at the end of the reference image list.
[0098] Here, reference image list L0 is a list of reference images used for inter-frame prediction of P slices or a list of reference images used as the first reference image list in inter-frame prediction of B slices. Reference image list L1 is a second reference image list used for inter-frame prediction of B slices.
[0099] Therefore, the reference image list L0 can be composed of short-term reference images preceding the current image, inter-layer reference images, short-term reference images following the current image, and long-term reference images. The reference image list L1 can be composed of short-term reference images following the current image, short-term reference images preceding the current image, long-term reference images, and inter-layer reference images.
[0100] In this context, a prediction slice (P-slice) is a slice on which intra-frame prediction is performed or inter-frame prediction is performed using up to one motion vector and a reference picture index per prediction block. A double prediction slice (B-slice) is a slice on which intra-frame prediction is performed or prediction is performed using up to two motion vectors and a reference picture index per prediction block. In this respect, an intra-frame slice (I-slice) is a slice on which intra-frame prediction is applied only.
[0101] Additionally, when performing inter-frame prediction for the current block based on a list of reference images including inter-layer reference images, the list of reference images may include multiple inter-layer reference images derived from multiple layers.
[0102] When the reference image list includes multiple inter-layer reference images, these reference images can be interleaved within reference image lists L0 and L1. For example, suppose the reference image list used for inter-frame prediction of the current block includes two inter-layer reference images (inter-layer reference images ILRP). i And interlayer reference images ILRP j In this case, ILRP is in the reference image list L0. i It can be located after a short reference image preceding the current image, and ILRP j It can be located at the end of the list. Furthermore, in the reference image list L1, ILRP... i It can be at the end of the list, and ILRP j It can be placed after a short reference image following the current image.
[0103] In this case, the reference image list L0 can include short-term reference images and interlayer reference images (ILRP) preceding the current image. i Short-term reference image, long-term reference image, and interlayer reference image (ILRP) following the current image. j The order in which they are composed. The reference image list L1 can be composed of short-term reference images following the current image, inter-layer reference images (ILRP). j The current image includes short-term reference images, long-term reference images, and interlayer reference images (ILRP). i The order of the components.
[0104] Furthermore, one of the two interlayer reference images can be an interlayer reference image derived from a resolution-dependent scalable layer, and the other can be an interlayer reference image derived from a layer providing a different view. In this case, for example, suppose ILRP... i It is an inter-layer reference image derived from layers providing different resolutions, and ILRP j This is an inter-layer reference image derived from layers that provide different views. Then, in the case of scalable video encoding that only supports scalability other than the view, the reference image list L0 can be composed of short-term reference images preceding the current image and inter-layer reference images ILRP. i The current image is composed of short-term reference images following it and long-term reference images in that order. On the other hand, the reference image list L1 can be composed of short-term reference images following the current image, short-term reference images preceding the current image, long-term reference images, and inter-layer reference images (ILRP). j The order of the components.
[0105] Furthermore, for inter-layer prediction, the information of the inter-layer reference image can consist of only sample values, only motion information (motion vectors), or both sample values and motion information. When the reference image index indicates an inter-layer reference image, the predictor 220-1 uses only the sample values of the inter-layer reference image, the motion information (motion vectors) of the inter-layer reference image, or both the sample values and motion information of the inter-layer reference image, based on the information received from the encoding device.
[0106] When using only sample values from the inter-layer reference image, predictor 220-1 can derive the predicted sample for the current block from samples of the block specified by motion vectors in the inter-layer reference image. Without considering scalable video coding of the view, the motion vectors in the inter-frame prediction (inter-layer prediction) using the inter-layer reference image can be set to fixed values (e.g., 0).
[0107] When using only motion information from inter-layer reference images, predictor 220-1 can use the motion vector specified in the inter-layer reference images as a motion vector predictor to derive the motion vector of the current block. Alternatively, predictor 220-1 can use the motion vector specified in the inter-layer reference images as the motion vector of the current block.
[0108] When using both samples from the inter-layer reference image and motion information, the predictor 220-1 can use samples related to the current block in the inter-layer reference image and motion information (motion vectors) specified in the inter-layer reference image to predict the current block.
[0109] When applying inter-layer prediction, the encoding device can send to the decoding device a reference index indicating an inter-layer reference image in the reference image list, and also send to the decoding device information specifying which information (sample information, motion information, or both) from the inter-layer reference images to use (i.e., information specifying the dependency type of the dependency related to the inter-layer prediction between the two layers).
[0110] Figure 4 This is a diagram illustrating, schematically, the configuration of a video / image decoding apparatus to which embodiments of the present disclosure may be applied.
[0111] Reference Figure 4 The decoding device 300 may include and be configured with an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above may be constituted by hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded image buffer (DPB) or may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0112] When the input includes a bitstream containing video / image information, the decoding device 300 can interact with... Figure 2The processing of video / image information in the encoding apparatus correspondingly reconstructs the image. For example, the decoding apparatus 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding apparatus 300 can perform decoding using processing units applied in the encoding apparatus. Therefore, the decoding processing unit can be, for example, an encoding unit, and the encoding unit can be segmented from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units can be derived from the encoding unit. The reconstructed image signal decoded and output by the decoding apparatus 300 can be reproduced by a reproduction apparatus.
[0113] Decoding device 300 can receive data in bitstream form from... Figure 2 The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The signaling / receiving information and / or syntax elements described later in this disclosure can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, context adaptive variable-length coding (CAVLC), or context adaptive arithmetic coding (CABAC) and outputs the quantized values of the syntax elements and transform coefficients of the residuals required for image reconstruction. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine the context model using information about the target syntax element, decoding information about the target block, or information about symbols / bins decoded in previous stages, and perform arithmetic decoding on the bin by predicting the occurrence probability of the bin based on the determined context model, generating a symbol corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using the information of the decoded symbol / bin for the context model of the next symbol / bin. The prediction-related information in the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantization transform coefficients and related parameter information) from which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320.
[0114] The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, filtering information from the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, the receiving unit (not shown) for receiving signals output from the encoding device can be further configured as an internal / external element of the decoding device 300, or the receiving unit can be a component of the entropy decoder 310. Furthermore, the decoding device according to this disclosure can be referred to as a video / image / picture decoding device, and the decoding device can be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0115] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order executed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0116] Inverse transformer 322 performs an inverse transformation on the transform coefficients to obtain the residual signal (residual block, residual sample array). In this disclosure, at least one of quantization / dequantization and / or transform / inverse transform may be omitted. When quantization / dequantization is omitted, the quantized transform coefficients may be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients may be referred to as coefficients or residual coefficients, or for consistency, they may still be referred to as transform coefficients.
[0117] In this disclosure, quantization transform coefficients and transform coefficients can be referred to as transform coefficients and scaling transform coefficients, respectively. In this case, residual information can include information about the transform coefficients, and this information about the transform coefficients can be transmitted as a signal using residual coding syntax. Transform coefficients can be derived based on the residual information (or the information about the transform coefficients), and scaling transform coefficients can be derived through the inverse transform (scaling) of the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaling transform coefficients. This can also be applied / expressed in other parts of this disclosure.
[0118] Predictor 330 can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from entropy decoder 310, and can determine the specific intra-frame / inter-frame prediction mode.
[0119] Predictor 330 can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can predict blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as games, for example, screen content coding (SCC). IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction because a reference block is derived in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying a palette mode, sample values within the frame can be signaled based on information about the palette table and palette index.
[0120] Intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples may be located near or far from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Intra-predictor 331 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0121] Inter-frame predictor 332 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.
[0122] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331). If the block to be processed has no residual (e.g., when a skip mode is applied), the prediction block can be used as the reconstruction block.
[0123] Adder 340 can be called a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, or it can be filtered and output as described below, or it can be used for inter-frame prediction of the next image.
[0124] In addition, Luminance Mapping and Chromaticity Scaling (LMCS) can be applied during image decoding.
[0125] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 360 (specifically, the DPB of memory 360). Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0126] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks from which motion information in the current image is derived (or decoded) and / or motion information of reconstructed blocks in the image. The stored motion information can be sent to inter-frame predictor 332 for use as motion information of spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and can transmit the reconstructed samples to intra-frame predictor 331.
[0127] In this disclosure, the embodiments described in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding apparatus 200 may be the same as or correspond to the filter 350, inter-frame predictor 332 and intra-frame predictor 331.
[0128] Figure 5 This is a diagram illustrating the configuration of a multi-layer-based video / image decoding apparatus to which embodiments of the present disclosure may be applied.
[0129] Figure 5 The decoding device may include Figure 4 The decoding device. In Figure 5In this context, the rearranger can be omitted or included in the dequantizer. The graph will be primarily described based on multi-level predictions. The remainder may include information about... Figure 4 The content of the description.
[0130] For ease of description, Figure 5 The example assumes a multi-layer structure consisting of two layers. However, embodiments of this disclosure are not limited to the specific example, and it should be noted that multi-layer structures to which embodiments of this disclosure are applied may include two or more layers.
[0131] Reference Figure 5 The decoding device 500 includes a decoder 500-1 for layer 1 and a decoder 500-0 for layer 0.
[0132] The decoder 500-1 of layer 1 may include an entropy decoder 510-1, a residual processor 520-1, a predictor 530-1, an adder 540-1, a filter 550-1, and a memory 560-1.
[0133] The decoder 500-0 of layer 0 may include an entropy decoder 510-0, a residual processor 520-0, a predictor 530-0, an adder 540-0, a filter 550-0, and a memory 560-0.
[0134] When a bitstream including image information is sent from the encoding device, the demultiplexer 505 can demultiplex the information of each layer and deliver the information to the decoding device for each layer.
[0135] Entropy decoders 510-1 and 510-0 can perform decoding based on the encoding method used in the encoding device. For example, when CABAC is used in the encoding device, entropy decoders 510-1 and 510-0 can also perform entropy decoding based on CABAC.
[0136] When the prediction mode used for the current block is the intra-prediction mode, predictors 530-1 and 530-0 can perform intra-prediction on the current block based on the neighboring reconstructed samples in the current image.
[0137] When the prediction mode used for the current block is inter-frame prediction mode, predictors 530-1 and 530-0 can perform inter-frame prediction on the current block based on information from at least one of the images included before or after the current image. Information received from the encoding device can be examined, and some or all of the motion information required for inter-frame prediction can be derived based on the examined information.
[0138] When the skip mode is applied as an inter-frame prediction mode, residuals may not be sent from the coding device, and prediction blocks may be used as reconstruction blocks.
[0139] Furthermore, the predictor 530-1 of layer 1 can perform inter-frame prediction or intra-frame prediction using only information from within layer 1, or it can perform inter-layer prediction using information from another layer (layer 0).
[0140] The information of the current layer predicted using information from different layers (i.e., predicted by inter-layer prediction) includes at least one of texture, motion information, cell information, and predetermined parameters (e.g., filtering parameters).
[0141] Furthermore, the information from different layers used for prediction of the current layer (i.e., for inter-layer prediction) may include at least one of texture, motion information, cell information, and predetermined parameters (e.g., filtering parameters).
[0142] In inter-layer prediction, the current block can be a block within the current image of the current layer (e.g., layer 1) and can be the target block to be decoded. The reference block can be a block within an image (reference image) of the layer (reference layer, e.g., layer 0) referenced for the prediction of the current block, belonging to the same access unit (AU) as the image (current image) to which the current block belongs, and can be a block corresponding to the current block.
[0143] An example of inter-layer prediction is inter-layer motion prediction, which uses motion information from a reference layer to predict the motion information of the current layer. Based on inter-layer motion prediction, the motion information of the current block can be predicted based on the motion information of the reference block. In other words, when deriving motion information based on inter-frame prediction modes (described later), motion information candidates can be derived using motion information from an inter-layer reference block rather than temporally neighboring blocks.
[0144] When applying interlayer motion prediction, predictor 530-1 can scale and use motion information from the reference block of the reference layer (i.e., the interlayer reference block).
[0145] In another example of inter-layer prediction, inter-layer texture prediction can use the texture of the reconstructed reference block as the predicted value for the current block. In this case, predictor 530-1 can scale the texture of the reference block by upsampling. Inter-layer texture prediction can be called inter-layer (reconstructed) sample prediction or simply inter-layer prediction.
[0146] In interlayer parameter prediction (another example of interlayer prediction), parameters derived from the reference layer can be reused in the current layer, or the parameters of the current layer can be derived based on the parameters used in the reference layer.
[0147] In interlayer residual prediction (another example of interlayer prediction), residual information from another layer can be used to predict the residual of the current layer, and prediction of the current block can be performed based on the predicted residual.
[0148] In inter-layer differential prediction (another example of inter-layer prediction), prediction of the current block can be performed using the difference between the images obtained by upsampling or downsampling the reconstructed image of the current layer and the reconstructed image of the reference layer.
[0149] In inter-layer syntax prediction (another example of inter-layer prediction), the syntax information of a reference layer can be used to predict or generate the texture of the current block. In this case, the syntax information of the referenced layer can include information about intra-frame prediction modes and motion information.
[0150] When predicting a specific block, multiple prediction methods using inter-layer prediction can be used across multiple layers.
[0151] Here, as examples of interlayer prediction, interlayer texture prediction, interlayer motion prediction, interlayer cell information prediction, interlayer parameter prediction, interlayer residual prediction, interlayer difference prediction, and interlayer syntax prediction have been described; however, the interlayer prediction applicable to this disclosure is not limited to the examples above.
[0152] For example, inter-layer prediction can be applied as an extension of inter-frame prediction for the current layer. In other words, inter-frame prediction for the current block can be performed by including a reference image derived from a reference layer in a reference image that can be referenced for inter-frame prediction of the current block.
[0153] When a reference image index received from the encoding device or derived from a neighboring block indicates an inter-layer reference image within the reference image list, the predictor 530-1 can use the inter-layer reference image to perform inter-layer prediction. For example, when the reference image index indicates an inter-layer reference image, the predictor 530-1 can derive sample values of the region specified by the motion vector in the inter-layer reference image as the prediction block for the current block.
[0154] In this case, inter-layer reference images can be included in the reference image list for the current block. Using inter-layer reference images, predictor 530-1 can perform inter-frame prediction for the current block.
[0155] Here, the inter-layer reference image can be a reference image constructed by sampling the reconstructed image of a reference layer to correspond to the current layer. Therefore, when the reconstructed image of the reference layer corresponds to the image of the current layer, the reconstructed image of the reference layer can be used as the inter-layer reference image without resampling. For example, when the width and height of the samples in the reconstructed image of the reference layer are the same as the width and height of the samples in the reconstructed image of the current layer; and when the offsets between the top-left, top-right, bottom-left, and bottom-right of the reference layer image and the top-left, top-right, bottom-left, and bottom-right of the current layer image are 0, the reconstructed image of the reference layer can be used as the inter-layer reference image of the current layer without resampling.
[0156] Furthermore, the reconstructed image of the reference layer used to derive the inter-layer reference image can be an image belonging to the same AU as the current image to be encoded. When inter-frame prediction of the current block is performed by including inter-layer reference images in the reference image list, the positions of the inter-layer reference images within reference image lists L0 and L1 can differ. For example, in the case of reference image list L0, the inter-layer reference image can be located after a short-term reference image preceding the current image, and in the case of reference image list L1, the inter-layer reference image can be located at the end of the reference image list.
[0157] Here, reference image list L0 is a list of reference images used for inter-frame prediction of P slices or a list of reference images used as the first reference image list in inter-frame prediction of B slices. Reference image list L1 is a second reference image list used for inter-frame prediction of B slices.
[0158] Therefore, the reference image list L0 can be composed of short-term reference images preceding the current image, inter-layer reference images, short-term reference images following the current image, and long-term reference images. The reference image list L1 can be composed of short-term reference images following the current image, short-term reference images preceding the current image, long-term reference images, and inter-layer reference images.
[0159] In this context, a prediction slice (P-slice) is a slice on which intra-frame prediction is performed or inter-frame prediction is performed using up to one motion vector and a reference picture index per prediction block. A double prediction slice (B-slice) is a slice on which intra-frame prediction is performed or prediction is performed using up to two motion vectors and a reference picture index per prediction block. In this respect, an intra-frame slice (I-slice) is a slice on which intra-frame prediction is applied only.
[0160] Additionally, when performing inter-frame prediction for the current block based on a list of reference images including inter-layer reference images, the list of reference images may include multiple inter-layer reference images derived from multiple layers.
[0161] When the reference image list includes multiple interlayer reference images, these interlayer reference images can be interleaved within reference image lists L0 and L1. For example, suppose there are two interlayer reference images (interlayer reference image ILRP). i And interlayer reference images ILRP j This is included in the list of reference images used for inter-frame prediction of the current block. In this case, ILRP is in the reference image list L0. i It can be located after a short reference image preceding the current image, and ILRP j It can be located at the end of the list. Furthermore, in the reference image list L1, ILRP... i It can be at the end of the list, and ILRP jIt can be placed after a short reference image following the current image.
[0162] In this case, the reference image list L0 can include short-term reference images and interlayer reference images (ILRP) preceding the current image. i Short-term reference image, long-term reference image, and interlayer reference image (ILRP) following the current image. j The order in which they are composed. The reference image list L1 can be composed of short-term reference images following the current image, inter-layer reference images (ILRP). j The current image includes short-term reference images, long-term reference images, and interlayer reference images (ILRP). i The order of the components.
[0163] Furthermore, one of the two interlayer reference images can be an interlayer reference image derived from a resolution-dependent scalable layer, and the other can be an interlayer reference image derived from a layer providing a different view. In this case, for example, suppose ILRP... i It is an inter-layer reference image derived from layers providing different resolutions, and ILRP j This is an inter-layer reference image derived from layers that provide different views. Then, in the case of scalable video encoding that only supports scalability other than the view, the reference image list L0 can be composed of short-term reference images preceding the current image and inter-layer reference images ILRP. i The current image is composed of short-term reference images following it and long-term reference images in that order. On the other hand, the reference image list L1 can be composed of short-term reference images following the current image, short-term reference images preceding the current image, long-term reference images, and inter-layer reference images (ILRP). j The order of the components.
[0164] Furthermore, for inter-layer prediction, the information of the inter-layer reference image can consist of only sample values, only motion information (motion vectors), or both sample values and motion information. When the reference image index indicates an inter-layer reference image, the predictor 530-1 uses only the sample values of the inter-layer reference image, the motion information (motion vectors) of the inter-layer reference image, or both the sample values and motion information of the inter-layer reference image, based on the information received from the encoding device.
[0165] When using only sample values from the inter-layer reference image, predictor 530-1 can derive the predicted sample for the current block from the sample of the block specified by the motion vector in the inter-layer reference image. Without considering scalable video coding of the view, the motion vector in the inter-frame prediction (inter-layer prediction) using the inter-layer reference image can be set to a fixed value (e.g., 0).
[0166] When using only motion information from inter-layer reference images, predictor 530-1 can use the motion vector specified in the inter-layer reference images as a motion vector predictor to derive the motion vector of the current block. Alternatively, predictor 530-1 can use the motion vector specified in the inter-layer reference images as the motion vector of the current block.
[0167] When using both samples from the inter-layer reference image and motion information, the predictor 530-1 can use samples from the inter-layer reference image that are relevant to the current block and motion information (motion vectors) specified in the inter-layer reference image to predict the current block.
[0168] The decoding device can receive reference indices from the encoding device indicating interlayer reference images within a list of reference images, and perform interlayer prediction based on the received reference indices. Furthermore, the decoding device can receive information from the encoding device specifying which information (sample information, motion information, or both) will be used from the interlayer reference images; that is, information specifying the type of dependency related to the interlayer prediction between the two layers.
[0169] Furthermore, in the video / image encoding according to this disclosure, the image processing unit can have a hierarchical structure. An image can be divided into one or more tiles, patches, slices, and / or tile groups. A slice may include one or more patches. A patch may include one or more CTU rows within a tile. A slice may include an integer number of patches of the image. A tile group may include one or more tiles. A tile may include one or more CTUs. A CTU may be divided into one or more CUs. A tile represents a rectangular area of a CTU within a specific tile column and a specific tile row in the image. A tile group may include an integer number of tiles based on a tile raster scan in the image. A slice header may carry information / parameters applicable to the corresponding slice (patches within a slice). In the case where the encoding / decoding device has a multi-core processor, the encoding / decoding process for tiles, slices, patches, and / or tile groups can be processed in parallel. In this disclosure, slices or tile groups can be used interchangeably. That is, a tile group header can be referred to as a slice header. Here, a slice can have one of the following slice types: intra-frame (I) slices, prediction (P) slices, and double prediction (B) slices. When predicting tiles in an I slice, inter-frame prediction may not be used, and intra-frame prediction may be used only. Of course, even in this case, notification can be performed by encoding the original sample values without prediction. For tiles in a P slice, either intra-frame or inter-frame prediction can be used, and if inter-frame prediction is used, single prediction may be used only. Furthermore, for tiles in a B slice, either intra-frame or inter-frame prediction can be used, and if inter-frame prediction is used, up to double prediction can be used to a maximum extent.
[0170] The encoding device may determine the size of tiles / tile groups, tiles, slices, and the maximum and minimum encoding units by taking into account encoding efficiency or parallel processing or by the characteristics of the video image (e.g., resolution), and may include information about the content or information that enables the acquisition of the content in the bitstream.
[0171] The decoding device can obtain information indicating the tiles / groups of tiles, blocks, and slices of the current image, as well as whether the CTUs within the tiles are divided into multiple coded units. Efficiency can be improved by ensuring that such information is obtained (transmitted) only under specific conditions.
[0172] Furthermore, as described above, an image can include multiple slices, and a slice can include a slice header and slice data. In this case, an image header can be further added to multiple slices within an image (slice header and slice dataset). The image header (image header syntax) can include information / parameters generally applicable to the image. The slice header (slice header syntax) can include information / parameters that can be applied commonly to the slice. APS (APS syntax) or PPS (PPS syntax) can include information / parameters that can be applied commonly to one or more slices or images. SPS (SPS syntax) can include information / parameters that can be applied commonly to one or more sequences. VPS (VPS syntax) can include information / parameters that can be applied commonly to multiple layers. DCI can include information / parameters related to decoding capabilities.
[0173] The advanced syntax (HLS) in this manual may include at least one of the following: APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, image header syntax, and slice header syntax.
[0174] Additionally, for example, information about the segmentation and configuration of tiles / tile groups / plots / slices can be configured in the encoding device based on high-level syntax, and then delivered (or transmitted) to the decoding device in bitstream format.
[0175] Figure 6 A schematic example of an image decoding process to which embodiments of the present disclosure may be applied is shown.
[0176] In image / video coding, images / video pictures can be encoded / decoded according to the decoding order. The image order, corresponding to the output order of the decoded images, can be configured differently from the decoding order. Furthermore, when performing inter-frame prediction based on the configured image order, both forward and backward prediction can be performed.
[0177] exist Figure 6 In the middle, the S600 can be composed of the above... Figure 4The entropy decoder 310 of the decoding apparatus described herein executes step S610, which can be executed by the predictor 330, step S620 by the residual processor 320, step S630 by the adder 340, and step S640 by the filter 350. Step S600 can include the information decoding process described herein, step S610 can include the inter-frame / intra-frame prediction process described herein, step S620 can include the residual processing process described herein, step S630 can include the block / picture reconstruction process described herein, and step S640 can include the loop filtering process described herein.
[0178] Reference Figure 6 As mentioned above Figure 4 As described, the image decoding process generally includes a process of obtaining image / video information from the bitstream (S600) through decoding, an image reconstruction process (S610 to S630), and a loop filtering process for reconstructing the image (S640). The image reconstruction process can be performed based on prediction samples and residual samples obtained by performing an inter-frame / intra-frame prediction process (S610) and a residual processing (or processing) process (S620, dequantization and inverse transform process of quantization transform coefficients). By performing a loop filtering process on the reconstructed image generated by the image reconstruction process, a modified reconstructed image can be generated, and the modified reconstructed image can be output as a decoded image and then stored in the decoded image buffer or memory 360 of the decoding device for use as a reference image during the inter-frame prediction process when the image is decoded in a later process. In some cases, the loop filtering process can be skipped. In this case, the reconstructed image can be output as a decoded image and then stored in the decoded image buffer or memory 360 of the decoding device for use as a reference image during the inter-frame prediction process when the image is decoded in a later process. As described above, the loop filtering process (S640) may include a deblocking filtering process, a sample adaptive offset (SAO) process, an adaptive loop filter (ALF) process, and / or a bilateral filter process, and may skip part or all of the loop filtering process. Furthermore, one or a portion of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and the bilateral filter process may be applied sequentially, or all of the deblocking filtering process, the sample adaptive offset (SAO) process, the adaptive loop filter (ALF) process, and the bilateral filter process may be applied sequentially. For example, the SAO process may be performed after the deblocking filtering process has been applied to the reconstructed image. Alternatively, for example, the ALF process may be performed after the deblocking filtering process has been applied to the reconstructed image. This can also be performed in the encoding apparatus.
[0179] Figure 7A schematic example of an image encoding process to which embodiments of the present disclosure can be applied is shown.
[0180] exist Figure 7 In China, the S700 can be composed of the above... Figure 2 The predictor 220 of the coding apparatus described herein executes S710, which may be executed by the residual processor 230, and S720 may be executed by the entropy encoder 240. S700 may include the inter-frame / intra-frame prediction process described herein, S710 may include the residual processing process described herein, and S720 may include the information encoding process described herein.
[0181] Reference Figure 7 As mentioned above Figure 2 As described above, the image encoding process typically includes encoding information for image reconstruction (e.g., prediction information, residual information, segmentation information, etc.) and outputting the encoded information as a bitstream, as well as generating a reconstructed image for the current image and optionally applying loop filtering to the reconstructed image. The encoding device can derive residual samples (which are modified) from the quantized transform coefficients using dequantizer 234 and inverse transformer 235, and then generate a reconstructed image based on the prediction samples as the output of S700 and the (modified) residual samples. The reconstructed image generated as described above can be the same as the reconstructed image generated in the decoding device. A modified reconstructed image can be generated by performing a loop filtering process on the reconstructed image, and the modified reconstructed image is then stored in the decoding image buffer or memory 270 of the decoding device. Furthermore, as in the decoding device, the modified reconstructed image can be used as a reference image during the inter-frame prediction process when encoding the image. As described above, in some cases, part or all of the loop filtering process can be skipped. When the loop filtering process is performed, the (loop) filtering related information (parameters) can be encoded in the entropy encoder 240 and then sent in the form of a bit stream, and the decoding device can perform the loop filtering process based on the filtering related information by using the same method as the encoding device.
[0182] By performing the aforementioned loop filtering process, noise (such as occlusion artifacts and ringing artifacts) occurring during image / moving image encoding can be reduced, and subjective / objective visual quality can be enhanced. Furthermore, by having both the encoding and decoding units perform the loop filtering process, they can obtain the same prediction results, increasing the reliability of image encoding and reducing the size (or amount) of data to be sent for image encoding.
[0183] As described above, the image reconstruction process can be performed in both the decoding and encoding apparatuses. Reconstructed blocks can be generated for each block unit based on intra-frame prediction / inter-frame prediction, and a reconstructed image including these blocks can be generated. When the current image / slice / patch group is an I-type image / slice / patch group, the blocks included in the current image / slice / patch group can be reconstructed based solely on intra-frame prediction. Furthermore, when the current image / slice / patch group is a P-type or B-type image / slice / patch group, the blocks included in the current image / slice / patch group can be reconstructed based on either intra-frame prediction or inter-frame prediction. In this case, inter-frame prediction can be applied to portions of the blocks within the current image / slice / patch group, and intra-frame prediction can be applied to the remaining blocks. The color components of the image can include both luma and chroma components. And, unless expressly limited (or constrained) in this specification, the methods and implementations described herein can be applied to both luma and chroma components.
[0184] Figure 8 An example of a layered structure for encoded images / videos is shown.
[0185] Reference Figure 8 The encoded image / video is divided into a VCL (Video Coding Layer) responsible for the image / video decoding process and itself, a subsystem for sending and storing encoded information, and a Network Abstraction Layer (NAL) that exists between the VCL and the subsystems and is responsible for network adaptation functions.
[0186] VCL can generate VCL data that includes compressed image data (slice data), or generate parameter sets or additional supplemental enhancement information (SEI) messages required for the image decoding process, such as picture parameter sets (picture parameter sets: PPS), sequence parameter sets (sequence parameter sets: SPS), and video parameter sets (video parameter sets: VPS).
[0187] In NAL, a NAL cell can be generated by adding header information (NAL cell header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc., generated in VCL. The NAL cell header may include NAL cell type information, which is specified based on the RBSP data included in the corresponding NAL cell.
[0188] As shown in the figure, NAL units can be divided into VCL NAL units and non-VCL NAL units based on the RBSP generated in VCL. VCL NAL units can refer to NAL units that include information about the image (slice data), while non-VCL NAL units can refer to NAL units that contain information required for decoding the image (parameter set or SEI message).
[0189] The VCL NAL units and non-VCL NAL units described above can be transmitted over a network by attaching header information according to the subsystem's data standard. For example, NAL units can be transformed into predetermined standard data formats (such as H.266 / VVC file format, Real-time Transport Protocol (RTP), Transport Stream (TS), etc.) and transmitted over various networks.
[0190] As described above, in a NAL cell, the NAL cell type can be specified according to the RBSP data structure included in the corresponding NAL cell, and information about the NAL cell type can be stored in the NAL cell header and signaled.
[0191] For example, NAL cells can be broadly classified into VCL NAL cell type and non-VCL NAL cell type depending on whether the NAL cell includes information about the image (slice data). VCL NAL cell type can be classified according to the nature and type of the image included in the VCL NAL cell, while non-VCL NAL cell type can be classified according to the type of parameter set.
[0192] The following is an example of a NAL cell type specified based on the type of the parameter set included in a non-VCL NAL cell type.
[0193] -DCI (Decoding Capability Information) NAL Unit: Includes the type of NAL unit for DCI.
[0194] -VPS (Video Parameter Set) NAL Unit: Includes the type of NAL unit for the VPS.
[0195] -SPS (Sequence Parameter Set) NAL Unit: The type of NAL unit that includes SPS.
[0196] -PPS (Image Parameter Set) NAL Unit: The type of NAL unit including PPS.
[0197] -APS (Adaptive Parameter Set) NAL Unit: The type of NAL unit including APS.
[0198] -PH (Picture Header) NAL Unit: Types of NAL units including PH.
[0199] The NAL unit type described above has syntax information specific to the NAL unit type, which can be stored in the NAL unit header and signaled. For example, this syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.
[0200] Furthermore, as described above, an image can include multiple slices, and a slice can include a slice header and slice data. In this case, an image header can be further added to multiple slices within an image (slice header and slice dataset). The image header (image header syntax) can include information / parameters common to the image. The slice header (slice header syntax) can include information / parameters common to the slice. APS (APS syntax) or PPS (PPS syntax) can include information / parameters common to one or more slices or images. SPS (SPS syntax) can include information / parameters common to one or more sequences. VPS (VPS syntax) can include information / parameters common to multiple layers. DCI (DCI syntax) can include information / parameters related to decoding capabilities.
[0201] In this specification, High-Level Syntax (HLS) may include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DCI syntax, image header syntax, and slice header syntax. Furthermore, in this specification, Low-Level Syntax (LLS) may include, for example, slice data syntax, CTU syntax, and transform unit syntax.
[0202] In this specification, the image / video information encoded by the encoding device to the decoding device and then signaled in bitstream format may include not only information related to intra-frame segmentation, intra / inter-frame prediction information, residual information, loop filtering information, etc., but also slice header information, image header information, APS information, PPS information, SPS information, VPS information, and / or DCI information. Furthermore, the image / video information may further include general constraint information and / or NAL unit header information.
[0203] Furthermore, as described above, the video / image information in this specification may include higher-level signaling, and video / image encoding methods may be performed based on the video / image information.
[0204] The encoded image may include one or more slices. Parameters describing the encoded image can be signaled in the image header, and parameters describing the slices can be signaled in the slice header. The image header (PH) is carried within its own NAL cell type. The slice header exists at the beginning of the NAL cell, which includes the slice payload (slice data).
[0205] Additionally, the encoded image may include slices of another NAL unit type. The image should reference the image parameter set that includes the `mixed_nalu_type_in_pic_flag` syntax element.
[0206] If the value of mixed_nalu_type_in_pic_flag is equal to 1, this indicates that each picture in the reference PPS has one or more VCL NAL units, the VCL NAL units do not have the same value as nal_unit_type, and the picture is not an IRAP picture. Furthermore, if the value of mixed_nalu_type_in_pic_flag is equal to 0, this indicates that each picture in the reference PPS has one or more VCL NAL units, and the VCL NAL units of each picture in the reference PPS have the same value as nal_unit_type.
[0207] When the value of no_mixed_nalu_type_in_pic_constraint_flag is equal to 1, the value of mixed_nalu_type_in_pic_flag is equal to 0.
[0208] For each slice with a nal_unit_type value of nalUnitTypeA, in the range from IDR_W_RADL to CRA_NUT, in the image picA that includes one or more slices with a nal_unit_type value of another (i.e., the mixed_nalu_type_in_pic_flag value of image picA is equal to 1), the following holds true.
[0209] - The slice should belong to the subpic A with a corresponding subpic_treat_as_pic_flag[i] value equal to 1.
[0210] - Slices should not be subpicks of picA, which include VCL NAL units with a nal_unit_type different from nalUnitTypeA.
[0211] - For all PUs in the following PUs within the Coding Layer Video Sequence (CLVS), the RefPicList[0] or RefPicList[1] of the slice within subpicA should not include the picture that is in the valid entry preceding picA in decoding order.
[0212] Additionally, the following content applies to the VCL NAL unit of a specific image.
[0213] - If the value of mixed_nalu_type_in_pic_flag is equal to 0, then the nal_unit_type value should be the same as all encoded slice NAL units within the picture. The picture or PU has the same NAL unit type as the encoded slice NAL units of that picture or PU.
[0214] - Otherwise (if the value of mixed_nalu_type_in_pic_flag is equal to 1), one or more VCL NAL units should all have a specific value of nal_unit_type in the range from IDR_W_RADL to CRA_NUT, and all other VCLNAL units should have a specific value of nal_unit_type in the range from TRAIL_NUT to RSV_VCL_6.
[0215] The signaling of the multi-layered information within the video parameter set will be described in detail below.
[0216] The following shows the available layer set (Output Layer Set (OLS)), profile, level and grade (PTL), information about the OLS, DPB information, HRD information, etc. that can be decoded for multi-layer bitstreams.
[0217] [Table 1]
[0218]
[0219]
[0220] The VPS RBSP in Table 1 should be available for the decoding process before being referenced, and the VPS RBSP should include at least one AU with a temporary identifier (temporalId) equal to 0 or provided by an external means.
[0221] All VPSNAL units within a encoded video sequence (CVS) that each have a specific value for vps_video_parameter_set_id should have the same content.
[0222] The `vps_video_parameter_set_id` provides an identifier specific to the VPS, allowing it to be referenced by another syntax element. The `vps_video_parameter_set_id` value should be greater than 0.
[0223] vps_max_layers_minus1+1 indicates the maximum number of layers allowed within each CVS of the reference VPS.
[0224] Incrementing `vps_max_sublayer_minus1` by 1 indicates the maximum number of time-based sublayers that can exist within the layers of the reference VPS across its individual CVS. The value of `vps_max_sublayer_minus1` should be in the range of 0 to 6.
[0225] When the `vps_all_layers_same_num_sublayer_flag` value is equal to 1, it indicates that the number of time sublayers is the same for all layers within each CVS of the reference VPS. Conversely, when the `vps_all_layers_same_num_sublayer_flag` value is equal to 0, it indicates that layers within each CVS of the reference VPS may or may not have the same number of time sublayers. When the `vps_all_layers_same_num_sublayer_flag` syntax element is not present in the VPS syntax, the `vps_all_layers_same_num_sublayer_flag` value is inferred to be equal to 1 (or deduced to be 1).
[0226] If the value of `vps_all_independent_layers_flag` is equal to 1, this indicates that all layers within the CVS are encoded independently without using inter-layer prediction. If the value of `vps_all_independent_layers_flag` is equal to 1, this indicates that one or more layers within the CVS can use inter-layer prediction. When the `vps_all_independent_layers_flag` syntax element is not present in the VPS syntax, the value of `vps_all_independent_layers_flag` is inferred to be equal to 1.
[0227] vps_layer_id[i] indicates the nuh_layer_id value of the i-th layer. For two non-negative integer values between m and n, where m is less than n, the value of vps_layer_id[m] should be less than the value of vps_layer_id[n].
[0228] When the value of `vps_independent_layer_flag[i]` is equal to 1, this indicates that the layer with index `i` does not use inter-layer prediction. And when the value of `vps_independent_layer_flag[i]` is equal to 0, this indicates that the layer with index `i` can use inter-layer prediction, and the syntax element `vps_direct_ref_layer_flag[i][j]` exists within the VPS, where `j` ranges from 0 to `i-1` (inclusive). When the `vps_independent_layer_flag` syntax element does not exist in the VPS syntax, the value of `vps_independent_layer_flag` is inferred to be equal to 1.
[0229] When the value of vps_max_tid_ref_present_flag[i] is equal to 1, it indicates that the syntax element vps_max_tid_il_ref_pics_plus1[i][j] exists. And when the value of vps_max_tid_ref_present_flag[i] is equal to 0, it indicates that the syntax element vps_max_tid_il_ref_pics_plus1[i][j] does not exist.
[0230] If the value of `vps_direct_ref_layer_flag[i][j]` is 0, this indicates that the layer with index `j` is not a direct reference layer to the layer with index `i`. And, if the value of `vps_direct_ref_layer_flag[i][j]` is 1, this indicates that the layer with index `j` is a direct reference layer to the layer with index `i`. If `vps_direct_ref_layer_flag[i][j]` does not exist for `i` and `j` in the range from 0 to `vps_max_layers_minus1`, then the value of `vps_direct_ref_layer_flag[i][j]` is inferred to be 0. When the value of `vps_independent_layer_flag[i]` is 0, there should exist one or more values for `j` in the range from 0 to `i-1` (inclusive) such that `vps_direct_ref_layer_flag[i][j]` is allowed to be 1.
[0231] As described below, we derive (or infer) NumDirectRefLayers[i], DirectRefLayerIdx[i][d], NumRefLayers[i], RefLayerIdx[i][r], and LayerUsedAsRefLayerFlag[j].
[0232] If the value of `vps_max_tid_il_ref_pics_plus1[i][j]` is equal to 0, this indicates that the j-th layer image, which is neither an IRAP image nor a GDR image with a `ph_recovery_poc_cnt` value of 0, will not be used as the ILRP for image decoding of the i-th layer image. When the value of `vps_max_tid_il_ref_pics_plus1[i][j]` is greater than 0, this indicates that when decoding the i-th layer image, the j-th layer image with a TemporalId greater than `vps_max_tid_il_ref_pics_plus1[i][j] - 1` will not be used as the ILRP. When the `vps_max_tid_il_ref_pics_plus1` syntax element is not present in the VPS, the value of `vps_max_tid_il_ref_pics_plus1[i][j]` is inferred to be equal to `vps_max_sublayer_minus1 + 1`.
[0233] If `vps_each_layer_is_an_ols_flag` equals 1, this indicates that each OLS comprises only one layer, and the layer within each layer of the reference VPS's CVS is the OLS itself, which is the unique output layer. If `vps_each_layer_is_an_ols_flag` equals 0, this indicates that one or more OLSs comprise two or more layers. If `vps_max_layers_minus1` equals 0, then the value of `vps_each_layer_is_an_ols_flag` can be inferred to be equal to 1. Otherwise, if `vps_all_independent_layers_flag` equals 0, then the value of `vps_each_layer_is_an_ols_flag` can be inferred to be equal to 0.
[0234] If the value of vps_ols_mode_idc is equal to 0, this indicates that the total number of OLS indicated by the VPS is equal to vps_max_layers_minus1+1, and this also indicates that the i-th OLS includes layers with layer indices ranging from 0 to i, and that for each OLS only the highest layer in that OLS is an output layer.
[0235] If the value of vps_ols_mode_idc is equal to 1, this indicates that the total number of OLS indicated by the VPS is equal to vps_max_layers_minus1+1, and this also indicates that the i-th OLS includes layers with layer indices ranging from 0 to i, and that all layers of the OLS are output layers for each OLS.
[0236] If the value of vps_ols_mode_idc is equal to 2, this indicates that the total number of OLS indicated by the VPS is explicitly signaled, the output layer is explicitly signaled for each OLS, and the other layers are direct or indirect reference layers to the output layer of the OLS.
[0237] The value of vps_ols_mode_idc should be in the range of 0 to 2.
[0238] If the value of vps_all_independent_layers_flag is equal to 1, and if the value of vps_each_layer_is_an_ols_flag is equal to 0, then the value of vps_ols_mode_idc is inferred to be equal to 2.
[0239] vps_num_output_layer_sets_minus1+1 indicates the total number of OLS as indicated by the VPS when the value of vps_ols_mode_idc is equal to 2.
[0240] The variable TotalNumOlss is derived (or inferred) as shown below, which indicates the total number of OLS indicated by the VPS.
[0241] [Table 2]
[0242]
[0243] When the value of vps_ols_output_layer_flag[i][j] is equal to 1, this indicates that when the value of vps_ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is the output layer of the i-th OLS. When the value of vps_ols_output_layer_flag[i][j] is equal to 0, this indicates that when the value of vps_ols_mode_idc is equal to 2, the layer with nuh_layer_id equal to vps_layer_id[j] is not the output layer of the i-th OLS.
[0244] The following results in the following variables: NumOutputLayersInOls[i] indicating the number of output layers in the i-th OLS, NumSubLayersInLayerInLayerInOLS[i][j] indicating the number of sublayers of the j-th layer in the i-th OLS, OutputLayerIdInOls[i][j] indicating the nuh_layer_id value of the j-th output layer in the i-th OLS, and LayerUsedAsOutputLayerFlag[k] indicating whether the k-th layer is used as an output layer in at least one OLS.
[0245] [Table 3]
[0246]
[0247]
[0248] For each value of i in the range from 0 to vps_max_layers_minus1, the values of LayerUsedAsRefLayerFlag[i] and LayerUsedAsOutputLayerFlag[i] are both not equal to 0. That is, there should be no layer that is not the output layer of at least one OLS or a layer that is not a direct reference layer of another layer. There should be one or more layers that serve as the output layer of each OLS. In other words, for each value of i in the range from 0 to TotalNumOlsl-1 (inclusive), the value of NumOutputLayersInOls[i] should be equal to or greater than 1.
[0249] The following results in the following variables: NumLayersInOls[i] indicating the number of layers in the i-th OLS; LayerIdInOls[i][j] indicating the nuh_layer_id value of the j-th layer in the i-th OLS; NumMultiLayerOlss indicating the number of multi-layer OLS (i.e., OLS with two or more layers); and MultiLayerOlsIdx[i] indicating the index used for the multi-layer OLS list or the i-th OLS when NumLayersInOls[i] is greater than 0.
[0250] [Table 4]
[0251]
[0252] The 0th OLS includes only the lowest layer (i.e., the layer with nuh_layer_id equal to vps_layer_id[0]), and in the case of the 0th OLS, the output includes only the layer.
[0253] As shown below, the variable OlsLayerIdx[i][j] is derived (or inferred), which indicates the OLS layer index of the layer with nuh_layer_id equal to LayerIdInOls[i][j].
[0254] [Table 5]
[0255]
[0256] The lowest layer of each OLS should be an independent layer. That is, for each value of i in the range from 0 to TotalNumOlss-1 (inclusive), the value of vps_independent_layer_flag[GeneralLayerIdx[LayerIdInOls[i][0]]] should be equal to 1.
[0257] Each layer should be included in one or more OLS indicated by the VPS. That is, there should be one or more pairs of i and j values such that for each layer with a specific nuh_layer_id nuhLayerId that is equal to one of the values of vps_layer_id[k], the LayerIdInOls[i][j] value can be equal to nuhLayerId, and k is in the range from 0 to vps_max_layers_minus1. In this document, i is in the range from 0 to TotalNumOlss-1, and j is in the range from 0 to NumLayersInOls[i]-1 (inclusive).
[0258] vps_num_ptls_minus1+1 indicates the number of profile_tier_level() syntax structures within the VPS. The value of vps_num_ptls_minus1 should be less than TotalNumOlss.
[0259] If the value of vps_pt_present_flag[i] is equal to 1, it indicates that the profile, hierarchy, and general constraint information exist in the i-th profile_tier_level() syntax structure within the VPS. If the value of vps_pt_present_flag[i] is equal to 0, it indicates that the profile, hierarchy, and general constraint information do not exist in the i-th profile_tier_level() syntax structure within the VPS. The value of vps_pt_present_flag[0] is inferred to be equal to 1. If vps_pt_present_flag[i] is equal to 0, the profile, hierarchy, and general constraint information of the i-th profile_tier_level() syntax structure within the VPS is derived (or inferred) to be the same as that of the (i-1)-th profile_tier_level() syntax structure within the VPS.
[0260] `vps_ptl_max_temporal_id[i]` indicates the TemporalId of the highest sublayer representation, where level information exists in the `i`th `profile_tier_level()` syntax structure. The value of `vps_ptl_max_temporal_id[i]` should be in the range of 0 to `vps_max_sublayer_minus1`. When the `vps_ptl_max_temporal_id` syntax element does not exist in the VPS, the value of `vps_ptl_max_temporal_id[i]` is inferred to be equal to `vps_max_sublayer_minus1`.
[0261] vps_ptl_alignment_zero_bit should be equal to 0.
[0262] `vps_ols_ptl_idx[i]` specifies the index of the `profile_tier_level()` syntax structure applied to the `i`th OLS within the list of `profile_tier_level()` syntax structures in the VPS. When the `vps_ols_ptl_idx` syntax element exists in the VPS, the value of `vps_ols_ptl_idx[i]` should be in the range from 0 to `vps_num_ptls_minus1` (inclusive).
[0263] When the vps_ols_ptl_idx syntax element does not exist in the VPS, the value of vps_ols_ptl_idx[i] is derived (or inferred) as described below.
[0264] - If the value of vps_num_ptls_minus1 is equal to 0, then the value of vps_ols_ptl_idx[i] is inferred to be equal to 0.
[0265] Otherwise (vps_num_ptls_minus1 is greater than 0, and vps_num_ptls_minus1+1 equals TotalNumOlss), then the value of vps_ols_ptl_idx[i] is inferred to be equal to i.
[0266] If NumLayersInOls[i] equals 1, then the profile_tier_level() syntax structure applied to the i-th OLS also exists in the SPS referenced by the layer within the i-th OLS. When NumLayersInOls[i] equals 1, the profile_tier_level() syntax structure signaled in the VPS and the profile_tier_level() syntax structure signaled in the SPS for the i-th OLS should be the same, according to bitstream compliance requirements.
[0267] Each profile_tier_level() syntax structure within a VPS should be referenced by at least one vps_ols_ptl_idx[i] value, where i ranges from 0 to TotalNumOlss-1 (inclusive of endpoints).
[0268] (When present) vps_num_dpb_params_minus1+1 indicates the number of dpb_parameters() syntax structures within the VPS. The value of vps_num_dpb_params_minus1 should be in the range from 0 to NumMultiLayerOlss-1 (inclusive).
[0269] As shown below, the variable VpsNumDpbParams is derived (or inferred), which indicates the number of dpb_parameters() syntax structures within the VPS.
[0270] [Table 6]
[0271]
[0272] The `vps_sublayer_dpb_params_present_flag` flag controls whether the `dpb_parameters()` syntax element within the VPS contains the `max_dec_pic_buffering_minus1[]`, `max_num_reorder_pics[]`, and `max_latency_increase_plus1[]` syntax elements. If these elements are not present, the value of `vps_sub_dpb_params_info_present_flag` is inferred to be 0.
[0273] `vps_dpb_max_temporal_id[i]` indicates the TemporalId of the highest sublayer representation, where DPB parameters can exist in the i-th `dpb_parameters()` syntax structure within the VPS. The value of `vps_dpb_max_temporal_id[i]` should be in the range of 0 to `vps_max_sublayer_minus1`. When it does not exist, the value of `vps_dpb_max_temporal_id[i]` is inferred to be equal to `vps_max_sublayer_minus1`.
[0274] vps_ols_dpb_pic_width[i] indicates the width of the individual image storage buffers used in the i-th multilayer OLS, in units of luminance samples.
[0275] vps_ols_dpb_pic_height[i] indicates the height of each image storage buffer used in the i-th multi-layer OLS, in units of luminance samples.
[0276] vps_ols_dpb_chroma_format[i] indicates the maximum allowed value of sps_chroma_format_idc for all SPS references within the CVS used for the i-th multi-level OLS.
[0277] vps_ols_dpb_bitdepth_minus8[i] indicates the maximum allowed value of sps_bit_depth_minus8 for all SPS referenced by the CLVS within the CVS used for the i-th multilayer OLS.
[0278] In order to decode the i-th multi-layer OLS, the decoding device can securely allocate memory to the DPB according to the syntax element values of the syntax elements vps_ols_dpb_pic_width[i], vps_ols_dpb_pic_height[i], vps_ols_dpb_chroma_format[i], and vps_ols_dpb_bitdepth_ols_dpb_bitdepth_ols_dpb_bitdepth.
[0279] `vps_ols_dpb_params_idx[i]` indicates the index of the `dpb_parameters()` syntax structure applied to the `i`-th level OLS for the list of `dpb_parameters()` syntax structures within the VPS. When present, the value of `vps_ols_dpb_params_idx[i]` should be in the range from 0 to `VpsNumDpbParams-1` (inclusive).
[0280] If it does not exist, then infer vps_ols_dpb_params_idx[i] as described below.
[0281] - When VpsNumDpbParams equals 1, the value of vps_ols_dpb_params_idx[i] equals 0.
[0282] Otherwise (VpsNumDpbParams is greater than 1 and equal to NumMultiLayerOlss), the value of vps_ols_dpb_params_idx[i] is inferred to be equal to i.
[0283] In the case of a single-layer OLS, the applicable dpb_parameters() syntax structure exists in the SPS referenced by the layer within the OLS.
[0284] Each dpb_parameters() syntax structure within a VPS should be referenced by at least one vps_ols_dpb_params_idx[i] value, where i ranges from 0 to NumMultiLayerOlss-1 (inclusive of endpoints).
[0285] If the value of `vps_general_hrd_params_present_flag` is equal to 1, it indicates that the VPS includes the `general_hrd_parameters()` syntax structure and other HRD parameters. If the value of `vps_general_hrd_params_present_flag` is equal to 0, it indicates that the VPS does not include the `general_hrd_parameters()` syntax structure or other HRD parameters. When it does not exist, the value of `vps_general_hrd_params_present_flag` is inferred to be equal to 0.
[0286] When the value of NumLayersInOls[i] is equal to 1, the general_hrd_parameters() syntax structure and the ols_hrd_parameters() syntax structure applied to the i-th OLS exist in the SPS referenced by the layer within the i-th OLS.
[0287] If the value of `vps_sublayer_cpb_params_present_flag` is equal to 1, this indicates that the i-th `ols_hrd_parameters()` syntax structure within the VPS includes HRD parameters for the sublayer representation with a TemporalId ranging from 0 to `vps_hrd_max_tid[i]` (inclusive). And, if the value of `vps_sublayer_cpb_params_present_flag` is equal to 0, this indicates that the i-th `ols_hrd_parameters()` syntax structure within the VPS includes only HRD parameters for the sublayer representation with a TemporalId equal to `vps_hrd_max_tid[i]`. If the value of `vps_max_sublayer_minus1` is equal to 0, then the value of `vps_sublayer_cpb_params_present_flag` is inferred to be equal to 0.
[0288] When the value of `vps_sublayer_cpb_params_present_flag` is equal to 0, the HRD parameters for a sublayer representation with a TemporalId ranging from 0 to `vps_hrd_max_tid[i]-1` (inclusive) are inferred to be equal to the sublayer representation with a TemporalId equal to `vps_hrd_max_tid[i]`. This includes the HRD parameters from the `fixed_pic_rate_general_flag[i]` syntax element to the `sublayer_hrd_parameters(i)` syntax structure, which is immediately followed by the `if(general_vcl_hrd_params_present_flag)` condition within the `ols_hrd_parameters` syntax structure.
[0289] Incrementing `vps_num_ols_hrd_params_minus1` by 1 indicates the number of `ols_hrd_parameters()` syntax structures present within the VPS when the value of `vps_general_hrd_params_present_flag` is equal to 1. The value of `vps_num_ols_hrd_params_minus1` should be in the range from 0 to `NumMultiLayerOlss-1` (inclusive).
[0290] `vps_hrd_max_tid[i]` indicates the TemporalId of the highest sublayer representation that has HRD parameters included in the `i`th `ols_hrd_parameters()` syntax structure. The value of `vps_hrd_max_tid[i]` should be in the range of 0 to `vps_max_sublayer_minus1`. When it does not exist, the value of `vps_hrd_max_tid[i]` is inferred to be equal to `vps_max_sublayer_minus1`.
[0291] `vps_ols_hrd_idx[i]` indicates the index of the `ols_hrd_parameters()` syntax structure applied to the `i`-th level of the OLS, for the list of `ols_hrd_parameters()` syntax structures within the VPS. The value of `vps_ols_hrd_idx[i]` should be in the range of 0 to `vps_num_ols_hrd_params_minus1`.
[0292] If it does not exist, then infer (or deduce) vps_ols_hrd_idx[i] as follows.
[0293] - If the value of vps_num_ols_hrd_params_minus1 is equal to 0, then the value of vps_ols_hrd_idx[i] is inferred to be equal to 0.
[0294] Otherwise (vps_num_ols_hrd_params_minus1+1 is greater than 1 and equal to NumMultiLayerOlss), the value of vps_ols_hrd_idx[i] is inferred to be equal to i.
[0295] In the case of a single-layer OLS, the applicable ols_hrd_parameters() syntax structure exists in the SPS referenced by the layer within the OLS.
[0296] Each ols_hrd_parameters() syntax structure within a VPS should be referenced by at least one vps_ols_hrd_idx[i] value, where i ranges from 0 to NumMultiLayerOlss-1 (inclusive of endpoints).
[0297] If the value of vps_extension_flag is 0, it indicates that the vps_extension_data_flag syntax element does not exist in the VPS RBSP syntax structure. Conversely, if the value of vps_extension_flag is 1, it indicates that the vps_extension_data_flag syntax element exists in the VPS RBSP syntax structure.
[0298] vps_extension_data_flag can have a random value.
[0299] In the case of a single-layer bitstream, the presence of a VPS is optional. If no VPS exists, the value of the sps_video_parameter_set_id syntax element is equal to 0, and the parameter value is inferred as described below.
[0300] If the value of sps_video_parameter_set_id is equal to 0, the following applies.
[0301] - The SPS does not reference the VPS, and the VPS is not referenced when decoding individual CLVSs that reference the SPS.
[0302] The value of -vps_max_layers_minus1 is inferred to be equal to 0.
[0303] The value of -vps_max_sublayer_minus1 is inferred to be equal to 6.
[0304] - The CVS should consist of only one layer (i.e., all VCL NAL units within the CVS should have the same nuh_layer_id value).
[0305] The value of -GeneralLayerIdx[nuh_layer_id] is inferred to be equal to 0.
[0306] The value of -vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1.
[0307] When performing video / image encoding, parameter sets (PPS, SPS, VPS, etc.) can be shared between layers. That is, a VCL NAL unit within a specific layer can reference the parameter set of another layer. When using parameter set sharing functions, the following constraints apply.
[0308] spsLayerId is set to the nuh_layer_id of a specific (or particular) SPS NAL cell, and vclLayerId is set to the nuh_layer_id value of a specific VCL NAL cell. A specific VCL NAL cell does not reference a specific SPS NAL cell except when spsLayerId is less than or equal to vclLayerId, and when not all OLSs include layers with the same nuh_layer_id as spsLayerId, as indicated by a VPS that includes layers with the same nuh_layer_id as vclLayerId.
[0309] The ppsLayerId is set to the nuh_layer_id of a specific PPS NAL cell, and the vclLayerId is set to the nuh_layer_id value of a specific VCL NAL cell. A specific VCL NAL cell should not reference a specific PPS NAL cell except when the ppsLayerId is less than or equal to the vclLayerId, and when not all OLSs include layers with the same nuh_layer_id as the ppsLayerId, as indicated by a VPS that includes layers with the same nuh_layer_id as the vclLayerId.
[0310] The apsLayerId is set to the nuh_layer_id of a specific APS NAL cell, and the vclLayerId is set to the nuh_layer_id value of a specific VCL NAL cell. A specific VCL NAL cell should not reference a specific APS NAL cell, except when the apsLayerId is less than or equal to the vclLayerId, and when not all OLSs include layers with the same nuh_layer_id as the apsLayerId, as indicated by a VPS that includes layers with the same nuh_layer_id as the vclLayerId.
[0311] Furthermore, when the VPS does not exist in the CVS (i.e., when the bitstream is a single-layer bitstream), all VCL NAL units within the bitstream should have the same nuh_layer_id, and all VCL NAL units should reference the parameter set within the same layer. However, since such a constraint is not included in the aforementioned constraints, it introduces additional complexity in decoding devices designed to handle only single-layer bitstreams.
[0312] Furthermore, OLS information is not established when the VPS does not exist in the CVS. Therefore, some parameters required for decoding, such as TotalNumOlss, NumLayersInOls[], NumOutputLayersInOls[], etc., must not be derived (or inferred) or initialized. This causes problems in the operation of the decoding device.
[0313] The following figures are illustrated to provide a detailed description of this specification. Detailed terminology for the apparatus (or device) or for the signals / information specified in the figures is merely exemplary. Therefore, the technical features of this specification are not limited to the detailed terminology used in the following figures.
[0314] This manual provides the following methods to solve the above problems. Each method can be applied independently or in combination.
[0315] For example, when a VPS does not exist for CVS (i.e., when the value of sps_video_parameter_set_id is equal to 0), the following constraints can be applied.
[0316] a) The layer identifier (nuh_layer_id) of all VCL NAL cells within the CVS of the reference SPS has the same value as the layer identifier (nuh_layer_id) of the SPS.
[0317] b) The layer identifier (nuh_layer_id) value of all VCL NAL cells is the same as the layer identifier (nuh_layer_id) value of the parameter set referenced by all VCL NAL cells.
[0318] In this document, the parameter set referenced by the VCL NAL unit includes a set of parameters used for decoding the video / image information disclosed in this specification. For example, the parameter set may include APS, PPS, SPS, VPS, etc.
[0319] Alternatively, the aforementioned constraints can be expressed as follows.
[0320] a) All VCL NAL cells within CVS have the same layer identifier (nuh_layer_id) value.
[0321] b) The layer identifier (nuh_layer_id) of the parameter set referenced by each VCL NAL unit within CVS is the same as the layer identifier (nuh_layer_id) of the VCLNAL unit.
[0322] Alternatively, the aforementioned constraints can be expressed as follows.
[0323] a) All VCL NAL cells within CVS and the parameter set referenced by VCL NAL cells have the same layer identifier (nuh_layer_id) value.
[0324] Alternatively, when the VCL NAL unit reference has a parameter with a different layer identifier (nuh_layer_id) than the layer identifier (nuh_layer_id) of the VCL NAL unit, the aforementioned constraint can be expressed (or represented) such that the value of sps_video_parameter_set_id is equal to 0.
[0325] Additionally, for example, when the VPS does not exist in the CVS (i.e., when the sps_video_parameter_set_id value is equal to 0), only one output layer set exists within the CVS, and the output layer set includes only one layer within the CVS, and that layer can be derived (or inferred) as the output layer of the output layer set.
[0326] In the absence of a VPS for CVS, the values of the following parameters can be derived (or inferred) as described below.
[0327] a) TotalNumOlss is inferred to be equal to 1.
[0328] b) NumLayersInOls[0] is inferred to be equal to 1.
[0329] c) NumOutputLayersInOls[0] is inferred to be equal to 1.
[0330] d) OutputLayerIdInOls[0][0] is inferred to be equal to nuh_layer_id of SPS.
[0331] According to the implementation method, Table 7 shown below can be applied to the sps_video_parameter_set_id syntax element.
[0332] [Table 7]
[0333]
[0334] Referring to Table 7, when the value of the sps_video_parameter_set_id syntax element is equal to 0 or greater (i.e., when the bitstream including the sps_video_parameter_set_id syntax element is a multi-level bitstream), the sps_video_parameter_set_id syntax element indicates the value of the vps_video_parameter_set_id syntax element of the VPS referenced by the SPS within the corresponding bitstream.
[0335] When the `sps_video_parameter_set_id` syntax element is equal to 0 (i.e., when the bitstream including the `sps_video_parameter_set_id` syntax element is a single-layer bitstream), the corresponding bitstream may not include the VPS. Therefore, the SPS within the corresponding bitstream does not reference the VPS, and the VPS is not referenced when decoding each CLVS that references the SPS.
[0336] Furthermore, the syntax element (vps_max_layers_minus1) indicating the maximum number of layers within CVS is derived (or deduced) to be equal to 0, and the syntax element (vps_max_sublayer_minus1) indicating the number of time sublayers that can exist in CVS can be deduced to be equal to 6.
[0337] Additionally, when the value of the sps_video_parameter_set_id syntax element is equal to 0, CVS includes only one layer, and the following applies to that CVS.
[0338] - The nuh_layer_id syntax element value of all VCL NAL units of the SPS that references this CVS is the same as the nuh_layer_id syntax element value of the SPS.
[0339] - The value of the nuh_layer_id syntax element of all VCL NAL units within this CVS is the same as the value of the nuh_layer_id syntax element of the parameter set referenced by the VCL NAL unit.
[0340] In this paper, the parameter set can include APS, PPS, SPS, VPS, etc. Therefore, when CVS includes only one layer, the nuh_layer_id (apsLayerId) of VPS, the nuh_layer_id (ppsLayerId) of PPS, the nuh_layer_id (spsLayerId) of SPS, and the nuh_layer_id (vps_layer_id) of VPS are the same as the nuh_layer_id of the VCL NAL unit.
[0341] Additionally, the value of the syntax element related to inter-layer prediction (vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]) is inferred to be equal to 1. That is, inter-layer prediction is not used.
[0342] When the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referenced by the CLVS with a nuhLayerId having a specific nuh_layer_id value has the same nuh_layer_id as the nuhLayerId.
[0343] The value of sps_video_parameter_set_id is the same for all SPS referenced by CLVS within CVS.
[0344] According to another implementation, Table 8 shown below can be applied to the sps_video_parameter_set_id syntax element.
[0345] [Table 8]
[0346]
[0347] Referring to Table 8, when the value of the sps_video_parameter_set_id syntax element is equal to 0 or greater, the sps_video_parameter_set_id syntax element indicates the value of the vps_video_parameter_set_id syntax element used for the VPS referenced by the SPS in the corresponding bitstream.
[0348] When the sps_video_parameter_set_id syntax element is equal to 0, the SPS in the corresponding bitstream does not reference the VPS, and the VPS is not referenced when decoding each CLVS of the referenced SPS.
[0349] Furthermore, the syntax element (vps_max_layers_minus1) indicating the maximum number of layers within CVS is derived (or deduced) to be equal to 0, and the syntax element (vps_max_sublayer_minus1) indicating the number of time sublayers that can exist in CVS can be deduced to be equal to 6.
[0350] Additionally, when the value of the sps_video_parameter_set_id syntax element is equal to 0, CVS includes only one layer, and the following applies to that CVS.
[0351] - All VCL NAL units within this CVS have the same nuh_layer_id value.
[0352] - The nuh_layer_id of the parameter set referenced by each VCL NAL unit within this CVS is the same as the nuh_layer_id of the VCL NAL unit.
[0353] Additionally, the value of the syntax element (vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]) associated with inter-layer prediction is inferred to be equal to 1.
[0354] When the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referenced by the CLVS with a nuhLayerId having a specific nuh_layer_id value has the same nuh_layer_id as the nuhLayerId.
[0355] The value of sps_video_parameter_set_id is the same for all SPS referenced by CLVS within CVS.
[0356] According to another implementation, Table 9 shown below can be applied to the sps_video_parameter_set_id syntax element.
[0357] [Table 9]
[0358]
[0359] Referring to Table 9, when the value of the sps_video_parameter_set_id syntax element is equal to 0 or greater, the sps_video_parameter_set_id syntax element indicates the value of the vps_video_parameter_set_id syntax element of the VPS referenced by the SPS in the corresponding bitstream.
[0360] When the sps_video_parameter_set_id syntax element is equal to 0, the SPS in the corresponding bitstream does not reference the VPS, and the VPS is not referenced when decoding each CLVS of the referenced SPS.
[0361] Furthermore, the syntax element (vps_max_layers_minus1) indicating the maximum number of layers within CVS is derived (or deduced) to be equal to 0, and the syntax element (vps_max_sublayer_minus1) indicating the number of time sublayers that can exist in CVS can be deduced to be equal to 6.
[0362] Additionally, when the value of the sps_video_parameter_set_id syntax element is equal to 0, CVS comprises only one layer, and all VCL NAL units within CVS and the parameter set referenced by the VCL NAL units have the same nuh_layer_id value.
[0363] Additionally, the value of the syntax element (vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]) associated with inter-layer prediction is inferred to be equal to 1.
[0364] When the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referenced by the CLVS with a specific nuh_layer_id value nuhLayerId has the same nuh_layer_id as nuhLayerId.
[0365] The value of sps_video_parameter_set_id is the same for all SPS referenced by CLVS within CVS.
[0366] As another implementation, when the CVS comprises only one layer, the following constraints may be applied, as shown in Table 10 below.
[0367] [Table 10]
[0368]
[0369] In other words, `spsLayerId` is set to the `nuh_layer_id` of a specific (or particular) SPS NAL unit, and `vclLayerId` is set to the `nuh_layer_id` value of a specific VCL NAL unit. A specific VCL NAL unit does not reference a specific SPS NAL unit except when `spsLayerId` is less than or equal to `vclLayerId`, when the `sps_video_parameter_set_id` value is not equal to 0, and when not all OLS in an OLS indicated by a VPS that includes layers with the same `nuh_layer_id` as `vclLayerId`.
[0370] The ppsLayerId is set to the nuh_layer_id of a specific PPS NAL unit, and the vclLayerId is set to the nuh_layer_id value of a specific VCL NAL unit. A specific VCL NAL unit should not reference a specific PPS NAL unit except when the ppsLayerId is less than or equal to the vclLayerId, when the sps_video_parameter_set_id value is not equal to 0, and when not all OLS in an OLS indicated by a VPS that includes layers with the same nuh_layer_id as the vclLayerId.
[0371] The apsLayerId is set to the nuh_layer_id of a specific APS NAL unit, and the vclLayerId is set to the nuh_layer_id value of a specific VCL NAL unit. A specific VCL NAL unit should not reference a specific APS NAL unit except when the apsLayerId is less than or equal to the vclLayerId, and when not all OLS in an OLS indicated by a VPS that includes layers with the same nuh_layer_id as the vclLayerId.
[0372] Additionally, when CVS comprises only one layer, the following constraints may be applied, as shown in Table 11 below.
[0373] [Table 11]
[0374]
[0375]
[0376] Referring to Table 11, when the sps_video_parameter_set_id syntax element is equal to 0, the SPS in the corresponding bitstream does not reference the VPS, and when decoding each CLVS of the referenced SPS, the VPS is not referenced.
[0377] Additionally, the syntax element indicating the maximum number of layers within CVS (vps_max_layers_minus1) is derived (or deduced) to be equal to 0, and the syntax element indicating the number of time sublayers that can exist in CVS (vps_max_sublayer_minus1) is deduced to be equal to 6.
[0378] Additionally, the CVS comprises only one layer, and the following applies to this CVS. That is, all VCL NAL units within this CVS have the same nuh_layer_id value.
[0379] Additionally, TotalNumOlss, which indicates the total number of OLS specified by the VPS, is inferred to be equal to 1, and NumLayersInOls[0], which indicates the number of layers within the 0th OLS, is inferred to be equal to 1.
[0380] Additionally, NumOutputLayersInOls[0], which indicates the number of output layers within the 0th OLS, is inferred to be equal to 1, and OutputLayerIdInOls[0][0], which indicates the nuh_layer_id value of the 0th output layer within the 0th OLS, is inferred to be equal to 1.
[0381] Additionally, the value of GeneralLayerIdx[nuh_layer_id] is inferred to be equal to nuh_layer_id, and the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is inferred to be equal to 1. When the value of vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, the SPS referenced by the CLSV with a specific nuh_layer_id value nuhLayerId has the same nuh_layer_id as nuhLayerId.
[0382] The sps_video_parameter_set_id value is the same across all SPS referenced by CLVS within CVS.
[0383] Furthermore, as described below, the variable PictureOutputFlag of the current image can be derived (or inferred).
[0384] - PictureOutputFlag is set to 0 when the current layer is not an output layer (i.e., when nuh_layer_id is not the same as OutputLayerIdInOls[TargetOlsIdx][i] for i values in the range of 0 to NumOutputLayersInOls[TargetOlsIdx]-1 (inclusive of the endpoint), or when one of the following conditions is true.
[0385] - The current image is a RASL image, and the NoOutputBeforeRecoveryFlag of the related IRAP image is equal to 1.
[0386] - The current image is a GDR image with a NoOutputBeforeRecoveryFlag equal to 1, or a reconstructed image of a GDR image with a NoOutputBeforeRecoveryFlag equal to 1.
[0387] Otherwise, PictureOutputFlag is set to the same as ph_pic_output_flag.
[0388] Furthermore, the decoding device can output images that do not belong to the output layer. For example, when the AU cannot use images from the output layer, or when only one output layer exists (e.g., due to loss or down-layer switching), among all images that the AU can use, the decoding device can set the PictureOutputFlag to 1 for the image with the highest nuh_layer_id value and a ph_pic_output_flag equal to 1. For all other images that the AU can use, the decoding device can set the PictureOutputFlag to 0.
[0389] Figure 9 and Figure 10 General examples of video / image encoding methods and related components according to embodiments of the present disclosure are shown respectively.
[0390] Figure 9 The publicly disclosed video / image coding methods can be derived from... Figure 2 , Figure 3 and Figure 10 The (video / image) encoding apparatus 200 disclosed herein shall be used to perform this. More specifically, for example, Figure 9S900 and S910 can be executed by the predictor 220 of the encoding device 200, and S920 can be executed by the entropy encoder 240 of the encoding device 200. Figure 9 The video / image encoding methods disclosed herein may include the embodiments described above.
[0391] More specifically, refer to Figure 9 and Figure 10 The predictor 220 of the encoding device can perform at least one of inter-frame prediction or intra-frame prediction on the current block in the current image (S900), and then generate a prediction sample (prediction block) and prediction information about the current block based on the prediction (S910).
[0392] When performing intra-frame prediction, predictor 220 can predict the current block by referring to samples in the current image (neighboring samples of the current block). Predictor 220 can determine the prediction mode to be applied to the current block by using the prediction mode applied to neighboring samples.
[0393] When performing inter-frame prediction, predictor 220 can generate prediction information and predicted blocks for the current block by performing inter-frame prediction based on the motion information of the current block. The prediction information described above may include information related to the prediction mode, information related to motion information, etc. The information related to motion information may include candidate selection information (e.g., merge index, mvp_flag, or mvp_index), which is information used to derive motion vectors. In addition, the information related to motion information may include the aforementioned information about motion vector difference (MVD) and / or reference image index information. Furthermore, the information related to motion information may include information indicating whether L0 prediction, L1 prediction, or dual prediction is applied. For example, predictor 220 can derive motion information of the current block in the current image based on motion estimation. To this end, by using the original block in the original image corresponding to the current block, predictor 220 can search for highly relevant similar reference blocks in fractional pixels within a defined search range in the reference image. Then, predictor 220 can derive motion information from the searched reference blocks. The similarity of blocks can be derived based on the difference between sample values based on phase. For example, block similarity can be calculated based on the sum of absolute differences (SAD) between the current block (or current block template) and a reference block (or reference block template). In this case, motion information can be derived based on the reference block with the minimum SAD within the search region. The derived motion information can be signaled to the decoding device using various methods based on inter-frame prediction modes.
[0394] The residual processor 230 of the encoding device can generate residual samples and residual information based on the predicted samples generated from the predictor 220 and the original images (original blocks, original samples). In this document, residual information refers to information related to the residual samples, and may include information related to the (quantization) transformation coefficients of the residual samples.
[0395] The adder (or reconstructor) of the encoding device can generate reconstructed samples (reconstructed images, reconstructed blocks, reconstructed sample arrays) by adding the residual samples generated in the residual processor 230 and the predicted samples generated in the predictor 220.
[0396] The entropy encoder 240 of the encoding device can encode image information including prediction information generated in the predictor 220, residual information generated in the residual processor 230, etc. (S920). In this document, the image information may further include information about VCL NAL units and information about HLS, and can be transmitted (or sent) to the decoding device in the form of a bitstream. A bitstream is a sequence of bits configured as a stream of NAL units or bytes forming an expression constituting one or more access units (AUs) of CVS. In the case of a single-layer bitstream, the bitstream can be formed by one CVS, and in this case, CVS can be used with the same meaning as a bitstream.
[0397] Information related to HLS may include information / syntax related to the parameter set used to decode image / video information. For example, the parameter set may include APS, PPS, SPS, VPS, etc. SPS may include the sps_video_parameter_set_id syntax element.
[0398] In this implementation, the layer identifier of the VCL NAL unit included in the bitstream and the layer identifier of the parameter set referenced by the VCL NAL unit can be derived based on the value of the sps_video_parameter_set_id syntax element.
[0399] For example, when the value of the sps_video_parameter_set_id syntax element is greater than 0, that is, when the value of the sps_video_parameter_set_id syntax element is not equal to 0, the sps_video_parameter_set_id syntax element can indicate the value of the identifier (vps_video_parameter_set_id syntax element) of the VPS referenced by the SPS.
[0400] When the value of the `sps_video_parameter_set_id` syntax element is equal to 0, the value of the `nuh_layer_id` syntax element of all VCLNAL units within CVS can be equal to the value of the `nuh_layer_id` syntax element of the SPS. Furthermore, the value of the `nuh_layer_id` syntax element of a VCL NAL unit can be equal to the value of the `nuh_layer_id` syntax element of the parameter set referenced by the VCL NAL unit.
[0401] Alternatively, when the value of the sps_video_parameter_set_id syntax element is equal to 0, all VLSNAL units within CVS can have the same NAL unit header layer identifier (nuh_layer_id) value, and the NAL unit header layer identifier (nuh_layer_id) of the parameter set referenced by each VCL NAL unit within CVS can be the same as the NAL unit header layer identifier (nuh_layer_id) of the VCL NAL unit.
[0402] Alternatively, when the value of the sps_video_parameter_set_id syntax element is equal to 0, all VCLNAL units within CVS and the parameter set referenced by the VCL NAL unit have the same NAL unit header layer identifier (nuh_layer_id).
[0403] Furthermore, according to this embodiment, information related to the OLS can be derived based on the value of the `sps_video_parameter_set_id` syntax element. This OLS-related information may include the aforementioned `TotalNumOlss`, `NumLayersInOls[i]`, `NumOutputLayersInOls[i]`, `OutputLayerIdInOls[i][j]`, etc. In this document, `TotalNumOlss` indicates the total number of OLSs specified by the video parameter set. `NumOutputLayersInOls[i]` indicates the number of output layers within the i-th OLS. `OutputLayerIdInOls[i][j]` indicates the value of the layer identifier (nuh_layer_id) in the NAL unit header of the j-th output layer within the i-th OLS.
[0404] For example, when the value of the sps_video_parameter_set_id syntax element is equal to 0, the value of at least one of TotalNumOlss, NumLayersInOls[0], NumOutputLayersInOls[0], and OutputLayerIdInOls[0][0] can be inferred to be equal to 1. Additionally, the value of GeneralLayerIdx[nuh_layer_id] can be inferred to be the same as nuh_layer_id.
[0405] Therefore, according to this specification, even in the case of a single-layer bitstream (where the VSP is absent in the CVS), the encoding efficiency of a decoding device designed to handle only a single-layer bitstream can be improved because the layer identifier of the VPS referenced by the VCL NAL unit can be derived (or inferred). Furthermore, even if the VSP is absent in the CVS, problems that may occur during the decoding process in the absence of the VSP can be prevented because information related to the OLS can be derived (or inferred) or initialized.
[0406] Figure 11 and Figure 12 General examples of video / image decoding methods and related components according to embodiments of the present disclosure are shown respectively.
[0407] Figure 11 The publicly disclosed video / image decoding method can be used by Figure 4 , Figure 5 and Figure 12 The (video / image) decoding device 300 disclosed herein shall be used to perform this operation. More specifically, for example, Figure 11 S1100 can be executed by the entropy decoder 310 of the decoding device. S1110 can be executed by the predictor 330 of the decoding device, and S1120 can be executed by the adder 340 of the decoding device. Figure 11 The video / image decoding methods disclosed herein may include the embodiments described above.
[0408] Reference Figure 11 and Figure 12The entropy decoder 310 of the decoding device can obtain image information including VCL NAL units from the bitstream (S1100). In addition to information related to VCL NAL units, the image information may also include prediction information, residual information, HLS-related information, loop filtering-related information, etc. Prediction information may include inter-frame / intra-frame prediction differentiation information, intra-frame prediction mode related information, inter-frame prediction mode related information, etc. HLS-related information may include information / syntax related to the parameter set used for decoding the image / video information. In this document, the parameter set may include APS, PPS, SPS, VPS, etc. SPS may include the sps_video_parameter_set_id syntax element.
[0409] The entropy decoder 310 of the decoding device can derive (or infer) the layer identifier of the VCL NAL unit, the layer identifier of the parameter set referenced by the VCL NAL unit, and / or information related to OLS based on the value of the sps_video_parameter_set_id syntax element.
[0410] For example, when the value of the sps_video_parameter_set_id syntax element parsed from the bitstream is greater than 0, the entropy decoder 310 of the decoding device can determine (or infer) that the value of the sps_video_parameter_set_id syntax element is equal to the value of the identifier of the VPS (vps_video_parameter_set_id syntax element) referenced by the SPS. However, if the value of the sps_video_parameter_set_id syntax element is equal to 0, the entropy decoder 310 of the decoding device can determine (or infer) that the values of the nuh_layer_id syntax elements of all VCL NAL units of the SPS within the CVS are the same as the values of the nuh_layer_id syntax elements of the SPS, and can determine (or infer) that the values of the nuh_layer_id syntax elements of the VCL NAL units are the same as the values of the nuh_layer_id syntax elements of the parameter set referenced by the VCLNAL units.
[0411] As another example, if the value of the sps_video_parameter_set_id syntax element parsed from the bitstream is greater than 0, the entropy decoder 310 of the decoding device can deduce (or infer) that all VLS NAL units within the CVS have the same NAL unit header layer identifier (nuh_layer_id) value, and can deduce (or infer) that the NAL unit header layer identifier (nuh_layer_id) of the parameter set referenced by each VCL NAL unit within the CVS is the same as the NAL unit header layer identifier (nuh_layer_id) of the VCL NAL unit.
[0412] As yet another example, if the value of the sps_video_parameter_set_id syntax element parsed from the bitstream is greater than 0, the entropy decoder 310 of the decoding device can deduce (or infer) that all VCL NAL units within the CVS have the same NAL unit header layer identifier (nuh_layer_id) as the parameter set referenced by the VCLNAL units.
[0413] Additionally, the entropy decoder 310 of the decoding device can derive OLS-related information based on the value of the sps_video_parameter_set_id syntax element parsed from the bitstream. OLS-related information may include TotalNumOlss, indicating the total number of OLS specified by the video parameter set; NumLayersInOls[i], indicating the number of layers within the i-th OLS; NumOutputLayersInOls[i], indicating the number of output layers within the i-th OLS; and OutputLayerIdInOls[i][j], indicating the value of the layer identifier (nuh_layer_id) in the NAL unit header of the j-th output layer within the i-th OLS.
[0414] For example, if the value of the sps_video_parameter_set_id syntax element parsed from the bitstream is equal to 0, then the value of at least one of TotalNumOlss, NumLayersInOls[0], NumOutputLayersInOls[0], and OutputLayerIdInOls[0][0] can be inferred to be equal to 1. Additionally, the value of GeneralLayerIdx[nuh_layer_id] can be inferred to be the same as nuh_layer_id.
[0415] More specifically, the predictor 330 of the decoding device can perform inter-frame prediction and / or intra-frame prediction on the current block within the current image based on prediction information obtained from the bitstream to generate a prediction sample for the current block (S1110). Subsequently, the residual processor 320 of the decoding device can generate residual samples based on residual information obtained from the bitstream. The adder 340 of the decoding device can generate reconstructed samples based on the prediction samples generated in the predictor 330 and the residual samples generated in the residual processor 320, and then can generate a reconstructed image (reconstructed block) based on the reconstructed samples (S1120).
[0416] Subsequently, loop filtering processes (such as deblocking filtering, SAO, and / or ALF processes) can be applied to the reconstructed image as needed to enhance subjective / objective image quality.
[0417] Although the method has been described based on a flowchart listing the steps or blocks in the sequence of the above embodiments, the steps of this disclosure are not limited to a particular order, and specific steps may be performed in different steps or in a different order or simultaneously relative to those described above. Furthermore, those skilled in the art will understand that the steps in the flowchart are not exclusive, and one or more steps may be included or removed from the flowchart without affecting the scope of this disclosure.
[0418] The methods described above according to this disclosure may be in the form of software, and the encoding and / or decoding devices according to this disclosure may be included in devices for performing image processing, such as TVs, computers, smartphones, set-top boxes, display devices, etc.
[0419] When the embodiments of this disclosure are implemented in software, the above methods can be implemented by modules (processes or functions) that perform the above functions. Modules can be stored in memory and executed by a processor. Memory can be installed inside or outside the processor and can be connected to the processor via various known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, embodiments of this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in the corresponding figures can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information about the implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0420] Furthermore, the decoding and encoding devices using embodiments of this disclosure can be included in multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, and real-time communication devices, such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, image telephony video devices, vehicle terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, or ship terminals), and medical video devices; and can be used to process image signals or data. For example, OTT video devices can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0421] Furthermore, the processing methods applying the embodiments of this disclosure can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having data structures according to embodiments of this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices for storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Computer-readable recording media also include media embodied in the form of a carrier wave (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored in a computer-readable recording medium or transmitted via wired or wireless communication networks.
[0422] Furthermore, the embodiments of this disclosure can be embodied in computer program products based on program code, and the program code can be executed on a computer according to the embodiments of this disclosure. The program code can be stored on a computer-readable medium.
[0423] Figure 13 Examples of content streaming systems to which embodiments of the present disclosure can be applied are provided.
[0424] Reference Figure 13 Content streaming systems that utilize embodiments of this disclosure typically include encoding servers, streaming servers, network servers, media storage devices, user devices, and multimedia input devices.
[0425] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, or camcorders into digital data to generate a bitstream, which is then sent to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, or camcorders generate bitstreams directly, the encoding server can be omitted.
[0426] A bit stream can be generated by an encoding method or bit stream generation method that applies the embodiments of this disclosure, and the streaming server can temporarily store the bit stream during the sending or receiving of the bit stream.
[0427] A streaming server sends multimedia data to a user device via a web server based on a user request, and the web server acts as a medium for notifying the user of services. When a user requests a desired service from the web server, the web server delivers the request to the streaming server, and the streaming server sends multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server is used to control the commands / responses between devices within the content streaming system.
[0428] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0429] For example, user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0430] In a content streaming system, each server can function as a distributed server, and in this case, data received by each server can be processed in a distributed manner.
Claims
1. A video decoding device, the video decoding device comprising: Memory; as well as At least one processor connected to the memory, wherein the at least one processor is configured to: Obtain image information, including the Video Coding Layer (VCL) and Network Abstraction Layer (NAL) units, from the bitstream; By performing inter-frame or intra-frame prediction on the current block within the current image based on the image information, a prediction sample for the current block is generated; and Reconstruct the current block based on the predicted samples. The image information includes the sps_video_parameter_set_id syntax element. Specifically, when the value of the sps_video_parameter_set_id syntax element is greater than 0, the sps_video_parameter_set_id syntax element indicates the value of the identifier for the video parameter set referenced by the sequence parameter set. Wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, i) the maximum number of allowed layers within each coded video sequence CVS is inferred to be equal to 1, and ii) the total number of output layer sets OLS specified by the video parameter set is inferred to be equal to 1, and Wherein, based on the fact that the value of the sps_video_parameter_set_id syntax element is equal to 0, without referring to the parameter specifying whether at least one OLS includes one or more layers, the number of layers in the 0th OLS is inferred to be equal to 1.
2. The video decoding apparatus according to claim 1, wherein, The bit stream is a single-layer bit stream.
3. The video decoding apparatus according to claim 1, wherein, Based on the fact that the value of the sps_video_parameter_set_id syntax element is equal to 0, the number of output layers within the OLS is inferred to be equal to 1.
4. The video decoding device according to claim 1, wherein, Based on the value of the sps_video_parameter_set_id syntax element, the layer identifier of the VCL NAL unit and the layer identifier of the parameter set referenced by the VCL NAL unit are inferred.
5. The video decoding apparatus according to claim 1, wherein, Based on the fact that the value of the sps_video_parameter_set_id syntax element is equal to 0, the values of the layer identifiers of all VCL NAL units referenced by the sequence parameter set are the same as the values of the layer identifiers of the sequence parameter set.
6. The video decoding apparatus according to claim 4, wherein, Based on the fact that the value of the sps_video_parameter_set_id syntax element is equal to 0, the value of the layer identifier of the VCL NAL unit is the same as the value of the layer identifier of the parameter set.
7. The video decoding apparatus according to claim 1, wherein, Based on the fact that the value of the sps_video_parameter_set_id syntax element is equal to 0, the layer identifier of the NAL unit header of the parameter set referenced by each VCL NAL unit in the image information is the same as the layer identifier of the NAL unit header of the VCL NAL unit.
8. The video decoding apparatus according to claim 1, wherein, Based on the VCL NAL unit reference within the VCL NAL unit having a parameter set with a layer identifier different from the layer identifier of the VCL NAL unit, the value of the sps_video_parameter_set_id syntax element is inferred to be non-zero.
9. A video encoding apparatus, the video encoding apparatus comprising: Memory; as well as At least one processor connected to the memory, wherein the at least one processor is configured to: Perform inter-frame prediction or intra-frame prediction on the current block within the current image; Generate prediction information for the current block based on the inter-frame prediction or the intra-frame prediction; and The image information, including the predicted information, is encoded. The image information includes the Video Coding Layer (VCL) network abstraction layer (NAL) unit and the sps_video_parameter_set_id syntax element. Specifically, when the value of the sps_video_parameter_set_id syntax element is greater than 0, the sps_video_parameter_set_id syntax element indicates the value of the identifier for the video parameter set referenced by the sequence parameter set. Wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, i) the maximum number of allowed layers within each coded video sequence CVS is inferred to be equal to 1, and ii) the total number of output layer sets OLS specified by the video parameter set is inferred to be equal to 1, and Wherein, based on the fact that the value of the sps_video_parameter_set_id syntax element is equal to 0, without referring to the parameter specifying whether at least one OLS includes one or more layers, the number of layers in the 0th OLS is inferred to be equal to 1.
10. The video encoding apparatus according to claim 9, wherein, The image information consists of only one layer.
11. The video encoding apparatus according to claim 9, wherein, Based on the value of the sps_video_parameter_set_id syntax element, the layer identifier of the VCL NAL unit and the layer identifier of the parameter set referenced by the VCL NAL unit are inferred.
12. The video encoding apparatus according to claim 11, wherein, Based on the fact that the value of the sps_video_parameter_set_id syntax element is equal to 0, the value of the layer identifier of the VCL NAL unit is the same as the value of the layer identifier of the parameter set.
13. An apparatus for transmitting data relating to image information, the apparatus comprising: At least one processor is configured to generate a bitstream for the image information, wherein the bitstream is generated based on the following operations: performing inter-frame prediction or intra-frame prediction on a current block within a current image; generating prediction information for the current block based on the inter-frame prediction or intra-frame prediction; and encoding the image information including the prediction information; and A transmitter configured to transmit the data comprising the bit stream. The image information includes the Video Coding Layer (VCL) network abstraction layer (NAL) unit and the sps_video_parameter_set_id syntax element. Specifically, when the value of the sps_video_parameter_set_id syntax element is greater than 0, the sps_video_parameter_set_id syntax element indicates the value of the identifier for the video parameter set referenced by the sequence parameter set. Wherein, based on the value of the sps_video_parameter_set_id syntax element being equal to 0, i) the maximum number of allowed layers within each coded video sequence CVS is inferred to be equal to 1, and ii) the total number of output layer sets OLS specified by the video parameter set is inferred to be equal to 1, and Wherein, based on the fact that the value of the sps_video_parameter_set_id syntax element is equal to 0, without referring to the parameter specifying whether at least one OLS includes one or more layers, the number of layers in the 0th OLS is inferred to be equal to 1.