Image encoding method and apparatus therefor
Patent Information
- Application Number
- CN202180060988.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-05-22
- Filing Date
- 2021-05-21
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2041-05-21
AI Technical Summary
因此,当使用诸如传统有线/无线宽带线路这样的介质发送图像数据或者使用现有存储介质存储图像/视频数据时,其传输成本和存储成本增加
[0020] According to this document, the overall image/video compression efficiency can be increased.
Smart Images

Figure CN116195247B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to image coding techniques, and most specifically, to methods and apparatus for encoding images based on a point of view (POC) in an image coding system. Background Technology
[0002] Today, the demand for high-resolution and high-quality images / videos, such as 4K, 8K, or even higher ultra-high-definition (UHD) images / videos, continues to grow across various sectors. As image / video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases compared to traditional image data. Therefore, transmission and storage costs increase when using media such as traditional wired / wireless broadband lines to transmit image data or when using existing storage media to store image / video data.
[0003] In addition, there is increasing interest and demand for immersive media such as virtual reality (VR), artificial reality (AR) content, or holograms, and broadcasting of images / videos with image features that differ from those of real images such as game images is on the rise.
[0004] Therefore, there is a need for efficient image / video compression technology for effectively compressing, transmitting, storing, and reproducing information with high-resolution, high-quality images / videos that have the various characteristics described above. Summary of the Invention
[0005] Technical issues
[0006] The technical aspect of this disclosure is to provide methods and apparatus for improving image coding efficiency.
[0007] This disclosure also provides methods and apparatus for improving the efficiency of POC decoding of images.
[0008] This disclosure also provides methods and apparatus for improving inter-frame prediction efficiency by not using reference images that have been removed from the system.
[0009] This disclosure is also used to reduce errors and stabilize the network by limiting the POC value between the current image and the reference image.
[0010] Technical solution
[0011] In one aspect, an image decoding method performed by a decoding device is provided. The method includes the following steps: receiving POC information and information about a reference image from a bitstream; deriving the POC value of a current image and the POC value of the reference image based on the POC information; constructing a reference image list based on the POC value of the current image and the POC value of the reference image; deriving a prediction sample of the current block by performing inter-frame prediction on the current block based on the reference image list; and generating a reconstructed image based on the prediction sample, wherein the POC information includes the maximum LSB value of the POC, and the information about the reference image includes a non-reference image flag related to whether the image is not used as a reference image, wherein the value of the non-reference image flag of the previous image used to derive the POC value of the current image is 0, and wherein the difference between the POC value of the current image and the POC value of the previous image is less than half of the maximum LSB value of the POC.
[0012] The layer ID used for the current image and the layer ID used for the previous image can be the same, and the time ID derived from the time layer identification information used for the previous image can be 0.
[0013] The preceding image may be neither a RASL image nor a RADL image.
[0014] The POC value of the current image can be derived based on the variable POCMsb and the POC LSB information value of the current image. The variable POCMsb can be derived based on the cyclic existence flag related to the existence of the POC MSB cyclic value and the POC LSB cyclic value that is signaled based on the cyclic existence flag value.
[0015] When the value of the cycle existence flag of the current image is 0 and the current image is not a CLVSS image, the variable POCMsb of the current image is derived based on the variable POCMsb of the previous image.
[0016] On the other hand, an image encoding method performed by an encoding device is provided. The method includes the following steps: deriving the POC value of a current image and the POC value of a reference image; performing inter-frame prediction on the current block using the reference image; and encoding POC information and information about the reference image, wherein the POC information includes the maximum LSB value of the POC, and the information about the reference image includes a non-reference image flag related to whether the image is not used as a reference image, wherein the value of the non-reference image flag of the previous image used to derive the POC value of the current image is 0, and wherein the difference between the POC value of the current image and the POC value of the previous image is less than half of the maximum LSB value of the POC.
[0017] According to another embodiment of this document, a digital storage medium can be provided for storing image data including encoded image information and / or bitstream generated according to an image encoding method performed by an encoding device.
[0018] According to yet another embodiment of this document, a digital storage medium can be provided that stores image data including encoded image information and / or bitstreams that cause a decoding device to perform an image decoding method.
[0019] Technical effect
[0020] According to this document, the overall image / video compression efficiency can be increased.
[0021] According to this disclosure, the POC decoding efficiency of images can be improved.
[0022] According to this disclosure, reference images that have been removed from the system are not used, thus improving the efficiency of inter-frame prediction.
[0023] According to this disclosure, the POC value between the current image and the reference image is limited, thereby reducing the occurrence of errors and stabilizing the network.
[0024] The effects achievable through the specific examples in this specification are not limited to those listed above. For example, various technical effects may exist that can be understood or derived from this specification by one of ordinary skill in the art. Therefore, the specific effects of this specification are not limited to those explicitly described herein, but may include various effects that can be understood or derived from the technical features of this specification. Attached Figure Description
[0025] Figure 1 This is a schematic diagram illustrating the configuration of a video / image encoding device to which this disclosure can be applied.
[0026] Figure 2 This is a schematic diagram illustrating the configuration of a video / image decoding device to which this disclosure can be applied.
[0027] Figure 3 An exemplary hierarchical structure for encoded images / videos is shown.
[0028] Figure 4 The temporal layer structure of NAL units in a bitstream that supports temporal hierarchy is illustrated.
[0029] Figure 5 This is a diagram illustrating an encoding method for image information performed by an encoding device according to an example of this disclosure.
[0030] Figure 6This is a diagram illustrating a method for decoding image information performed by a decoding device according to an example of this disclosure.
[0031] Figure 7 This is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.
[0032] Figure 8 This is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.
[0033] Figure 9 Examples of video / image coding systems applicable to this disclosure are illustrated schematically.
[0034] Figure 10 The structure of a content streaming system applying this disclosure is illustrated. Detailed Implementation
[0035] This disclosure may be modified in various forms, and specific embodiments thereof will be described and illustrated in the accompanying drawings. However, these embodiments are not intended to limit this disclosure. The terminology used in the following description is for the purpose of describing specific embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions, provided that they are clearly read differently. Terms such as “comprising” and “having” are intended to indicate the presence of the features, numbers, steps, operations, elements, components or combinations thereof used in the following description, and therefore it should be understood that the possibility of having or adding one or more different features, numbers, steps, operations, elements, components or combinations thereof is not excluded.
[0036] Furthermore, the elements in the figures described in this disclosure are drawn independently for the purpose of illustrating different specific functions, but this does not mean that these elements are implemented by independent hardware or independent software. For example, two or more of these elements may be combined to form a single element, or a single element may be divided into multiple elements. Embodiments in which elements are combined and / or divided are part of this disclosure without departing from its concept.
[0037] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Furthermore, throughout the drawings, similar reference numerals are used to indicate similar elements, and identical descriptions of similar elements will be omitted.
[0038] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Multi-Functional Video Coding) standard (ITU-T Rec.H.266), the next generation video / image coding standard after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), EVC (Essential Video Coding) standard, AVS2 standard, etc.).
[0039] Various implementations related to video / image encoding may be provided in this document, and these implementations may be combined and performed together unless otherwise specified.
[0040] In this document, video can refer to a collection of images over time. Generally, an image refers to a unit of image representing a specific time period, and a slice / tile is a unit that constitutes a slice / image. A slice / image may include one or more coding tree units (CTUs). An image may consist of one or more slices / tiles. An image may consist of one or more groups of tiles. A group of tiles may include one or more tiles.
[0041] A pixel, or image unit, can refer to the smallest unit that makes up a picture (or image). Alternatively, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when the pixel value is transformed to the frequency domain, it can refer to the transform coefficients in the frequency domain.
[0042] A unit can represent a basic unit of image processing. A unit may include a specific region and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the context, units and terms such as block, region, etc., may be used interchangeably. Typically, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients consisting of M columns and N rows.
[0043] In this document, the terms “ / ” and “,” should be interpreted as indicating “and / or”. For example, the expression “A / B” can mean “A and / or B”. Additionally, “A, B” can mean “A and / or B”. Furthermore, “A / B / C” can mean “at least one of A, B, and / or C”. Additionally, “A / B / C” can mean “at least one of A, B, and / or C”.
[0044] Furthermore, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" can include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".
[0045] In this disclosure, "at least one of A and B" may mean "only A", "only B" or "both A and B". Furthermore, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as "at least one of A and B".
[0046] Additionally, in this disclosure, "at least one of A, B, and C" may mean "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0047] Additionally, the brackets used in this disclosure may mean "for example". Specifically, when indicated as "prediction (intra-frame prediction)", it may mean that "intra-frame prediction" is proposed as an example of "prediction". That is, "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" may be proposed as an example of "prediction". Furthermore, when indicated as "prediction (i.e., intra-frame prediction)", it may also mean that "intra-frame prediction" is proposed as an example of "prediction".
[0048] The technical features described independently in one of the figures in this disclosure may be implemented individually or simultaneously.
[0049] Figure 1 This is an illustrative diagram illustrating the configuration of a video / image encoding apparatus to which embodiments of this disclosure apply. Hereinafter, the term "video encoding apparatus" may include an image encoding apparatus.
[0050] Reference Figure 1 The encoding device 100 may include and be configured with an image segmenter 110, a predictor 120, a residual processor 130, an entropy encoder 140, an adder 150, a filter 160, and a memory 170. The predictor 120 may include an inter-frame predictor 121 and an intra-frame predictor 122. The residual processor 130 may include a transformer 132, a quantizer 133, an inverse quantizer 134, and an inverse transformer 135. The residual processor 130 may also include a subtractor 131. The adder 150 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 110, predictor 120, residual processor 130, entropy encoder 140, adder 150, and filter 160 described above may be constituted by one or more hardware components (e.g., an encoder chipset or processor). Additionally, the memory 170 may include a decoded image buffer (DPB) and may also be constituted by a digital storage medium. The hardware components may also include the memory 170 as an internal / external component.
[0051] Image segmenter 110 can segment an input image (or picture or frame) input to encoding device 100 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). In this case, coding units can be recursively segmented from coding tree units (CTUs) or maximum coding units (LCUs) according to a quadtree-binary-trinary tree (QTBTTT) structure. For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree structure and / or a ternary structure. Alternatively, a binary tree structure can be applied first. The encoding process according to this disclosure can be performed based on the final coding unit that is no longer segmented. In this case, the maximum coding unit can be used as the final coding unit based on coding efficiency according to image characteristics, or, if necessary, the coding unit can be recursively segmented into deeper coding units, and the coding unit with the optimal size can be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes, which will be described later. As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be separated or divided from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving the residual signal from the transform coefficients.
[0052] In some cases, a unit can be used interchangeably with terms such as a block or region. Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can typically represent a pixel or pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. A sample can be used as a term corresponding to a picture (or image) of pixels or pictographs.
[0053] Encoding device 100 generates a residual signal (residual block, residual sample array) by subtracting the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 121 or intra-frame predictor 122 from the input image signal (original block, original sample array), and the generated residual signal is sent to converter 132. In this case, as illustrated, the unit for subtracting the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) within encoding device 100 can be referred to as subtractor 131. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block including the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction on a current block or CU basis. The predictor can generate various information about the prediction, such as prediction mode information, to transmit the generated information to entropy encoder 140, as described later in the description of each prediction mode. The information about the prediction can be encoded by entropy encoder 140 and can be output in the form of a bitstream.
[0054] Intra-predictor 122 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples may be located near or far from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, the directional modes can include, for example, 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-predictor 122 can determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.
[0055] Inter-frame predictor 121 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be called a juxtaposed reference block, a co-located CU (colCU), etc., and the reference image including the temporally neighboring block may be called a juxtaposed image (colPic). For example, inter-frame predictor 121 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 121 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0056] Predictor 120 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can perform prediction on blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as in games, etc. IBC essentially performs prediction in the current frame, but it can be performed similarly to inter-frame prediction by deriving a reference block in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying a palette mode, sample values in the frame can be signaled based on information about the palette index and the palette table.
[0057] The predicted signal generated by the predictor (including inter-frame predictor 121 and / or intra-frame predictor 122) can be used to generate a reconstructed signal or a residual signal. Transformer 132 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform technique can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, when the relationship information between pixels is illustrated as a graph, GBT refers to the transform obtained from that graph. CNT refers to the transform obtained based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform processing can also be applied to square pixel blocks of the same size, or to blocks of variable size that are not square.
[0058] Quantizer 133 can quantize the transform coefficients and send the quantized transform coefficients to entropy encoder 140, which can encode the quantized signal (information about the quantized transform coefficients) to output an encoded quantized signal as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 133 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients. Entropy encoder 140 can perform various encoding methods such as exponential Golomb coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 140 can also encode information necessary for video / image reconstruction (e.g., values of syntax elements, etc.) other than the quantized transform coefficients, either together or separately. The encoded information (e.g., encoded video / image information) can be sent or stored as a bitstream in units of Network Abstraction Layer (NAL) units. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. In this document, syntax elements and / or information transmitted / signed from the encoding device to the decoding device may be included in the video / image information. The video / image information may be included in the bitstream by encoding using the encoding process described above. The bitstream may be transmitted over a network or stored in a digital storage medium. In this document, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, and SSD. A transmitter (not shown) for transmitting the signal output from the entropy encoder 140 or a memory (not shown) for storing the signal may be configured as an internal / external element of the encoding device 100, or the transmitter may also be included in the entropy encoder 140.
[0059] The quantized transform coefficients output from quantizer 133 can be used to generate a prediction signal. For example, dequantizer 134 and inverse transformer 135 can vectorize the transform coefficients and apply dequantization and inverse transform to recover the residual signal (residual block or residual sample). Adder 150 can add the reconstructed residual signal to the prediction signal output from inter-frame predictor 121 or intra-frame predictor 122 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If no residual exists in the block to be processed, such as when a skip mode is applied, the prediction block can be used as the reconstructed block. Adder 150 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current image, and, as described later, is also filtered for inter-frame prediction of the next image.
[0060] In addition, luminance mapping with chroma scaling (LMCS) can be applied during image encoding and / or reconstruction.
[0061] Filter 160 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 170 (specifically, the DPB of memory 170). Various filtering methods may include, for example, deblocking filtering, sample adaptive offset, adaptive loop filtering, bilateral filtering, etc. Filter 160 can generate various filtering-related information and send the generated information to entropy encoder 140, as described later in the description of the various filtering methods. The filtering-related information can be encoded by entropy encoder 140 and output as a bitstream.
[0062] The modified reconstructed image sent to memory 170 can be used as a reference image in inter-frame predictor 121. When inter-frame prediction is applied via the encoding device, prediction mismatch between the encoding device 100 and the decoding device can be avoided, and encoding efficiency can be improved.
[0063] The DPB of memory 170 can store a modified reconstructed image used as a reference image in inter-frame predictor 121. Memory 170 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of reconstructed blocks in the image. The stored motion information can be sent to inter-frame predictor 121 and used as motion information for spatially or temporally neighboring blocks. Memory 170 can store reconstructed samples of reconstructed blocks in the current image and can transmit these reconstructed samples to intra-frame predictor 122.
[0064] Figure 2 This is a schematic diagram illustrating the configuration of a video / image decoding device to which embodiments of the present disclosure apply.
[0065] Reference Figure 2 The decoding device 200 may include and be configured with an entropy decoder 210, a residual processor 220, a predictor 230, an adder 240, a filter 250, and a memory 260. The predictor 230 may include an inter-frame predictor 232 and an intra-frame predictor 231. The residual processor 220 may include an inverse quantizer 221 and an inverse transformer 222. According to embodiments, the entropy decoder 210, residual processor 220, predictor 230, adder 240, and filter 250 described above may be constituted by one or more hardware components (e.g., a decoder chipset or processor). Additionally, the memory 260 may include a decoded image buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 260 as an internal / external component.
[0066] When the input includes a bitstream containing video / image information, the decoding device 200 can respond to... Figure 1 The image is reconstructed through processing video / image information in the illustrated encoding device. For example, decoding device 200 can deduce units / blocks based on block segmentation information obtained from the bitstream. Decoding device 200 can perform decoding using processing units applied to the encoding device. Therefore, the processing unit for decoding can be, for example, an encoding unit, and the encoding unit can be segmented from a CTU or LCU according to a quadtree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by decoding device 200 can be reproduced by a reproduction device.
[0067] Decoding device 200 can receive data from bitstreams. Figure 1The signal output by the encoding device illustrated herein can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can deduce the information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information) by parsing the bitstream. The video / image information may also include information about various parameter sets such as the Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. The syntax elements and / or information to be signaled / received, which will be described later in this document, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 210 can decode the information within the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and output the values of the syntax elements necessary for image reconstruction and the quantized values of the residual correlation transform coefficients. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine the context model using information about the syntax element to be decoded, as well as decoding information of neighboring blocks and the block to be decoded, or information about symbols / bins decoded in the previous stage, and generate symbols corresponding to the values of each syntax element by predicting the generation probability of bins based on the determined context model and performing arithmetic decoding of bins. At this point, the CABAC entropy decoding method can determine the context model and then update the context model using information about the decoded symbols / bins for the context model of the next symbol / bin. Prediction information from the information decoded by the entropy decoder 210 can be provided to the predictors (inter-frame predictor 232 and intra-frame predictor 231), and the residual values (i.e., quantization transform coefficients and related parameter information) from the entropy decoding performed by the entropy decoder 210 can be input to the residual processor 220. The residual processor 220 can derive the residual signals (residual blocks, residual samples, residual sample matrices). Additionally, filtering information from the information decoded by the entropy decoder 210 can be provided to the filter 250. Furthermore, the receiver (not shown) for receiving the signal output from the encoding device can be further configured as an internal / external component of the decoding device 200, or the receiver can also be a component of the entropy decoder 210. Additionally, the decoding device according to this disclosure can be referred to as a video / image / picture decoding device, and the decoding device can also be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 210, and the sample decoder may include at least one of an inverse quantizer 221, an inverse transformer 222, an adder 240, a filter 250, a memory 260, an inter-frame predictor 232, and an intra-frame predictor 231.
[0068] The dequantizer 221 can dequantize the quantized transform coefficients to output transform coefficients. The dequantizer 221 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order executed by the encoding device. The dequantizer 221 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0069] The inverse transformer 222 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0070] Predictor 230 can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the predictions output from entropy decoder 210, and determine a specific intra-frame / inter-frame prediction mode.
[0071] The predictor can generate a prediction signal based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Alternatively, the predictor can perform prediction on blocks based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as in games with screen content coding (SCC). IBC essentially performs prediction in the current frame, but similarly to inter-frame prediction by deriving a reference block in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. The palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be signaled by being included in the video / image information.
[0072] The intra-predictor 231 can refer to samples within the current image to predict the current block. The referenced samples can be located near the current block or away from it, depending on the prediction pattern. The prediction pattern in intra-prediction can include various non-directional and various directional patterns. The intra-predictor 231 can also use prediction patterns applied to neighboring blocks to determine the prediction pattern applied to the current block.
[0073] Inter-frame predictor 232 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 232 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode for the current block.
[0074] Adder 240 can add the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 232 and / or intra-frame predictor 231) to generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array). If the block to be processed has no residual when the skip mode is applied, the prediction block can be used as the reconstruction block.
[0075] Adder 240 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and as described later, it can also be output by filtering or used for inter-frame prediction of the next image.
[0076] In addition, Luminance Mapping with Chroma Scaling (LMCS) can also be applied to image decoding processing.
[0077] Filter 250 can apply filtering to the reconstructed signal, thereby improving the subjective / objective image quality. For example, filter 250 can apply various filtering methods to the reconstructed image to generate a modified reconstructed image, and send the modified reconstructed image to memory 260, specifically, the DPB of memory 260. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bidirectional filtering, etc.
[0078] The (modified) reconstructed image stored in the DPB of memory 260 can be used as a reference image in inter-frame predictor 232. Memory 260 can store motion information of blocks in which motion information within the current image is derived (or decoded) and / or motion information of blocks within previously reconstructed images. The stored motion information can be transmitted to inter-frame predictor 260 to be used as motion information for spatially or temporally neighboring blocks. Memory 260 can store reconstructed samples of reconstructed blocks within the current image and transmit the stored reconstructed samples to intra-frame predictor 231.
[0079] In this document, the exemplary embodiments described in the filter 160, inter-frame predictor 121 and intra-frame predictor 122 of the encoding device 100 can be equally applied to or correspond to the filter 250, inter-frame predictor 232 and intra-frame predictor 231 of the decoding device 200, respectively.
[0080] As described above, prediction is performed to increase compression efficiency during video encoding. Accordingly, a prediction block can be generated that includes prediction samples for the current block, which is the target block for encoding. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in both the encoding and decoding devices, and the encoding device can improve image encoding efficiency by signaling to the decoding device information about the residual between the original block and the prediction block (residual information), which is not the original sample values of the original block themselves. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed image including the reconstructed block.
[0081] Residual information can be generated through transformation and quantization processes. For example, an encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients. This allows it to signal the associated residual information to the decoding device (via a bitstream). Here, the residual information can include the value information, position information, transform technique, transform kernel, quantization parameters, etc., of the quantized transform coefficients. The decoding device can perform quantization / dequantization processes based on the residual information and derive residual samples (or residual sample blocks). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive the residual block by performing dequantization / inverse transform on the quantized transform coefficients for inter-frame prediction reference of the next image, and can generate a reconstructed image based on this.
[0082] Furthermore, in a VVC system, there exists a signaling mechanism that allows system-level entities to know whether an image is used as a reference image for another image. Using this information, system-level entities can remove images under specific circumstances. That is, system-level entities can remove images marked as not being used as a reference image for another image. For example, when network congestion occurs, a Media Identification Network (MIN) router can drop network packets carrying the encoded bits of images marked as not being used as reference images for another image.
[0083] Table 1 below shows the content's identifier information.
[0084] [Table 1]
[0085]
[0086] As shown in Table 1, when the value of ph_non_ref_pic_flag is 1, it means that the image associated with the image header is not used as a reference image, and when the value is 0, it means that the image associated with the image header can be used or not used as a reference image.
[0087] Furthermore, the process for deriving the POC value of the current image, as described in the VVC specification, is as follows.
[0088] This process allows us to derive the variable PicOrderCntVal, which represents the POC of the current image.
[0089] To derive the variable PicOrderCntVal, image information signaled using a more advanced syntax is required, as detailed below.
[0090] The nuh_layer_id is signaled in the NAL cell header and is used to identify the layer to which a VCL NAL cell belongs or to which a non-VCL NAL cell is applied.
[0091] Figure 3 An exemplary hierarchical structure for encoded images / videos is shown. For example... Figure 3 As shown, the encoded image / video is divided into a Video Coding Layer (VCL) that handles the decoding of the image / video and the image / video itself, a lower layer system that sends and stores the encoded information, and a Network Abstraction Layer (NAL) that plays a role in network adaptation and is located between the VCL and the lower layer system.
[0092] VCL can generate VCL data that includes compressed image data (slice data), or it can generate parameter sets or supplementary enhancement information (SEI) messages that are needed to supplement the decoding process of images, such as picture parameter sets (PPS), sequence parameter sets (SPS), video parameter sets (VPS).
[0093] In NAL, NAL cells can be generated by adding header information (NAL cell header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, the RBSP is referred to as slice data, parameter set, SEI message, etc., generated in VCL. The NAL cell header can include NAL cell type information specified according to the RBSP data included in the NAL cell.
[0094] like Figure 3 As shown, based on the RBSP generated in the VCL, the NAL unit can be divided into VCL NAL units and non-VCL NAL units. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data), and a non-VCL NAL unit can refer to a NAL unit that includes information required for decoding the image (parameter set or SEI message).
[0095] The aforementioned VCL NAL units and non-VCL NAL units can be transmitted over the network according to the data specification header information of the lower system. For example, NAL units can be converted into predetermined data formats such as H.266 / VVC file format, RTP (Real-Time Transport Protocol), TS (Transport Stream), etc., and can be transmitted over various types of networks.
[0096] As described above, for a NAL cell, the NAL cell type can be specified according to the RBSP data structure included in the NAL cell, and the information of the NAL cell type can be stored in the NAL cell header and signaled.
[0097] For example, based on whether a NAL unit includes image information (slice data), NAL units can be classified into VCL NAL unit types and non-VCL NAL unit types. VCL NAL unit types can be classified according to the nature and type of the images included in the VCL NAL unit, while non-VCL NAL unit types can be classified according to the type of parameter set.
[0098] The aforementioned NAL unit type can have syntax information specific to the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be `nal_unit_type`, and the NAL unit type can be specified as the `nal_unit_type` value.
[0099] `vps_independent_layer_flag[i]` is a flag used to signal the video parameter set. When the value is 1, it means that the layer indexed as i is not used for inter-layer / inter-frame prediction, i.e., inter-layer prediction. When the value is 0, it means that the layer indexed as i is used for inter-layer prediction.
[0100] `sps_log2_max_pic_order_cnt_lsb_minus4` is a signal that is notified in the sequence parameter set and represents the value of the variable `MaxPicOrderCntLsb` used in the decoding process of the POC. The variable `MaxPicOrderCntLsb` can be specified as 2. (sps_log2_max_pic_order_cnt_lsb_minus4+4 ).
[0101] The value sps_poc_msb_cycle_len_minus1 plus 1 indicates the bit length of the syntax element ph_poc_msb_cycle_val.
[0102] `ph_pic_order_cnt_lsb` represents the POC value of the current image divided by the value of the variable `MaxPicOrderCntLsb`, and the length of `ph_pic_order_cnt_lsb` is `sps_log2_max_pic_order_cnt_lsb_minus4+4` bits. The value of `ph_pic_order_cnt_lsb` exists in the range of 0 to (`MaxPicOrderCntLsb-1`).
[0103] `ph_poc_msb_cycle_present_flag` is a flag indicating whether the syntax element `ph_poc_msb_cycle_val` exists in the image header. A value of 1 indicates that `ph_poc_msb_cycle_val` exists in the image header, and a value of 0 indicates that `ph_poc_msb_cycle_val` does not exist in the image header. The value of `ph_poc_msb_cycle_present_flag` is 0 when `vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]]` is 0 and the image exists in the current Active Entity (AU) of the current layer's reference layer.
[0104] `ph_poc_msb_cycle_val` represents the POC MSB cycle value of the current image. The length of the syntax element `ph_poc_msb_cycle_val` is `sps_poc_msb_cycle_len_minus1+1` bits.
[0105] If vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is 1 and image A exists in the current AU of the reference layer of the current layer, then the variable PicOrderCntVal is deduced to have the same value as PicOrderCntVal of image A, and all VCL NAL units in the current AU need to have the same ph_pic_order_cnt_lsb value.
[0106] Otherwise, when the current layer is not used for inter-layer prediction, the variable PicOrderCntVal for the current image can be derived as follows.
[0107] First, when ph_poc_msb_cycle_present_flag is 0 and the current image is not CLVSS (the beginning of a coding layer video sequence), the variables prevPicOrderCntLsb and prevPicOrderCntMsb can be derived as follows.
[0108] When the TemporalId is 0 and the previous image is not a RASL (Random Access Skip Preamble) or RADL (Random Access Decodeable Preamble) image, and is set to prevTid0Pic, the variable prevPicOrderCntLsb is the same as ph_pic_order_cnt_lsb of prevTid0Pic, and the variable prevPicOrderCntMsb is the same as PicOrderCntMsb of prevTid0Pic.
[0109] Here, TemporalId refers to a variable derived from the identification information of the time layer in a time-series bitstream (or time-series bitstream).
[0110] A time-level bitstream (or time-level bitstream) includes time-level time layer information. This time layer information can be identification information for the time layer specified according to the time level of the NAL unit. For example, as the time layer identification information, the temporal_id syntax information can be used, and this temporal_id syntax information can be stored in the NAL unit header in the encoding device and signaled to the decoding device. Hereinafter, in this disclosure, the time layer may be referred to as a sublayer, a time sublayer, or a time-level layer.
[0111] Figure 4 The temporal layer structure of NAL units in a bitstream that supports temporal hierarchy is illustrated.
[0112] If the bitstream supports time-leveling, the NAL units included in the bitstream have time-level identification information (e.g., temporal_id). For example, a time-level that includes NAL units with a temporal_id of 0 can provide the lowest time-leveling, and a time-level that includes NAL units with a temporal_id of 2 can provide the highest time-leveling.
[0113] exist Figure 4 In the diagram, the block indicated by 'I' is image I, and the block indicated by 'B' is image B. Additionally, arrow markers indicate whether one image references another.
[0114] like Figure 4 As shown, a NAL cell in a time layer with a temporal_id of 0 is a reference image that can be referenced by NAL cells in time layers with temporal_id of 0, 1, or 2. A NAL cell in a time layer with a temporal_id of 1 is a reference image that can be referenced by NAL cells in time layers with temporal_id of 1 or 2. A NAL cell in a time layer with a temporal_id of 2 can be a reference image that can be referenced by NAL cells in the same time layer (i.e., the time layer with temporal_id of 2), or it can be a non-reference image that is not referenced by another image.
[0115] like Figure 4 As shown, if the NAL unit of the time layer with temporal_id of 2, i.e. the highest time layer, is a non-reference image, then the NAL unit can be extracted (or removed) from the bitstream without affecting different images.
[0116] Furthermore, the variable PicOrderCntMsb of the current layer is derived as follows.
[0117] If the value of ph_poc_msb_cycle_present_flag is 1, then the variable PicOrderCntMsb is equal to the value of ph_poc_msb_cycle_val multiplied by MaxPicOrderCntLsb(ph_poc_msb_cycle_val*MaxPicOrderCntLsb).
[0118] Otherwise, that is, when the value of ph_poc_msb_cycle_present_flag is 0, if the current image is a CVSS image, then the variable PicOrderCntMsb equals 0.
[0119] If the value of ph_poc_msb_cycle_present_flag is 0 and the current image is not a CVSS image, the variable PicOrderCntMsb can be derived based on the following formula.
[0120] [Formula 1]
[0121]
[0122] Finally, the variable PicOrderCntVal, which is the POC value of the current image, is derived as the sum of the previously derived variables PicOrderCntMsb and ph_pic_order_cnt_lsb.
[0123] Here, since all CVSS images that do not have a ph_poc_msb_cycle_val value have a variable PicOrderCntMsb value of 0, PicOrderCntVal is equal to ph_pic_order_cnt_lsb.
[0124] The PicOrderCntVal value can be -2 31 Up to 2 31 Within the range of -1, two encoded images with the same nuh_layer_id in a CVS cannot have the same PicOrderCntVal value.
[0125] Additionally, all images in a specific AU should have the same PicOrderCntVal value.
[0126] Furthermore, there is an issue related to non-reference images during the aforementioned POC decoding process. In POC decoding, the inference of PicOrderCntMsb can be delayed based on the POC value er of the image designed with prevTid0Pic. Since the prevTid0Pic of the current image should be the same image in both encoding and decoding processes, the POC value is the same.
[0127] However, when determining prevTid0Pic, since ph_non_ref_pic_flag has a value of 1, it doesn't consider whether prevTid0Pic is an image that can be removed by the system entity. When an image exists that is designated as the prevTid0Pic for the current image, because it has been removed by the system, the decoding device might unintentionally use another image as the prevTid0Pic for decoding the POC for the current image. As a result, the decoding device might derive an incorrect POC value.
[0128] To address this problem, various implementation methods are proposed in this disclosure. Each of these implementation methods can be applied independently or in combination to the image decoding and encoding process.
[0129] 1. During the POC decoding process, the image selected as prevTid0Pic can be restricted to images whose ph_non_ref_pic_flag is not equal to 1.
[0130] 2. The ph_non_ref_pic_flag value of an image with a TemporalId of 0 can be restricted to not being equal to 1.
[0131] 3. When a CLVS (i.e., a coding layer video sequence) has one or more temporal sublayers, any picture in the base temporal sublayer (i.e., TemporalId is 0) can be restricted to not having 1 as the value of ph_non_ref_pic_flag.
[0132] 4. When CLVS contains images that are not all intra-frame images (i.e., the value of intra_only_constraint_flag is equal to 1), images with a TemporalId of 0 can be restricted to not having 1 as the value of ph_non_ref_pic_flag.
[0133] 5. In cases where the CLVS contains images that are not all intra-pictures (i.e., the value of intra_only_constraint_flag is 1) and there are one or more temporal sub-layers in the CLVS, images with a TemporalId of 0 can be restricted to not having 1 as the value of ph_non_ref_pic_flag.
[0134] 6. For two consecutive image pairs with TemporalId of 0 and ph_non_ref_pic_flag of 0, it can be restricted such that the absolute POC difference between the images cannot exceed half the value of MaxPicOrderCntLsb.
[0135] In other words, to prevent incorrect derivation of the POC when using a non-reference image during POC decoding, in this disclosure, a non-reference image with a TemporalId of 0 may not be used during POC decoding. Furthermore, the difference in POC values between two consecutive images with a TemporalId of 0 and ph_non_ref_pic_flag of 0 may be limited to no more than half of MaxPicOrderCntLsb.
[0136] Figure 5This is a diagram illustrating an image information encoding method performed by an encoding device according to an example of this disclosure, and Figure 6 This is a diagram illustrating a method for decoding image information performed by a decoding device according to an example of this disclosure.
[0137] The encoding device can derive the POC value of the reference image to construct a reference image set, and can derive the POC value of the current image (step S510).
[0138] The encoding device can encode the POC information of the derived current image (step S520) and can encode image information including POC information (step S530).
[0139] In response to the operation performed in the encoding device, the decoding device can obtain image information including POC information from the bit stream (step S610), and can deduce the POC values of the reference image and the current image based on the POC information (step S620).
[0140] A reference image set can be constructed based on the derived POC values (step S630), and a reference image list can be derived based on the reference image set (step S640).
[0141] Based on the derived list of reference images, inter-frame prediction for the current image can be performed (step S650).
[0142] Image information, such as POC information, can be included in the HLS (High-Level Syntax). POC information can include POC-related information and syntax elements, and can include POC information related to the current image and / or a reference image. POC information can include at least one of ph_non_reference_picture_flag, ph_non_reference_picture_flag, ph_poc_msb_cycle_present_flag, and / or ph_poc_msb_cycle_val.
[0143] It can be omitted, such as Figure 5 and Figure 6 The reference image set or reference image list shown is used for derivation. For example, step S640 of deriving the reference image list can be omitted, and inter-frame prediction can be performed based on the reference image set.
[0144] Alternatively, according to another example, a list of reference images can be derived based on POC values, replacing steps S630 for deriving the set of reference images and S640 for deriving the list of reference images. For example, the POC value of the i-th reference image can be derived based on the POC difference indicated by POC information associated with the reference images. In this case, when i is 0, the POC information can represent the POC difference between the current image and the i-th reference image, and when i is greater than 0, the POC information can represent the POC difference between the i-th reference image and the (i-1)-th reference image. Reference images may include previous reference images with POC values smaller than the current image and / or subsequent reference images with POC values larger than the current image.
[0145] The embodiments proposed in this disclosure will be described in detail below.
[0146] [Implementation Method 1]
[0147] Table 2 corresponds to examples that implement the second example described above (where the ph_non_ref_pic_flag value of the image with a TemporalId of 0 is restricted to be not equal to 1). In Table 2, underlined portions are marked according to the implementation method and are based on the current VVC specification.
[0148] [Table 2]
[0149]
[0150] [Implementation Method 2]
[0151] Table 3 corresponds to examples that implement the third example described above (when CLVS has one or more time sub-layers, any picture in the base time sub-layer is restricted to not having 1 as the value of ph_non_ref_pic_flag). In Table 3, underlined portions are marked according to the implementation based on the current VVC specification.
[0152] [Table 3]
[0153]
[0154] [Implementation Method 3]
[0155] Table 4 corresponds to examples that implement the fourth example described above (when CLVS includes images that are not entirely intra-pictures, images with a TemporalId of 0 are restricted to not having 1 as the value of ph_non_ref_pic_flag). In Table 4, underlined portions are marked according to the implementation based on the current VVC specification.
[0156] [Table 4]
[0157]
[0158]
[0159] [Implementation Method 4]
[0160] Table 5 corresponds to examples implementing the fifth example described above (where, in the case that the CLVS contains images that are not all intra-pictures (i.e., the value of intra_only_constraint_flag is 1) and there are one or more temporal sub-layers in the CLVS, images with a TemporalId of 0 are restricted to not having a value of ph_non_ref_pic_flag of 1). In Table 5, underlined portions are marked according to the implementation based on the current VVC specification.
[0161] [Table 5]
[0162]
[0163] [Implementation Method 5]
[0164] Table 6 corresponds to an example implementing the sixth example above (for two consecutive image pairs whose TemporalId is 0 and ph_non_ref_pic_flag is 0, the absolute POC difference between the images is limited to no more than half the value of MaxPicOrderCntLsb). In Table 6, underlined portions are marked according to the implementation method and are based on the current VVC specification.
[0165] [Table 6]
[0166]
[0167]
[0168]
[0169] The following figures are provided to illustrate specific examples of this disclosure. Since the names of specific devices and signals / messages / fields shown in the figures are provided by way of example, the technical features of this disclosure are not limited to the specific names used in the following figures.
[0170] Figure 7 This is a flowchart illustrating the operation of a video decoding device according to an embodiment of the present disclosure.
[0171] Figure 7 Each step disclosed in the document is based on the above. Figures 3 to 6 Some of the content described above. Accordingly, the omissions or brief descriptions will be compared with the above. Figures 2 to 6A detailed description of the overlapping content described in the text.
[0172] According to the embodiment, the decoding device 200 can receive POC information and reference image information from the bitstream. The POC information may include the maximum LSB value of the POC, and the reference image information may include a non-reference image flag related to whether the image is used as a reference image (step S710).
[0173] The non-reference image flag can be ph_non_ref_pic_flag as shown in Table 1. A value of 1 indicates that the image associated with the image header cannot be used as a reference image, while a value of 0 indicates that the image associated with the image header can or may not be used as a reference image. That is, an image with a non-reference image flag value of 0 is not used as a reference image for another image. In other words, an image used as a reference image for another image has a non-reference image flag value of 1.
[0174] The received POC information may include vps_independent_layer_flag, sps_log2_max_pic_order_cnt_lsb_minus4, sps_poc_msb_cycle_len_minus1, ph_pic_order_cnt_lsb, ph_poc_msb_cycle_present_flag, ph_poc_msb_cycle_val, etc., and these types of information can be signaled in the image header or sequence parameter set. The syntax for signaling is described as described above.
[0175] The decoding device can deduce the POC values of the current image and the reference image based on the POC information for inter-frame prediction of the current image and for generating the reference image list (step S720).
[0176] The variable PicOrderCntVal, which indicates the POC value of the current image, can be derived as the sum of the variable PicOrderCntMsb (variable POCMsb), which indicates the MSB value of the current image, and the value of ph_pic_order_cnt_lsb (POC LSB information), which indicates the LSB of the current image that is signaled in the image header (variable PicOrderCntMsb + ph_pic_order_cnt_lsb).
[0177] If the current image's vps_independent_layer_flag value is 0 and the current layer is used as a reference image, the current image has the same POC value as the image included in the current AU in the reference layer.
[0178] Otherwise, if the current layer is not used in inter-layer prediction, the variable PicOrderCntVal for the current image can be derived based on the POC MSB cycle value (ph_poc_msb_cycle_val) signaled by the ph_poc_msb_cycle_present_flag value (cycle presence flag) and the cycle presence flag value. In this case, different derivation processes can be applied depending on the presence of the cycle presence flag value and whether the current image is a CLVSS (start of coding layer video sequence) image.
[0179] The first case corresponds to the situation where ph_poc_msb_cycle_present_flag is 0 and the current image is not a CLVSS image. For POC derivation, the variables prevPicOrderCntLsb and prevPicOrderCntMsb of the previous image can be derived, and the variable POCMsb of the current image can be derived based on the variable POCMsb of the previous image.
[0180] When the current image and nuh_layer_id are the same when TemporalId is 0, and the previous image other than a RASL (Random Access Skip Preamble) image or a RADL (Random Access Decodeable Preamble) image is set to prevTid0Pic, the variable prevPicOrderCntLsb is deduced to be the same as ph_pic_order_cnt_lsb of prevTid0Pic, and the variable prevPicOrderCntMsb is deduced to be the same as PicOrderCntMsb of prevTid0Pic.
[0181] In this case, the non-reference image flag of the previous image used to derive the POC value of the current image is 0, and there may be a restriction that the difference between the POC values of the current image and the previous image is less than half of the maximum LSB value of the POC.
[0182] Additionally, the layer IDs used for the current image and the previous image are the same, and the temporal ID (TemporalId) derived from the time layer identifier information used for the previous image is 0. The previous image derived from the POC used for the current image is neither a RASL image nor a RADL image.
[0183] Subsequently, the variable PicOrderCntVal can be derived based on the size of the current image's ph_pic_order_cnt_lsb and the variable prevPicOrderCntLsb of the previous image, as shown in Equation 1.
[0184] The second case corresponds to the situation where ph_poc_msb_cycle_present_flag is 0 and the current image is a CLVSS image. Since the value of PicOrderCntMsb is 0, the variable PicOrderCntVal can be deduced as the value of ph_pic_order_cnt_lsb.
[0185] The third case corresponds to the value of ph_poc_msb_cycle_present_flag being 1. In this case, the variable PicOrderCntMsb is derived as the value obtained by multiplying ph_poc_msb_cycle_val by MaxPicOrderCntLsb(ph_poc_msb_cycle_val*MaxPicOrderCntLsb). Consequently, the variable PicOrderCntVal is derived as the sum of the derived variable PicOrderCntMsb and the signaled value of ph_pic_order_cnt_lsb.
[0186] The decoding device can construct a list of reference images based on the POC value of the current image and the POC value of the reference image (step S730), and can derive the prediction sample of the current block by performing inter-frame prediction of the current block (step S740).
[0187] Furthermore, the decoding device 200 can decode the quantization transform coefficients of the current block from the bitstream and deduce the quantization transform coefficients of the target block based on the information of the quantization transform coefficients of the current block. The information of the quantization transform coefficients of the target block can be included in the Sequence Parameter Set (SPS) or slice header, and can include at least one of the following: information on whether a reduced transform (RST) is applied, information on the reduction factor, information on the minimum transform size for applying RST, information on the maximum transform size for applying RST, information on the size of the reduced inverse transform, and information indicating the transform index of any of the transform kernel matrices included in the transform set.
[0188] The decoding device 200 can derive the residual information of the current block, i.e., the transform coefficients, by performing dequantization on the quantized transform coefficients, and can arrange the derived transform coefficients in a predetermined scan order.
[0189] The transform coefficients derived from residual information can be either dequantization transform coefficients or quantization transform coefficients as described above. That is, the transform coefficients can be data that can be checked for non-zero data in the current block without considering whether quantization has been performed on it.
[0190] Decoding devices can derive residual samples by applying an inverse transform to the quantization transform coefficients.
[0191] Subsequently, the decoding device can generate a reconstructed image based on the residual samples and the predicted samples (step S750).
[0192] The following figures are provided to illustrate specific examples of this disclosure. Since the names of specific devices and signals / messages / fields shown in the figures are provided by way of example, the technical features of this disclosure are not limited to the specific names used in the following figures.
[0193] Figure 8 This is a flowchart illustrating the operation of a video encoding device according to an embodiment of the present disclosure.
[0194] Figure 8 Each step disclosed in the document is based on the above. Figures 3 to 6 Some of the content described above. Accordingly, the omissions or brief descriptions will be compared with the above. Figures 2 to 6 A detailed description of the overlapping content described in the text.
[0195] The encoding device 100 according to the embodiment can derive the POC values of the current image and the reference image (step S810), and can perform inter-frame prediction of the current block by using the derived POC values and the reference image (step S820).
[0196] The encoding device can encode and output POC information including the maximum LSB value and reference image information including a non-reference image flag related to whether the image is not used as a reference image. The value of the non-reference image flag of the previous image used to derive the POC value of the current image can be 0, and the difference between the POC values of the current image and the previous image can be constructed to be less than half of the maximum LSB value (step S830).
[0197] The POC information for the current image, the method used to derive the POC for the current image, the constraints on previous images, and the constraints on the POC values of previous images and their references. Figure 7 The descriptions of the decoding devices are essentially the same, and overlapping descriptions are omitted.
[0198] The encoding device can derive the residual samples of the current block based on the predicted samples and generate residual information through transformations. The residual information can include the transformation-related information / syntax elements mentioned above. The encoding device can encode image / video information including the residual information, thereby outputting the image / video information in bitstream format.
[0199] More specifically, the encoding device can generate information about the quantization transform coefficients and encode the information about the generated quantization transform coefficients.
[0200] In this disclosure, at least one of quantization / dequantization and / or transform / inverse transform can be omitted. When quantization / dequantization is omitted, the quantization transform coefficients can be referred to as transform coefficients. When transform / inverse transform is omitted, the transform coefficients can be referred to as coefficients or residual coefficients, or, for consistency of expression, may still be referred to as transform coefficients.
[0201] In the above embodiments, the method is illustrated using a series of steps or blocks based on a flowchart. However, this disclosure is not limited to the order of the steps, and a step may be performed in a different order or sequence than described above, or may occur simultaneously with another step. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and one or more steps in the flowchart may be incorporated or removed without affecting the scope of this disclosure.
[0202] The methods described above according to this disclosure can be implemented in software form, and the encoding and / or decoding devices according to this disclosure can be included in image processing devices such as TVs, computers, smartphones, set-top boxes, and display devices.
[0203] When the embodiments of this disclosure are implemented in software, the methods described above can be implemented as modules (processes, functions, etc.) to perform the functions described above. Modules can be stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor in various known ways. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. Furthermore, the functional units shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0204] Furthermore, the decoding and encoding devices using this disclosure can be included in multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, and medical video devices, and can be used to process video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0205] Furthermore, the processing methods of this disclosure can be generated in the form of a program executed by a computer and stored in a computer-readable recording medium. Multimedia data with a data structure according to this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices in which computer-readable data is stored. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Additionally, computer-readable recording media include media implemented in the form of a carrier wave (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks. Furthermore, embodiments of this disclosure can be implemented as computer program products by program code, and the program code can be executed on a computer by embodiments of this disclosure. The program code can be stored on a computer-readable medium.
[0206] Figure 9 Examples of video / image coding systems applicable to this disclosure are illustrated schematically.
[0207] Reference Figure 9 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0208] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0209] Video sources can be obtained through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, video / image archives including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc. In this case, the video / image capture process can be replaced by a process that generates related data.
[0210] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization to optimize compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0211] A transmitter can send encoded video / image information or data, output in bitstream form, to a receiver via a digital storage medium or network, either as a file or a stream. Digital storage media can include various media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to a decoding device.
[0212] Decoding devices can decode video / images by performing a series of processes such as inverse quantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0213] The renderer can render decoded video / images. The rendered video / images can then be displayed on a monitor.
[0214] Figure 10 The structure of a content streaming system applying this disclosure is illustrated.
[0215] In addition, the content streaming system using this disclosure may mainly include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0216] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then send it to a streaming server. As another example, in cases where multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. Furthermore, the streaming server can temporarily store the bitstream during the processing of sending or receiving the bitstream.
[0217] A streaming server sends multimedia data to a user device via a web server based on a user's request. The web server acts as an instrument to inform the user what services are available. When a user requests a desired service, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this regard, the content streaming system may include a separate control server, which in this case controls the commands / responses between the corresponding devices in the content streaming system.
[0218] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to smoothly provide streaming services.
[0219] For example, user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touchscreen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. Each server in the content streaming system can operate as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
[0220] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims can be combined to implement or perform in a device, and the technical features of the device claims can be combined to implement or perform in a method. Furthermore, the technical features of the method claims and the device claims can be combined to implement or perform in a device, and the technical features of the method claims and the device claims can be combined to implement or perform in a method.
Claims
1. An image decoding method performed by a decoding device, the method comprising: Receive image sequence count POC information and information about the reference image from the bitstream; Based on the POC information, the POC value for the current image and the POC value for the reference image are derived. A list of reference images is constructed based on the POC value of the current image and the POC value of the reference image. Predictive samples for the current block are derived by performing inter-frame prediction on the current block based on the reference image list. as well as The reconstructed image is generated based on the predicted samples. The POC information includes information for determining the maximum least significant bit (LSB) value of the POC, and the information regarding the reference image includes a non-reference image flag related to whether the image is not used as a reference image. Specifically, the specific previous image used to derive the POC value of the current image is determined to be a previous image with a non-reference image flag equal to 0, a time identifier ID equal to 0, and a layer ID equal to the layer ID of the current image. Wherein, the difference between the POC value of the current image and the POC value of the specific previous image is less than half of the maximum LSB value of the POC.
2. The method according to claim 1, wherein, The specific preceding image is neither a randomly accessed skipping preceding RADL image nor a randomly accessed decodeable preceding RADL image.
3. The method according to claim 1, wherein, The POC value of the current image is derived based on the variable POCMsb. Specifically, the variable POCMsb is derived based on the cyclic presence flag related to the existence of the most significant bit (MSB) of the POC and the cyclic value of the POC MSB, which is transmitted via signaling based on the cyclic presence flag value.
4. The method according to claim 3, wherein, When the value of the cycle presence flag for the current image is 0 and the current image is not a coding layer video sequence start (CLVSS) image, the variable POCMsb of the current image is derived based on the variable POCMsb of the specific previous image.
5. An image encoding method performed by an encoding device, the method comprising: Derive the image sequence count (POC) value for the current image and the POC value for the reference image; Perform inter-frame prediction on the current block using the reference image; as well as The POC information and information about the reference image are encoded. The POC information includes information for determining the maximum least significant bit (LSB) value of the POC, and the information regarding the reference image includes a non-reference image flag related to whether the image is not used as a reference image. Specifically, the specific previous image used to derive the POC value of the current image is determined to be a previous image with a non-reference image flag equal to 0, a time identifier ID equal to 0, and a layer ID equal to the layer ID of the current image. Wherein, the difference between the POC value of the current image and the POC value of the specific previous image is less than half of the maximum LSB value of the POC.
6. The method according to claim 5, wherein, The specific preceding image is neither a randomly accessed skipping preceding RADL image nor a randomly accessed decodeable preceding RADL image.
7. The method according to claim 5, wherein, The POC value of the current image is derived based on the variable POCMsb. The variable POCMsb is derived based on whether there is a cycle value of the most significant bit (MSB) of the POC for the current image and the cycle value of the POC MSB for the current image.
8. The method according to claim 7, wherein, When there is no POC MSB cycle value for the current image and the current image is not the start of the coding layer video sequence (CLVSS) image, the variable POCMsb of the current image is derived based on the variable POCMsb of the specific previous image.
9. A non-transitory computer-readable storage medium having a computer program and a bit stream stored thereon, wherein when the computer program is executed by a processor, it implements the image encoding method of claim 5 to generate the bit stream.
10. A method for transmitting data for image information, the method comprising: A bitstream for the image information is generated, wherein the bitstream is generated based on the following operations: deriving the image sequence count (POC) value for the current image and the POC value for a reference image; performing inter-frame prediction on the current block using the reference image; and encoding the POC information and information about the reference image; and Send the data including the bit stream. The POC information includes information for determining the maximum least significant bit (LSB) value of the POC, and the information regarding the reference image includes a non-reference image flag related to whether the image is not used as a reference image. Specifically, the specific previous image used to derive the POC value of the current image is determined to be a previous image with a non-reference image flag equal to 0, a time identifier ID equal to 0, and a layer ID equal to the layer ID of the current image. Wherein, the difference between the POC value of the current image and the POC value of the specific previous image is less than half of the maximum LSB value of the POC.