Image encoding / decoding method and apparatus and recording medium storing bitstream therein

By introducing Encoder Optimization Information (EOI) SEI messages into the bitstream, the problem of identifying and processing additional images during temporal upsampling in high-resolution and high-quality image encoding is solved, improving the efficiency and accuracy of image encoding and decoding.

CN122397256APending Publication Date: 2026-07-14LG ELECTRONICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LG ELECTRONICS INC
Filing Date
2025-07-02
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing image compression techniques struggle to effectively handle the encoding and decoding of high-resolution and high-quality images, especially during temporal upsampling optimization, where it is difficult to identify and process additional images.

Method used

Encoder Optimization Information (EOI) supplemental enhancement information (SEI) messages are introduced into the bitstream, containing source image identification information and temporal resampling type flags, to identify and process additional images generated during temporal upsampling.

Benefits of technology

It enables easy identification and processing of additional images generated during temporal upsampling in the bitstream, improving the efficiency and accuracy of image encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122397256A_ABST
    Figure CN122397256A_ABST
Patent Text Reader

Abstract

A method and apparatus for image decoding according to the present disclosure can receive a bitstream including encoded video pictures and can recover the encoded video pictures included in the bitstream. The bitstream can be configured to include an encoder optimization information (EOI) supplemental enhancement information (SEI) message. The EOI SEI message can include source picture identification information about whether a current picture is a source picture for a temporal upsampling optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to image encoding / decoding methods and apparatus, as well as recording media for storing bit streams. Background Technology

[0002] Recently, the demand for high-resolution and high-quality images, such as HD (high-definition) and UHD (ultra-high-definition) images, has been increasing in various application fields, and therefore, efficient image compression technology is being discussed.

[0003] There are various techniques, such as inter-frame prediction, which uses video compression technology to predict the pixel values ​​included in the current image from images before or after the current image; intra-frame prediction, which uses pixel information in the current image to predict the pixel values ​​included in the current image; and entropy coding technology, which assigns short symbols to values ​​that occur frequently and long symbols to values ​​that occur infrequently. These image compression techniques can be used to effectively compress image data and transmit or store it. Summary of the Invention

[0004] Technical issues

[0005] This disclosure provides a method and apparatus for configuring encoder optimization information.

[0006] This disclosure provides a method and apparatus for transmitting encoder-optimized information signals using signals.

[0007] Technical solution

[0008] The image decoding method and apparatus according to this disclosure can receive a bitstream including encoded video images and reconstruct the encoded video images included in the bitstream. The bitstream can be configured to include encoder optimization information (EOI) supplementary enhancement information (SEI) messages.

[0009] In the image decoding method and apparatus according to this disclosure, the EOI SEI message may include source image identification information related to whether the current image is a source image used for temporal upsampling optimization. The EOI SEI message can be obtained from the Network Abstraction Layer (NAL) unit of the bitstream.

[0010] In the image decoding method and apparatus according to the present disclosure, the source image identification information of the first value can indicate that the current image is a source image for temporal upsampling optimization, and the source image identification information of values ​​other than the first value can not indicate that the current image is a source image for temporal upsampling optimization.

[0011] In the image decoding method and apparatus according to this disclosure, source image identification information can be signaled based on at least one of a temporal resampling type flag or image quantity related information. The temporal resampling type flag can indicate the type of temporal resampling optimization. The image quantity related information can be information related to the number of images excluded from or added between encoded image pairs by the encoding system.

[0012] In the image decoding method and apparatus according to the present disclosure, a time resampling type flag with a value of 0 can indicate that the time resampling optimization is subsampling, and a time resampling type flag with a value of 1 can indicate that the time resampling optimization is upsampling.

[0013] In the image decoding method and apparatus according to the present disclosure, source image identification information can be transmitted by signaling based on the value of the time resampling type flag being 0 and the value of the image quantity related information being greater than 0.

[0014] In the image decoding method and apparatus according to the present disclosure, the EOI SEI message may further include a source image presence flag indicating whether source image identification information exists in the EOI SEI message.

[0015] In the image decoding method and apparatus according to the present disclosure, the source image identification information can be sent by signaling based on the source image presence flag indicating that the source image identification information exists in the EOI SEI message.

[0016] In the image decoding method and apparatus according to this disclosure, the EOI SEI message may include at least one of a temporal resampling type flag or image quantity information. The temporal resampling type flag may indicate the type of temporal resampling optimization. The image quantity information may be information related to the number of images excluded by the encoding system from the encoded image pair or added between the encoded image pairs.

[0017] In the image decoding method and apparatus according to this disclosure, when the temporal resampling type flag indicates that the temporal resampling optimization is upsampling and the value of the information related to the number of images is greater than 0, the EOI SEI message can be restricted to the access unit of images other than those added during the temporal upsampling process.

[0018] The image encoding method and apparatus of this disclosure can receive a video image to be encoded, encode the received video image to generate video information related to the video image, generate encoder optimization information (EOI) supplementary enhancement information (SEI) messages, and generate a bitstream including video information and EOI and SEI messages.

[0019] In the image encoding method and apparatus according to this disclosure, the EOI SEI message may include source image identification information related to whether the current image is a source image used for temporal upsampling optimization. The EOI SEI message may be encoded into the network abstraction layer (NAL) unit of the bitstream.

[0020] A computer-readable digital storage medium is provided that stores encoded video / image information, thereby causing an image decoding method to be performed by a decoding apparatus according to the present disclosure.

[0021] A computer-readable digital storage medium is provided according to the present disclosure for storing video / image information generated according to an image encoding method.

[0022] A method and apparatus for transmitting video / image information generated according to an image encoding method are provided according to the present disclosure.

[0023] Beneficial effects

[0024] As shown in this disclosure, by defining information for distinguishing additional images generated during temporal upsampling, additional images can be easily identified when one or more images need to be removed from the bitstream. Attached Figure Description

[0025] Figure 1 A video / image encoding system according to this disclosure is shown.

[0026] Figure 2 A schematic block diagram illustrating an encoding apparatus to which embodiments of the present disclosure are applicable and to perform encoding of video / image signals is shown.

[0027] Figure 3 A schematic block diagram of a decoding apparatus to which embodiments of the present disclosure are applicable and to perform decoding of video / image signals is shown.

[0028] Figure 4 The illustration shows a method for reconstructing video images performed by a decoding device 300 according to the present disclosure.

[0029] Figure 5 The illustration shows a schematic configuration of a decoding apparatus 300 that performs a method for reconstructing video images according to the present disclosure.

[0030] Figure 6 The figure illustrates a method for generating a bitstream performed by an encoding device 200 according to the present disclosure.

[0031] Figure 7 The figure shows a schematic configuration of an encoding apparatus 200 for performing a method for generating a bitstream according to the present disclosure.

[0032] Figure 8Examples of content streaming systems to which embodiments of this disclosure can be applied are shown. Detailed Implementation

[0033] Because this disclosure can be modified in various ways and has several embodiments, specific embodiments will be illustrated in the accompanying drawings and described in detail in the detailed description. However, this disclosure is not intended to be limited to the specific embodiments and should be understood to include all variations, equivalents, and substitutions included within the spirit and scope of this disclosure. Similar reference numerals are used for similar components in the description of each drawing.

[0034] Terms such as "first," "second," etc., may be used to describe various components, but components should not be limited by these terms. These terms are used only to distinguish one component from other components. For example, without departing from the scope of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. Terms and / or combinations of any one or more related statement items are included.

[0035] When a component is described as "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to another component, but there may also be another component in between. On the other hand, when a component is described as "directly connected" or "directly linked" to another component, it should be understood that there is no other component in between.

[0036] The terminology used in this application is for describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, it should be understood that terms such as “comprising” or “having” are intended to designate the presence of features, numbers, steps, operations, components, portions, or combinations thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, portions, or combinations thereof.

[0037] This disclosure relates to video / image coding. For example, the methods / exercises disclosed herein can be applied to methods disclosed in the Essential Video Coding (VVC) standard. Additionally, the methods / exercises disclosed herein can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding Standard 2 (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).

[0038] This specification presents various embodiments of video / image encoding, and unless otherwise stated, these embodiments may be combined with each other to perform the task.

[0039] Here, "video" can refer to a collection of images over time. "Image" generally refers to a unit representing an image within a specific time period, and a tile is a unit that forms part of an image during encoding. A tile can include at least one coding tree unit (CTU). An image can consist of at least one tile. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of an image. A tile column is a rectangular area of ​​CTUs with the same height as the image and a width assigned by the syntax requirements of the image parameter set. A tile row is a rectangular area of ​​CTUs with the same height assigned by the image parameter set and a width equal to the width of the image. CTUs within a tile can be arranged consecutively according to a CTU raster scan, and tiles within an image can be arranged consecutively according to a tile raster scan. A tile can include an integer number of complete tiles or an integer number of consecutive complete CTU rows that can be exclusively included within a single NAL unit of the image. Simultaneously, an image can be divided into at least two sub-images. A sub-image can be a rectangular area of ​​at least one tile within an image.

[0040] A pixel, cell, or pixel unit can refer to the smallest unit that makes up a picture (or image). Additionally, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.

[0041] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, units may be used interchangeably with terms such as block or region. In general, an MxN block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.

[0042] Here, "A or B" can refer to "A only", "B only", or "both A and B". In other words, "A or B" can be interpreted as "A and / or B". For example, "A, B or C" can refer to "A only", "B only", "C only", or "any combination of A, B and C".

[0043] The forward slash ( / ) or comma used in this article can refer to "and / or". For example, "A / B" can refer to "A and / or B". Therefore, "A / B" can refer to "A only", "B only", or "both A and B". For example, "A, B, C" can refer to "A, B, or C".

[0044] Here, "at least one of A and B" can refer to "only A", "only B" or "both A and B". Furthermore, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".

[0045] Additionally, here, "at least one of A, B, and C" can refer to "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can refer to "at least one of A, B, and C".

[0046] Additionally, the parentheses used in this document can refer to "for example". Specifically, when the indication is "prediction (intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction". In other words, "prediction" here is not limited to "intra-frame prediction", and "intra-frame prediction" can be cited as an example of "prediction". Furthermore, even when the indication is "prediction (i.e., intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction".

[0047] Here, a technical feature described individually in a single figure can be implemented individually or simultaneously.

[0048] Figure 1 A video / image encoding system according to this disclosure is shown.

[0049] refer to Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device).

[0050] A source device can transmit encoded video / image information or data to a receiving device in the form of a file or stream via digital storage media or a network. The source device may include a video source, an encoding apparatus, and a transmitting unit. The receiving device may include a receiving unit, a decoding apparatus, and a renderer. The encoding apparatus may be referred to as a video / image encoding apparatus, and the decoding apparatus may be referred to as a video / image decoding apparatus. A transmitter may be included in the encoding apparatus. A receiver may be included in the decoding apparatus. The renderer may include a display unit, and the display unit may consist of a separate device or an external component.

[0051] A video source can acquire video / images through the process of capturing, compositing, or generating video / images. A video source can include devices for capturing video / images and devices for generating video / images. Devices for capturing video / images can include at least one camera, video / image archives containing previously captured video / images, etc. Devices for generating video / images can include computers, tablets, smartphones, etc., and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc., and in this case, the process of capturing video / images can be replaced by the process of generating related data.

[0052] An encoding device can encode input video / images. The encoding device can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.

[0053] The transmitting unit can send encoded video / image information or data, output in bitstream form, to the receiving unit of the receiving device via digital storage media or a network, either as a file or through streaming. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating media files according to a predetermined file format and may include elements for transmission over broadcast / communication networks. The receiving unit can receive / extract the bitstream and send it to a decoding device.

[0054] Decoding devices can decode video / images by performing a series of processes such as inverse quantization, inverse transform, and prediction, which correspond to the operations of encoding devices.

[0055] The renderer can render decoded video / images. The rendered video / images can be displayed through a display unit.

[0056] Figure 2 A rough block diagram of an encoding apparatus that can be applied to embodiments of the present disclosure and perform encoding of video / image signals is shown.

[0057] refer to Figure 2The encoding device 200 may consist of an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transform 232, a quantizer 233, an inverse quantizer 234, and an inverse transform 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to an embodiment, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory 270 as an internal / external component.

[0058] Image segmenter 210 can partition an input image (or picture, frame) input to encoding device 200 into at least one processing unit. As an example, a processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively partitioned from the coding tree unit (CTU) or the largest coding unit (LCU) according to a quadtree-binary-trinary-tree (QTBTTT) structure.

[0059] For example, a coding unit can be segmented into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary structure can be applied later. Alternatively, a binary tree structure can be applied before the quadtree structure. The coding process according to this specification can be performed based on the final coding unit that is no longer segmented. In this case, based on coding efficiency according to image characteristics, the largest coding unit can be directly used as the final coding unit, or if necessary, the coding unit can be recursively segmented into deeper coding units, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding process can include processes such as prediction, transformation, and reconstruction, as described later.

[0060] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be divided or segmented from the aforementioned final encoding unit, respectively. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.

[0061] In some cases, a unit can be used interchangeably with terms such as block or region. Generally, an MxN block can represent a set of transform coefficients or samples consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Samples can be used as a term to correspond a picture (or image) to pixels or cells.

[0062] The encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from the inter-frame predictor 221 or the intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to the converter 232. In this case, the unit in the encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called the subtractor 231.

[0063] Predictor 220 can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a block of predictions including prediction samples for the current block. Predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction on a block or CU basis. Predictor 220 can generate various information about the prediction, such as prediction mode information, and send it to entropy encoder 240, as described later in the description of each prediction mode. The information about the prediction can be encoded in entropy encoder 240 and output as a bitstream.

[0064] Intra-predictor 222 can predict the current block by referencing samples within the current image. Depending on the prediction mode, the referenced samples can be located near the current block or at a distance from it. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The non-directional mode can include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional modes can include 33 or 65 directional modes. However, this is just an example, and more or fewer directional modes can be used depending on the configuration. Intra-predictor 222 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.

[0065] Inter-frame predictor 221 can derive a prediction block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes, and for example, for skip mode and merge mode, the inter-frame predictor 221 can use the motion information of neighboring blocks as the motion information of the current block. For skip mode, unlike merge mode, residual signals may not be sent. For motion vector prediction (MVP) mode, the motion vectors of surrounding blocks are used as motion vector predictors, and the motion vector difference is signaled to indicate the motion vector of the current block.

[0066] Predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as a combined intra-frame and inter-frame prediction (CIIP) mode. Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for prediction against a block. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) in games, etc. IBC essentially performs prediction within the current image, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current image. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. A palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, sample values ​​within the image can be signaled based on information about the palette table and palette index. The prediction signal generated by predictor 220 can be used to generate a reconstructed signal or a residual signal.

[0067] Transformer 232 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to the transform obtained from a graphic when the relationship information between pixels is expressed as a graphic. CNT refers to the transform obtained based on generating a prediction signal using all previously reconstructed pixels. Furthermore, the transform process can be applied to square pixel blocks of the same size or to non-square blocks of variable size.

[0068] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the quantized transform coefficients in block form into a one-dimensional vector form based on the coefficient scan order, and can generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.

[0069] The entropy encoder 240 can perform various encoding methods, such as exponential Columbus coding, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can encode information necessary for video / video image reconstruction (e.g., values ​​of syntax elements, etc.) in addition to the transform coefficients quantized together or individually.

[0070] Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the network abstraction layer (NAL) unit level. The video / image information may further include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Additionally, the video / image information may further include general constraint information. Here, information transmitted from the encoding device / signaled to the decoding device and / or syntax elements can be included in the video / image information. The video / image information can be encoded by the above-described encoding process and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmission units (not shown) for transmission and / or storage units (not shown) for storing signals output from the entropy encoder 240 can be configured as internal / external elements of the encoding device 200, or the transmission unit may also be included in the entropy encoder 240.

[0071] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients using dequantizer 234 and inverse transformer 235. Adder 250 can add the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed within the current image, and can also be used for inter-frame prediction of the next image by filtering, which will be described later. Meanwhile, a luminance mapping with chroma scaling (LMCS) can be applied during image encoding and / or reconstruction.

[0072] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods can include deblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various information about the filtering and send it to entropy encoder 240. The information about the filtering can be encoded in entropy encoder 240 and output as a bitstream.

[0073] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. When inter-frame prediction is applied through it, the encoding device can avoid prediction mismatch in encoding device 200 and decoding device, and can also improve encoding efficiency.

[0074] The DPB of memory 270 can store modified reconstructed images for use as reference images in inter-frame predictor 221. Memory 270 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of blocks in the pre-reconstructed image. The stored motion information can be sent to inter-frame predictor 221 to be used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 222.

[0075] Figure 3 A rough block diagram of a decoding apparatus that can be applied to embodiments of the present disclosure and perform decoding of video / image signals is shown.

[0076] refer to Figure 3 The decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 321.

[0077] According to an embodiment, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded image buffer (DPB) and can be configured by a digital storage medium. The hardware component may further include the memory 360 as an internal / external component.

[0078] When the input includes a bitstream containing video / image information, the decoding device 300 can respond to... Figure 2 The decoding device 300 reconstructs an image by processing video / image information in its encoding apparatus. For example, the decoding device 300 can derive units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 can perform decoding by using processing units applied in the encoding apparatus. Therefore, the processing unit for decoding can be an encoding unit, and the encoding unit can be segmented from encoding tree units or maximally encoded units according to a quadtree structure, binary tree structure, and / or ternary tree structure. At least one transform unit can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the decoding device 300 can be played back by a playback device.

[0079] Decoding device 300 can receive data in bitstream form from... Figure 2 The signal output by the encoding device and the received signal can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may further include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or the general constraint information. The information sent / received by the signal and / or the syntax elements described later herein can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, CABAC, etc., and output the values ​​of the syntax elements necessary for image reconstruction and the quantized values ​​of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine a context model using information about the syntax element to be decoded, decoding information of surrounding blocks and the block to be decoded, or information about symbols / bins decoded in the previous step, perform arithmetic decoding on the bins by predicting the occurrence probability of the bins based on the determined context model, and generate symbols corresponding to the value of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about the decoded symbols / bins for the context model used for the next symbol / bin. Among the information decoded in the entropy decoder 310, information about prediction is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values ​​of entropy decoding performed on them in the entropy decoder 310, i.e., the quantized transform coefficients and related parameter information, can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, information about filtering among the information decoded in the entropy decoder 310 can be provided to the filter 350. Meanwhile, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external element of the decoding device 300 or the receiving unit can be a component of the entropy decoder 310.

[0080] Furthermore, the decoding device according to this specification can be referred to as a video / image / picture decoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of an inverse quantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.

[0081] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order executed in the encoding device. The dequantizer 321 can obtain the transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).

[0082] The inverse transformer 322 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).

[0083] Predictor 320 can perform prediction on the current block and generate a prediction block including prediction samples for the current block. Predictor 320 can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from entropy decoder 310, and determine a specific intra-frame / inter-frame prediction mode.

[0084] Predictor 320 can generate prediction signals based on various prediction methods described later. For example, predictor 320 can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as a combined intra-frame and inter-frame prediction (CIIP) mode. Alternatively, the predictor can be based on an intra-block copy (IBC) prediction mode or a palette mode for block prediction. The IBC prediction mode or palette mode can be used for content image / video coding such as screen content coding (SCC) in games, etc. IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. Palette mode can be considered an example of intra-frame coding or intra-frame prediction. When a palette mode is applied, information about the palette table and palette index can be included in the video / image information and transmitted as a signal.

[0085] The intra-predictor 331 can predict the current block by referencing samples within the current image. Depending on the prediction mode, the referenced samples can be located near the current block or at a certain distance away from it. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The intra-predictor 331 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.

[0086] Inter-frame predictor 332 can derive a prediction block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information of neighboring blocks and the current block. Motion information may include motion vectors and a reference image index. Motion information may further include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and information about the prediction may include information indicating the inter-frame prediction mode used for the current block.

[0087] Adder 340 can add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331) to generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array). When there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the reconstruction block.

[0088] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, output by filtering as described later, or it can be used for inter-frame prediction of the next image. Meanwhile, a luminance map with chroma scaling (LMCS) can be applied during image decoding.

[0089] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and send the modified reconstructed image to memory 360, specifically the DPB of memory 360. Various filtering methods can include deblocking filtering, adaptive sampling offset, adaptive loop filtering, bilateral filtering, etc.

[0090] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame prediction unit 332. Memory 360 can derive (or decode) motion information of blocks from its current image and / or motion information of blocks in the pre-reconstructed image. The stored motion information can be sent to inter-frame predictor 332 as motion information for spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 331.

[0091] Here, the embodiments described in the encoding device 200’s filter 260, inter-frame predictor 221 and intra-frame predictor 222 can also be applied equally or correspondingly to the decoding device 300’s filter 350, inter-frame predictor 332 and intra-frame predictor 331, respectively.

[0092] Figure 4 The illustration shows a method for reconstructing video images performed by a decoding device 300 according to the present disclosure.

[0093] The S400 can receive bitstreams including encoded video images.

[0094] The S410 is a video image encoding that can reconstruct the bitstream.

[0095] Video information related to encoded video images can be extracted from the bitstream. The extracted video information can then be used to reconstruct the encoded video images.

[0096] The bitstream may include encoder optimization information (EOI). Encoder optimization information can be configured in supplementary enhancement information (SEI) messages, which will be referred to as EOI SEI messages. According to this disclosure, the EOI SEI message can be used to indicate whether the video is optimized for human or machine viewing. Additionally, the EOI SEI message can be used to indicate the type of optimization method applied to preprocessing or encoding.

[0097] The EOI SEI message may include an EOI cancellation flag (eoi_cancel_flag). When eoi_cancel_flag is 1, it indicates that the persistence of EOI SEI messages included in previous picture units (PUs) in output order is cancelled. When eoi_cancel_flag is 0, it indicates that optimization-related information applied during preprocessing or encoding follows. The optimization-related information included in the EOI SEI message will be described in detail below.

[0098] Optimization-related information may include the EOI persistence flag (eoi_persistence_flag). eoi_persistence_flag relates to the persistence of optimization information provided in the EOI SEI message. When eoi_persistence_flag is 0, it indicates that the optimization information is applied only to the current image. When eoi_persistence_flag is 1, it indicates that the optimization information is applied to the current image of the current layer and all subsequent images in output order.

[0099] Optimization-related information may include an EOI identifier for human viewing (eoi_for_human_viewing_idc). eoi_for_human_viewing_idc can be information used to identify the purpose and objective of the optimization. For example, a value of 3 for eoi_for_human_viewing_idc may indicate that the purpose of the optimization includes human viewing. A value of 2 for eoi_for_human_viewing_idc may indicate that the video is suitable for human viewing but is not optimized for human viewing. A value of 1 for eoi_for_human_viewing_idc may indicate that the video is not suitable for human viewing. A value of 0 for eoi_for_human_viewing_idc may indicate that whether the video is suitable for human viewing is unknown.

[0100] Optimization-related information may include an EOI identifier (eoi_for_machine_analysis_idc) used for machine analysis. eoi_for_machine_analysis_idc can be information used to identify the purpose and objective of the optimization. For example, a value of 3 for eoi_for_machine_analysis_idc may indicate that the purpose of the optimization includes machine analysis. A value of 2 for eoi_for_machine_analysis_idc may indicate that the video is suitable for machine analysis but is not optimized for it. A value of 1 for eoi_for_machine_analysis_idc may indicate that the video is not suitable for machine analysis. A value of 0 for eoi_for_machine_analysis_idc may indicate that whether the video is suitable for machine analysis is unknown.

[0101] You can restrict EOI SEI messages to include both eoi_for_human_viewing_idc and eoi_for_machine_analysis_idc, both of which have a value of 1.

[0102] Optimization-related information can include EOI type information (eoi_type). eoi_type can represent the type of optimization method. As an example, eoi_type can be defined as shown in Table 1 below.

[0103] [Table 1]

[0104] In Table 1, when (eoi_type & bitMask) is not 0, it indicates the type of optimization method applied corresponding to the bitmask value in Table 1. When eoi_type is greater than 0 and (eoi_type & bitMask) is 0, it indicates the type of optimization method not applied corresponding to the bitmask value. When eoi_type is 0, it indicates the use of the optimization determined by application.

[0105] Optimization-related information can include object-based type information (eoi_object_based_idc). eoi_object_based_idc can represent the object-based optimization type. As an example, eoi_object_based_idc can be defined as shown in Table 2.

[0106] [Table 2]

[0107] In Table 2, when (eoi_object_based_idc & bitMask) is not 0, it indicates that the object-based optimization type associated with the bitmask value in Table 2 is applied. When eoi_object_based_idc is greater than 0 and (eoi_object_based_idc & bitMask) is 0, it indicates that the object-based optimization type associated with the bitmask value is not applied. When eoi_object_based_idc is 0, it indicates that the object-based optimization type defined in the application is applied. The value of eoi_object_based_idc can be restricted to the range of 0 to 7. Values ​​of 8 to 65,535 for eoi_object_based_idc can be reserved for future use. When the value of eoi_object_based_idc is in the range of 8 to 65,535, decoders with specific specifications should ignore the corresponding eoi_object_based_idc value.

[0108] The `eoi_object_based_idc` can be signaled based on the first flag (`EoiObjectBasedFlag`). As an example, `eoi_object_based_idc` can be signaled based on a value of 1 for `EoiObjectBasedFlag`, and not signaled based on a value of 0 for `EoiObjectBasedFlag`. `EoiObjectBasedFlag` can specify whether `eoi_type` indicates that object-based optimizations will be included. The value of `EoiObjectBasedFlag` can be derived as shown in Equation 1 below.

[0109] [Formula 1]

[0110] Optimization-related information may include a temporal resampling type flag (eoi_temporal_resampling_type_flag). eoi_temporal_resampling_type_flag can indicate the type of temporal resampling optimization. As an example, eoi_temporal_resampling_type_flag can specify one of the predefined types of temporal resampling optimization. Predefined types of temporal resampling optimization can include at least one of subsampling or upsampling. When eoi_temporal_resampling_type_flag is 0, it indicates that the temporal resampling optimization is subsampling. When eoi_temporal_resampling_type_flag is 1, it indicates that the temporal resampling optimization is upsampling.

[0111] Optimization-related information can include image quantity information (eoi_num_int_pics). eoi_num_int_pics can be information related to the number of images excluded from or added between pairs of encoded images. For example, when eoi_num_int_pics is greater than 0, it can indicate (when eoi_temporal_resampling_type_flag is 0) that the number of images excluded by the encoding system between each pair of encoded images in output order, or (when eoi_temporal_resampling_type_flag is 1) that the number of images added by the encoding system between each pair of source images for encoding is constant. When eoi_temporal_resampling_type_flag is 0 and eoi_num_int_pics is greater than 0, eoi_num_int_pics can specify the number of images excluded by the encoding system between each pair of encoded images in output order. When eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_num_int_pics can specify the number of images that the encoding system adds between each pair of source images for encoding.

[0112] When `eoi_num_int_pics` is 0, it can indicate (when `eoi_temporal_resampling_type_flag` is 0) the number of pictures excluded by the encoding system between each pair of encoded pictures in output order, or (when `eoi_temporal_resampling_type_flag` is 1) the number of pictures added by the encoding system between each pair of source pictures for encoding, which is unknown or variable. The value of `eoi_num_int_pics` can be restricted to the range of 0 to 63.

[0113] At least one of the above-mentioned `eoi_temporal_resampling_type_flag` or `eoi_num_int_pics` can be signaled based on the second flag (`EoiTemporalResamplingFlag`). As an example, `eoi_temporal_resampling_type_flag` and `eoi_num_int_pics` can be signaled based on a value of 1 for `EoiTemporalResamplingFlag`, and `eoi_temporal_resampling_type_flag` and `eoi_num_int_pics` can be left unsigned based on a value of 0 for `EoiTemporalResamplingFlag`. `EoiTemporalResamplingFlag` can specify whether `eoi_type` represents temporal resampling optimization. The value of `EoiTemporalResamplingFlag` can be derived as shown in Equation 2 below.

[0114] [Equation 2]

[0115] Optimization-related information may include a spatial resampling type flag (eoi_spatial_resampling_type_flag). eoi_spatial_resampling_type_flag can be information used to identify the type of spatial resampling optimization. For example, when eoi_spatial_resampling_type_flag is 0, it can indicate that the spatial resampling optimization is subsampling. When eoi_spatial_resampling_type_flag is 1, it can indicate that the spatial resampling optimization is upsampling.

[0116] The `eoi_spatial_resampling_type_flag` can be signaled based on a third flag (`EoiSpatialResamplingFlag`). For example, `eoi_spatial_resampling_type_flag` can be signaled with a value of 1 for `EoiSpatialResamplingFlag`, and it can be left unsigned with a value of 0 for `EoiSpatialResamplingFlag`. `EoiSpatialResamplingFlag` can specify whether `eoi_type` indicates spatial resampling optimization. The value of `EoiSpatialResamplingFlag` can be derived as shown in Equation 3 below.

[0117] [Formula 3]

[0118] Optimization-related information may include a privacy protection type identifier (eoi_privacy_protection_type_idc). eoi_privacy_protection_type_idc can be information used to identify the type of privacy protection optimization. As an example, eoi_privacy_protection_type_idc can be defined as shown in Table 3.

[0119] [Table 3]

[0120] Optimization-related information can include privacy protection type information (eoi_privacy_protected_info_type). eoi_privacy_protected_info_type can represent the type of information being protected. As an example, eoi_privacy_protected_info_type can be defined as shown in Table 4 below.

[0121] [Table 4]

[0122] In Table 4, when `eoi_privacy_protected_info_type` is greater than 0 and `(eoi_privacy_protected_info_type & bitMask)` is not 0, it indicates that the information type corresponding to the bitmask value in Table 4 is protected. When `eoi_privacy_protected_info_type` is 0, it indicates that information of a type defined in the application is protected. The value of `eoi_privacy_protection_info_type` can be restricted to the range of 0 to 7. Values ​​of 8 to 255 for `eoi_privacy_protected_info_type` are reserved for future use, and decoders with specific specifications should ignore the corresponding `eoi_privacy_protected_info_type` when its value falls within the range of 8 to 255.

[0123] At least one of the aforementioned `eoi_privacy_protection_type_idc` or `eoi_privacy_protected_info_type` can be signaled based on the fourth flag (Eoi Privacy Protection Flag). As an example, `eoi_privacy_protection_type_idc` and `eoi_privacy_protected_info_type` can be signaled based on a value of 1 for `EoiPrivacyProtectionFlag`, and can be left unsigned based on a value of 0 for `EoiPrivacyProtectionFlag`. `EoiPrivacyProtectionFlag` can specify whether `eoi_type` represents privacy protection optimization. The value of `EoiPrivacyProtectionFlag` can be derived as shown in Equation 4 below.

[0124] [Formula 4]

[0125] When the eoi_persistence_flag of the EOI SEI message is 0, the value of EoiTemporalResamplingFlag can be restricted to 0.

[0126] As mentioned above, the EOI SEI message may include information indicating that temporal upsampling was applied before / during the encoding process. However, the EOI SEI message does not include information for the decoder / receiver of the bitstream to distinguish between the source image (i.e., the image used as input to the temporal upsampling process) and the additional images generated during the temporal upsampling process of the bitstream. This information can be useful when it is necessary to remove one or more images from the bitstream. This is because, in image removal, it is preferable to remove the additional images rather than the source images. The method for distinguishing between the source images and the additional images will be described below.

[0127] Example 1

[0128] Optimization-related information may include source image identification information (eoi_dist_from_src_pic). eoi_dist_from_src_pic can be related to whether the current image is a source image used for temporal upsampling optimization. When eoi_dist_from_src_pic is the first value, it indicates that the current image is a source image used for temporal upsampling optimization. On the other hand, when eoi_dist_from_src_pic is not the first value, it may not indicate that the current image is a source image used for temporal upsampling optimization.

[0129] As an example, when `eoi_dist_from_src_pic` is greater than 0, it can represent the distance from the last source image preceding the current image in output order (e.g., an image used as input for temporal upsampling or a newly generated image not resulting from the temporal upsampling process) to the current image as the number of images. In other words, when `eoi_dist_from_src_pic` is not 0, it can indicate that the current image is not a source image used for temporal upsampling optimization. On the other hand, when `eoi_dist_from_src_pic` is 0, it can indicate that the current image is a source image used for temporal upsampling optimization. The value of `eoi_dist_from_src_pic` can be restricted to the range of 0 to `eoi_num_int_pics`.

[0130] When temporal upsampling optimization is applied and the number of images to be added is available, the distance from the source image to the current image can be used to determine the source image and the images to be added in the temporal upsampling optimization.

[0131] A predefined flag (SrcPicFlag[i]) can be derived based on eoi_dist_from_src_pic. SrcPicFlag[i] indicates whether the i-th image within the persistence range of the EOI SEI message is the source image. As an example, SrcPicFlag[i] can be derived as shown in Table 5 below. In Table 5, NumPics represents the number of images within the persistence range of the EOI SEI message.

[0132] [Table 5]

[0133] As an example, eoi_dist_from_src_pic can be configured to be sent by a signal, as shown in Table 6 below.

[0134] [Table 6]

[0135] The `eoi_dist_from_src_pic` can be signaled based on at least one of `eoi_temporal_resampling_type_flag` or `eoi_num_int_pics`. Specifically, `eoi_dist_from_src_pic` can be signaled based on a value of 1 for `eoi_temporal_resampling_type_flag` and a value greater than 0 for `eoi_num_int_pics`. Conversely, `eoi_dist_from_src_pic` can be signaled without signaling based on a value of 0 for `eoi_temporal_resampling_type_flag` or a value not greater than 0 for `eoi_num_int_pics`.

[0136] Optimization-related information may include a source image presence flag (eoi_dist_from_src_pic_present_flag). In such cases, eoi_dist_from_src_pic can be signaled based on eoi_dist_from_src_pic_present_flag. eoi_dist_from_src_pic_present_flag indicates whether the source image identifier (eoi_dist_from_src_pic) exists in the EOI SEI message. As an example, when eoi_dist_from_src_pic_present_flag is 1, it indicates that eoi_dist_from_src_pic exists in the EOI SEI message. When eoi_dist_from_src_pic_present_flag is 0, it indicates that eoi_dist_from_src_pic does not exist in the EOI SEI message.

[0137] The `eoi_dist_from_src_pic` function can be signaled when `eoi_dist_from_src_pic_present_flag` is 1, and can be unsigned when `eoi_dist_from_src_pic_present_flag` is 0. When `eoi_dist_from_src_pic_present_flag` is 0, the distance from the source image to the current image can be unknown.

[0138] Even when eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, eoi_dist_from_src_pic can be restricted to being sent by a signal when eoi_dist_from_src_pic_present_flag is 1.

[0139] The `eoi_dist_from_src_pic_present_flag` can be signaled if `eoi_temporal_resampling_type_flag` is 1 and `eoi_num_int_pics` is greater than 0. Conversely, the `eoi_dist_from_src_pic_present_flag` can be refrained from being signaled if `eoi_temporal_resampling_type_flag` is 0 or `eoi_num_int_pics` is not greater than 0.

[0140] Example 2

[0141] Optimization-related information may include source image identification information (eoi_dist_from_src_pic). eoi_dist_from_src_pic may be related to whether the current image is the source image used for temporal upsampling optimization.

[0142] `eoi_dist_from_src_pic` can be represented as the distance from the last source image (e.g., an image used as input for temporal upsampling or a newly generated image not resulting from the temporal upsampling process) to the current image in output order. When `eoi_dist_from_src_pic` is greater than 0, the distance between the last source image and the current image can be defined as `(eoi_dist_from_src_pic%(eoi_num_int_pics+1))`. Here, `eoi_num_int_pics` is information related to the number of images excluded or added between the encoded images, which will be described in detail later. Otherwise (i.e., when `eoi_dist_from_src_pic` is 0), the distance from the source image to the current image is unknown. The value of `eoi_dist_from_src_pic` can be restricted to the range of 0 to `(eoi_num_int_pics+1)`.

[0143] When temporal upsampling optimization is applied and the number of images to be added is available, the distance from the source image to the current image can be used to determine the source image and the images to be added in the temporal upsampling optimization.

[0144] When the value of eoi_dist_from_src_pic is (eoi_num_int_pics+1), the current image can correspond to the source image.

[0145] A predefined flag (SrcPicFlag[i]) can also be derived based on eoi_dist_from_src_pic, which is the same as described in reference Table 5.

[0146] As an example, eoi_dist_from_src_pic can also be configured to be sent by a signal, as shown in Table 7 below.

[0147] [Table 7]

[0148] The eoi_dist_from_src_pic can be signaled based on at least one of eoi_temporal_resampling_type_flag or eoi_num_int_pics.

[0149] Specifically, `eoi_dist_from_src_pic` can be signaled if `eoi_temporal_resampling_type_flag` is 1 and `eoi_num_int_pics` is greater than 0. Alternatively, `eoi_dist_from_src_pic` can be signaled without a signal if `eoi_temporal_resampling_type_flag` is 0 or `eoi_num_int_pics` is not greater than 0.

[0150] Example 3

[0151] Optimization-related information may include the temporal resampling type flag (eoi_temporal_resampling_type_flag). eoi_temporal_resampling_type_flag can indicate the type of temporal resampling optimization.

[0152] Optimization-related information may include information related to the number of images (eoi_num_int_pics). eoi_num_int_pics can be information related to the number of images excluded or added between the encoded image pairs.

[0153] As an example, when eoi_num_int_pics is greater than 0, it can mean (when eoi_temporal_resampling_type_flag is 0) that the number of images excluded by the encoding system between each pair of encoded images in output order, or (when eoi_temporal_resampling_type_flag is 1) that the number of images added by the encoding system between each pair of source images for encoding is constant.

[0154] When `eoi_temporal_resampling_type_flag` is 0 and `eoi_num_int_pics` is greater than 0, `eoi_num_int_pics` can specify the number of pictures excluded by the encoding system between each pair of encoded pictures in output order. When `eoi_temporal_resampling_type_flag` is 1 and `eoi_num_int_pics` is greater than 0, `eoi_num_int_pics` can specify the number of pictures added by the encoding system between each pair of source pictures for encoding.

[0155] When `eoi_num_int_pics` is 0, it can indicate (when `eoi_temporal_resampling_type_flag` is 0) the number of images excluded by the encoding system between each pair of encoded images in output order, or (when `eoi_temporal_resampling_type_flag` is 1) the number of images added by the encoding system between each pair of source images for encoding is unknown or variable. The value of `eoi_num_int_pics` can be restricted to the range of 0 to 63.

[0156] However, when eoi_temporal_resampling_type_flag is 1 and eoi_num_int_pics is greater than 0, EOI SEI messages can be restricted to accessing units of images other than those added / generated during the temporal sampling process.

[0157] The encoder optimization information according to this disclosure can be configured in the SEI message of the bitstream. The SEI message can be included in the Network Abstraction Layer (NAL) unit of the bitstream. However, it is not limited thereto. As an example, the encoder optimization information according to this disclosure can be configured in the high-level syntax of the bitstream. Here, the high-level syntax can be at least one of Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header (PH), or Slice Header (SH). Alternatively, the encoder optimization information according to this disclosure can also be defined as a separate NAL unit type within the bitstream.

[0158] Figure 5 The illustration shows a schematic configuration of a decoding apparatus 300 that performs a method for reconstructing video images according to the present disclosure.

[0159] refer to Figure 5 The decoding device 300 may include a receiver 500, a video information extractor 510, and a video reconstructor 520.

[0160] Receiver 500 can receive bitstreams including encoded video images.

[0161] Video information extractor 510 can extract video information related to encoded video images from the bitstream. Additionally, video information extractor 710 can extract EOI and SEI messages from the bitstream, which are then compared with information obtained through reference... Figure 4 The description is the same.

[0162] The Video Reconstructor 520 can reconstruct encoded video images based on extracted video information.

[0163] Figure 6 The figure illustrates a method for generating a bitstream performed by an encoding device 200 according to the present disclosure.

[0164] The S600 can receive video images to be encoded.

[0165] The received video images can be encoded to generate video information associated with the video images S610.

[0166] It can generate a bitstream S420 that includes video information related to video images.

[0167] Additionally, an EOI SEI message can be generated that is applied to the bitstream, which is related to the reference... Figure 4 The description is the same. The generated EOI SEI message can be included in the bitstream.

[0168] Figure 7 The figure shows a schematic configuration of an encoding apparatus 200 for performing a method for generating a bitstream according to the present disclosure.

[0169] refer to Figure 7 The encoding device 200 may include a receiver 700, a video compressor 710, and a bitstream generator 720.

[0170] Receiver 700 can receive one or more video images to be encoded.

[0171] The video compressor 710 can encode one or more received video images to generate video information associated with the video images. The video compressor 710 can generate EOI and SEI messages that are applied to the bitstream.

[0172] Bitstream generator 720 can generate a bitstream that includes video information. Bitstream generator 720 can further generate a bitstream that includes the generated EOI SEI messages.

[0173] In the above embodiments, the method is described as a series of steps or blocks based on the flowchart. However, the corresponding embodiments are not limited to the order of the steps, and some steps may occur simultaneously or in a different order than the other steps described above. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and other steps may be included or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this disclosure.

[0174] The methods described above according to embodiments of the present disclosure can be implemented in software, and the encoding and / or decoding apparatus according to the present disclosure can be included in devices performing image processing, such as TVs, computers, smartphones, set-top boxes, display devices, etc.

[0175] In this disclosure, when the embodiments are implemented as software, the above methods can be implemented as modules (processes, functions, etc.) performing the above functions. Modules can be stored in memory and can be executed by a processor. Memory can be located inside or outside the processor and can be connected to the processor by various well-known means. The processor may include application-specific integrated circuits (ASICs), another chipset, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, the embodiments described herein can be implemented on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.

[0176] Furthermore, the decoding and encoding devices using embodiments of this disclosure can be included in multimedia broadcasting transmitting and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, devices for providing video-on-demand (VoD) services, over-the-top (OTT) devices, devices for providing internet streaming services, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, videophone video devices, transportation terminal devices (e.g., vehicle (including autonomous vehicles) terminals, aircraft terminals, ship terminals, etc.), and medical video devices, and can be used to process video signals or data signals. For example, over-the-top (OTT) devices can include game consoles, Blu-ray players, networked TVs, home theater systems, smartphones, tablets, digital video recorders (DVRs), etc.

[0177] Furthermore, the processing methods applying embodiments of this disclosure can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having data structures according to embodiments of this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical media storage devices. Additionally, computer-readable recording media include media implemented in carrier wave form (e.g., transmitted via the Internet). Furthermore, bitstreams generated by encoding methods can be stored in a computer-readable recording medium or transmitted via wired or wireless communication networks.

[0178] Furthermore, the embodiments of this disclosure can be implemented by a computer program product using program code, and this program code can be executed on a computer by the embodiments of this disclosure. The program code can be stored on a computer-readable medium.

[0179] Figure 8 Examples of content streaming systems to which embodiments of the present disclosure may be applied are shown.

[0180] refer to Figure 8 The content streaming system using embodiments of this disclosure may mainly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.

[0181] An encoding server generates a bitstream by compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, and then sends it to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders generate bitstreams directly, the encoding server can be omitted.

[0182] A bitstream can be generated by applying the encoding method or bitstream generation method of the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.

[0183] A streaming server sends multimedia data to a user's device via a web server based on the user's request, and the web server acts as a medium to notify the user what services are available. When a user requests a service from the web server, the web server delivers it to the streaming server, and the streaming server sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server, which in this case controls the commands / responses between each device in the content streaming system.

[0184] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store a bitstream for a certain period of time.

[0185] Examples of user equipment may include mobile phones, smartphones, laptops, digital broadcast terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs), digital TVs, desktop computers, digital signage, etc.).

[0186] In a content streaming system, each server can be operated as a distributed server, and in this case, data received from each server can be distributed and processed.

[0187] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined and implemented as a device, and the technical features of the device claims of this disclosure can be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this disclosure can be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this disclosure can be combined and implemented as a method.

Claims

1. A method comprising: Receive bitstreams including encoded video images; as well as Reconstruct the encoded video images included in the bitstream. The bitstream is configured to include encoder optimization information (EOI) supplemental enhancement information (SEI) messages. The EOI SEI message includes source image identification information related to whether the current image is a source image used for temporal upsampling optimization, and The EOI SEI message is obtained from the Network Abstraction Layer (NAL) unit of the bitstream.

2. The method according to claim 1, wherein, The first value of the source image identification information indicates that the current image is the source image used for the temporal upsampling optimization, and Wherein, the source image identification information of values ​​other than the first value does not indicate that the current image is the source image used for the temporal upsampling optimization.

3. The method according to claim 1, wherein, The source image identification information is transmitted by signaling based on at least one of the following: a time resampling type flag or image quantity related information. Wherein, the time resampling type flag indicates the type of time resampling optimization, and The image quantity information refers to the number of images excluded from or added between encoded image pairs by the encoding system.

4. The method according to claim 3, wherein, The time resampling type flag with a value of 0 indicates that the time resampling optimization is subsampling, and The time resampling type flag with a value of 1 indicates that the time resampling optimization is upsampling.

5. The method according to claim 4, wherein, Based on the fact that the value of the time resampling type flag is 0 and the value of the image quantity related information is greater than 0, the source image identification information is sent by signal.

6. The method according to claim 5, wherein, The EOI SEI message further includes a source image presence flag indicating whether the source image identification information exists in the EOI SEI message.

7. The method according to claim 6, wherein, Based on the presence flag of the source image, indicating that the source image identification information exists in the EOI SEI message, the source image identification information is sent by a signal.

8. The method according to claim 1, wherein, The EOI SEI message includes at least one of the following: a time resampling type flag or information related to the number of images. Wherein, the time resampling type flag indicates the type of time resampling optimization, and The image quantity information refers to the number of images excluded from or added between encoded image pairs by the encoding system.

9. The method according to claim 8, wherein, When the temporal resampling type flag indicates that the temporal resampling optimization is upsampling, and the value of the image quantity related information is greater than 0, the EOI SEI message is restricted to the access unit of images other than those added during the temporal upsampling process.

10. A method comprising: Receive the video and images to be encoded; The received video images are encoded to generate video information related to the video images; Generate encoder optimization information (EOI) supplemental enhancement information (SEI) messages; as well as Generate a bitstream including the video information and the EOI SEI message. The EOI SEI message includes source image identification information related to whether the current image is a source image used for temporal upsampling optimization, and The EOI SEI message is encoded into the Network Abstraction Layer (NAL) unit of the bitstream.

11. A computer-readable storage medium for storing a bit stream generated by the method of claim 10.

12. A method comprising: Generate a bitstream, wherein the bitstream is generated based on: receiving a video image to be encoded, encoding the received video image to generate video information associated with the video image, and generating encoder optimization information (EOI) supplementary enhancement information (SEI) messages; and Send data including the bit stream. The EOI SEI message includes source image identification information related to whether the current image is a source image used for temporal upsampling optimization, and The EOI SEI message is encoded into the Network Abstraction Layer (NAL) unit of the bitstream.