Video encoding / decoding method and apparatus based on supplemental enhancement information (SEI) message, and recording medium for storing bitstream
By introducing Source Picture Timing Information (SPTI) SEI messages into image encoding and decoding, the intra-frame and inter-frame prediction modes are optimized, solving the problem of high transmission and storage costs in high-resolution image encoding and achieving more efficient encoding and decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2024-10-04
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies suffer from high transmission and storage costs due to increased information volume during the encoding and decoding of high-resolution and high-quality images, and lack effective intra-frame prediction modes and supplementary enhancement information (SEI) messages.
Image encoding and decoding methods are optimized by introducing Source Picture Timing Information (SPTI) SEI messages, including intra-frame prediction mode and inter-frame prediction mode, and SPTI SEI messages are embedded in the bitstream to improve encoding efficiency.
It improves the efficiency of image encoding and decoding, supports intra-frame and inter-frame prediction modes, reduces transmission and storage costs, and provides an encoding scheme based on SEI messages.
Smart Images

Figure CN121970347A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an image encoding / decoding method and apparatus, and a recording medium for storing bitstreams, and more specifically, to an image encoding / decoding method and apparatus based on supplementary enhancement information (SEI) messages, and a recording medium for storing bitstreams generated by the image encoding method / apparatus of this disclosure. Background Technology
[0002] Recently, there has been an increasing demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, across various fields. As the resolution and quality of image data improve, the amount of information or bits transmitted increases relative to existing image data. This increase in the amount of information or bits transmitted leads to increased transmission and storage costs.
[0003] Therefore, there is a need for efficient image compression techniques to effectively send, store, and reproduce information about high-resolution and high-quality images. Summary of the Invention
[0004] Technical issues
[0005] The purpose of this disclosure is to provide an image encoding / decoding method and apparatus with improved encoding / decoding efficiency.
[0006] Another objective of this disclosure is to provide an image encoding / decoding method and apparatus for performing intra-frame prediction modes.
[0007] Another objective of this disclosure is to provide an image encoding / decoding method and apparatus for performing inter-frame prediction modes.
[0008] Another objective of this disclosure is to provide an image encoding / decoding method and apparatus based on Supplemental Enhancement Information (SEI) messages.
[0009] Another objective of this disclosure is to provide an image encoding / decoding method and apparatus that includes source picture timing information (SPTI) SEI messages in a bitstream.
[0010] Another object of this disclosure is to provide a non-transitory computer-readable recording medium for storing a bit stream generated by an image encoding method or apparatus according to this disclosure.
[0011] Another object of this disclosure is to provide a non-transitory computer-readable recording medium for storing a bitstream that is received and decoded by an image decoding apparatus according to this disclosure and used for image reconstruction.
[0012] Another object of this disclosure is to provide a method for transmitting a bit stream generated by an image encoding method or apparatus according to this disclosure.
[0013] The technical objectives to be achieved by this disclosure are not limited to those mentioned above, and those skilled in the art will clearly understand from the following description other technical objectives not mentioned herein.
[0014] Technical solution
[0015] According to embodiments of this disclosure, an image decoding method performed by an image decoding device may include: receiving a bitstream including image information; and generating a reconstructed image by reconstructing a current image based on the image information, wherein the image information may include a Source Image Timing Information (SPTI) SEI message related to a time interval between source images, the source image being associated with a corresponding decoded image of the source image.
[0016] According to embodiments of this disclosure, the SPTI SEI message may include specifying whether the SPTI SEI message applies only to the persistence information of the currently decoded image.
[0017] According to embodiments of this disclosure, the SPTI SEI message may include at least one of the following: first information specifying the maximum number of time sub-layers, second information specifying the scaling factor used in determining the time interval between source images, or third information specifying whether to synthesize the currently decoded image included in the time sub-layer.
[0018] According to embodiments of this disclosure, first information, second information, and third information can be obtained based on persistent information.
[0019] According to embodiments of this disclosure, first information, second information, and third information can be obtained based on the persistence information of a specified SPTI SEI message that is not only applied to the currently decoded image.
[0020] According to embodiments of this disclosure, second and third information can be further obtained based on the value of the first information.
[0021] According to embodiments of this disclosure, since the second information is not obtained, the value of the second information can be inferred to be equal to 1.
[0022] According to embodiments of this disclosure, based on the persistence information that a specified SPTI SEI message is applied only to the currently decoded image, the source picture interval can be set to be equal to the elemental source picture interval.
[0023] According to embodiments of this disclosure, an image encoding method performed by an image encoding device may include: generating image information associated with a current image, and encoding a bitstream including the image information, wherein the image information may include a Source Image Timing Information (SPTI) SEI message associated with a time interval between source images and a corresponding decoded image of the source image.
[0024] According to embodiments of this disclosure, a computer-readable recording medium for storing a bitstream generated by an image encoding method can be provided.
[0025] According to embodiments of this disclosure, a method for transmitting a bitstream generated by an image encoding method may include: generating image information associated with a current image; and encoding the bitstream including the image information, wherein the image information may include a Source Image Timing Information (SPTI) SEI message associated with a time interval between source images and a corresponding decoded image of the source image.
[0026] Beneficial effects
[0027] According to this disclosure, an image encoding / decoding method and apparatus with improved encoding / decoding efficiency can be provided.
[0028] According to this disclosure, an image encoding / decoding method and apparatus for performing intra-frame prediction modes can be provided.
[0029] According to this disclosure, an image encoding / decoding method and apparatus for performing inter-frame prediction modes can be provided.
[0030] According to this disclosure, an image encoding / decoding method and apparatus based on Supplemental Enhancement Information (SEI) messages can be provided.
[0031] According to this disclosure, an image encoding / decoding method and apparatus that includes Source Picture Timing Information (SPTI) SEI messages in a bitstream can be provided.
[0032] According to this disclosure, a non-transitory computer-readable recording medium may be provided for storing a bit stream generated by an image encoding method or apparatus according to this disclosure.
[0033] According to this disclosure, a non-transitory computer-readable recording medium may be provided for storing a bitstream that is received and decoded by an image decoding apparatus according to this disclosure and used for image reconstruction.
[0034] According to this disclosure, a method for transmitting a bit stream generated by an image encoding method or apparatus according to this disclosure can be provided.
[0035] The effects that can be obtained from this disclosure are not limited to those mentioned above, and those skilled in the art will clearly understand from the following description other effects not mentioned herein. Attached Figure Description
[0036] Figure 1 This is a view that schematically illustrates an example of a video coding system to which embodiments of this disclosure are applicable.
[0037] Figure 2 This is a schematic view illustrating an image encoding apparatus to which embodiments of the present disclosure are applicable.
[0038] Figure 3 This is a schematic view illustrating an image decoding apparatus to which embodiments of the present disclosure are applicable.
[0039] Figure 4 The illustration shows an example of a coding layer and structure to which embodiments of this disclosure apply.
[0040] Figure 5 This is a flowchart of the image encoding method according to the present disclosure.
[0041] Figure 6 This is a flowchart of the image decoding method according to the present disclosure.
[0042] Figure 7 This is an exemplary view showing a content streaming system to which embodiments of this disclosure are applicable. Detailed Implementation
[0043] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings to facilitate implementation by those skilled in the art. However, this disclosure can be implemented in various different forms and is not limited to the embodiments described herein.
[0044] In describing this disclosure, detailed descriptions of relevant known functions or constructions will be omitted if they unnecessarily obscure the scope of the disclosure. In the accompanying drawings, portions irrelevant to the description of this disclosure are omitted, and similar reference numerals are appended to similar portions.
[0045] In this disclosure, when a component is "connected," "coupled," or "linked" to another component, it may include not only direct connections but also indirect connections where intermediate components exist. Furthermore, when a component "comprises" or "has" other components, unless otherwise stated, it means that other components may be further included, rather than excluded.
[0046] In this disclosure, the terms first, second, etc., are used only for the purpose of distinguishing one component from other components and do not limit the order or importance of the components unless otherwise stated. Accordingly, within the scope of this disclosure, a first component in one embodiment may be referred to as a second component in another embodiment, and similarly, a second component in one embodiment may be referred to as a first component in another embodiment.
[0047] In this disclosure, the distinguishing components are intended to clearly describe each feature and do not imply that the components must be separate. That is, multiple components may be integrated and implemented in a single hardware or software unit, or a single component may be distributed and implemented in multiple hardware or software units. Therefore, unless otherwise specified, embodiments in which components are integrated or distributed are included within the scope of this disclosure.
[0048] In this disclosure, the components described in the various embodiments are not necessarily essential components, and some components may be optional. Therefore, embodiments comprising a subset of the components described in the embodiments are also included within the scope of this disclosure. Furthermore, embodiments that include other components besides those described in the various embodiments are also included within the scope of this disclosure.
[0049] This disclosure relates to the encoding and decoding of images, and unless redefined in this disclosure, the terms used herein may have the general meaning commonly used in the art to which this disclosure pertains.
[0050] In this disclosure, “video” can mean a collection of images that occur over time.
[0051] In this disclosure, "picture" generally refers to the basis of an image within a specific time period, and a slice / tile is a coding unit that constitutes a part of a picture. A picture may consist of one or more slices / tiles. In addition, slices / tiles may include one or more coding tree units (CTUs).
[0052] In this disclosure, "pixel" or "cell" can refer to the smallest unit that constitutes a picture (or image). Additionally, "sample" can be used as a term corresponding to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0053] In this disclosure, "unit" can refer to a basic unit of image processing. A unit may include at least one of a specific region of an image and information associated with that region. In some cases, the term "unit" may be used interchangeably with terms such as "sample array," "block," or "area." Generally, an M×N block may include a set (or array) of samples (or sample arrays) with M columns and N rows, or a set (or array) of transform coefficients.
[0054] In this disclosure, "current block" can mean one of "current coding block," "current coding unit," "coding target block," "decoding target block," or "processing target block." When performing prediction, "current block" can mean "current prediction block" or "prediction target block." When performing transform (inverse transform) / quantization (dequantization), "current block" can mean "current transform block" or "transform target block." When performing filtering, "current block" can mean "filter target block."
[0055] Additionally, in this disclosure, unless explicitly stated as a chroma block, "current block" may mean a block that includes both luma component blocks and chroma component blocks, or "the luma block of the current block." The luma component block of the current block can be expressed by an explicit description including luma component blocks such as "luma block" or "current luma block." Similarly, "the chroma component block of the current block" can be expressed by an explicit description including chroma component blocks such as "chroma block" or "current chroma block."
[0056] In this disclosure, the terms “ / ” or “,” can be interpreted as indicating “and / or”. For example, the expressions “A / B” and “A, B” can mean “A and / or B”. Furthermore, “A / B / C” and “A, B, C” can mean “at least one of A, B and / or C”.
[0057] In this disclosure, the term "or" should be interpreted to indicate "and / or". For example, expressing "A or B" can include 1) only "A", 2) only "B", or 3) both "A and B". In other words, in this disclosure, the term "or" should be interpreted to indicate "additionally or alternatively".
[0058] In this disclosure, "at least one of A, B, and C" can mean "only A", "only B", "only C", or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0059] As used in this disclosure, parentheses may mean "for example". For example, "prediction (intra-prediction)" may mean that "intra-prediction" is presented as an example of "prediction". In other words, "prediction" in this disclosure is not limited to "intra-prediction", and "intra-prediction" may be presented as an example of "prediction". In addition, even "prediction (i.e., intra-prediction)" may be constructed to present "intra-prediction" as an example of "prediction".
[0060] Overview of Video Encoding Systems
[0061] Figure 1 This is a view that schematically illustrates an example of a video coding system to which embodiments of this disclosure are applicable.
[0062] The video encoding system according to the embodiment may include an encoding device 10 and a decoding device 20. The encoding device 10 may deliver encoded video and / or image information or data to the decoding device 20 in the form of files or streams via a digital storage medium or network.
[0063] The encoding apparatus 10 according to an embodiment may include a video source generator 11, an encoding unit (also referred to as an encoder) 12, and a transmitter 13. The decoding apparatus 20 according to an embodiment may include a receiver 21, a decoding unit (also referred to as a decoder) 22, and a renderer 23. The encoding unit 12 may be referred to as a video / image encoding apparatus, and the decoding unit 22 may be referred to as a video / image decoding apparatus. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0064] The video source generator 11 can acquire video / images through a process of capturing, compositing, or generating video / images. The video source generator 11 may include a video / image capture device and / or a video / image generation device. The video / image capture device may include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device may include, for example, a computer, tablet computer, and smartphone, and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc. In this case, the video / image capture process can be replaced by a process of generating related data.
[0065] The encoding unit 12 can encode the input video / image. For compression and encoding efficiency, the encoding unit 12 can perform a series of processes such as prediction, transformation, and quantization. The encoding unit 12 is capable of outputting the encoded data (encoded video / image information) in the form of a bitstream.
[0066] Transmitter 13 can acquire encoded video / image information or data output as a bitstream and forward it to receiver 21 of decoding device 20 or another external object via digital storage medium or network as file or streaming data. Digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. Transmitter 13 can include elements for generating media files in a predetermined file format and may include elements for transmission via broadcast / communication networks. Transmitter 13 can be provided as a transmission device separate from encoding device 12, and in this case, the transmission device can include at least one processor that acquires encoded video / image information or data output as a bitstream; and a transmission unit for sending it as file or streaming data. Receiver 21 can extract / receive bitstreams from storage medium or network and send the bitstreams to decoding unit 22.
[0067] The decoding unit 22 can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of the encoding unit 12.
[0068] Renderer 23 can render decoded video / images. The rendered video / images can be displayed on a monitor.
[0069] Overview of image encoding devices
[0070] Figure 2 This is a view schematically illustrating an image encoding apparatus to which embodiments of the present disclosure may be applied.
[0071] like Figure 2 As shown, the image encoding apparatus 100 may include an image partitioner 110, a subtractor 115, a transformer 120, a quantizer 130, a dequantizer 140, an inverse transformer 150, an adder 155, a filter 160, a memory 170, an inter-frame prediction unit 180, an intra-frame prediction unit 185, and an entropy encoder 190. The inter-frame prediction unit 180 and the intra-frame prediction unit 185 may be collectively referred to as "prediction units". The transformer 120, quantizer 130, dequantizer 140, and inverse transformer 150 may be included in a residual processor. The residual processor may further include a subtractor 115.
[0072] In some embodiments, all or at least some of the components of the image encoding apparatus 100 may be configured by a single hardware component (e.g., an encoder or a processor). Additionally, the memory 170 may include a decoded image buffer (DPB) and may be configured by a digital storage medium.
[0073] Image partitioner 110 can partition an input image (or picture or frame) input to image encoding device 100 into one or more processing units. For example, a processing unit may be referred to as a coding unit (CU). A coding unit can be obtained by recursively partitioning a coding tree unit (CTU) or a maximum coding unit (LCU) according to a quadtree / binary tree / tritree (QT / BT / TT) structure. For example, a coding unit can be partitioned into multiple coding units of greater depth based on a quadtree structure, a binary tree structure, and / or a ternary tree structure. For the partitioning of a coding unit, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. The encoding process according to embodiments of the present disclosure can be performed based on the final coding unit that is no longer partitioned. The maximum coding unit can be used as the final coding unit, or a deeper coding unit obtained by partitioning the maximum coding unit can be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes. As another example, the processing unit of the encoding process may be a prediction unit (PU) or a transformation unit (TU). Prediction and transform units can be segmented or partitioned from the final coding unit. The prediction unit can be a unit for predicting samples, and the transform unit can be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.
[0074] The prediction unit (inter-frame prediction unit 180 or intra-frame prediction unit 185) can perform prediction on the block to be processed (the current block) and generate a prediction block including prediction samples for the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. The prediction unit can generate various information related to the prediction of the current block and send the generated information to the entropy encoder 190. The information about the prediction can be encoded in the entropy encoder 190 and output as a bitstream.
[0075] Intra-prediction unit (intra-predictor) 185 can predict the current block by referencing samples in the current image. Depending on the intra-prediction mode and / or intra-prediction technique, the reference samples may be located among the neighbors of the current block or may be placed separately. Intra-prediction modes may include multiple non-directional modes and multiple directional modes. Non-directional modes may include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, directional modes may include, for example, 33 or 65 directional prediction modes. However, this is merely an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-prediction unit 185 can determine the prediction mode to be applied to the current block by using prediction modes applied to neighboring blocks.
[0076] The inter-frame prediction unit (inter-frame predictor) 180 can derive a prediction block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc. The reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, the inter-frame prediction unit 180 can construct a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame prediction unit 180 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be transmitted. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be transmitted by signaling an encoded motion vector difference and an indicator for the motion vector predictor. The motion vector difference can refer to the difference between the motion vector of the current block and the motion vector predictor.
[0077] The prediction unit can generate a prediction signal based on various prediction methods and techniques described below. For example, the prediction unit can apply not only intra-frame prediction or inter-frame prediction, but also both intra-frame prediction and inter-frame prediction simultaneously to predict the current block. A prediction method that simultaneously applies intra-frame prediction and inter-frame prediction to predict the current block can be called combined intra-frame and inter-frame prediction (CIIP). Additionally, the prediction unit can perform intra-block copying (IBC) on the prediction of the current block. Intra-block copying can be used for content image / video coding in games, such as Screen Content Coding (SCC). IBC is a method of predicting the current image using a previously reconstructed reference block in the current image at a predetermined distance from the current block. When IBC is applied, the position of the reference block in the current image can be encoded as a vector (block vector) corresponding to the predetermined distance. IBC essentially performs prediction in the current image, but can be performed similarly to inter-frame prediction because the reference block is derived within the current image. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0078] The prediction signal generated by the prediction unit can be used to generate a reconstructed signal or a residual signal. Subtractor 115 generates a residual signal (residual block or residual sample array) by subtracting the prediction signal (prediction block or prediction sample array) output from the prediction unit from the input image signal (original block or original sample array). The generated residual signal can be sent to converter 120.
[0079] Transformer 120 can generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graphical Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented graphically. CNT refers to a transform obtained based on a prediction signal generated using all previously reconstructed pixels. Furthermore, the transform process can be applied to square pixel blocks of the same size or to blocks of variable size instead of square.
[0080] Quantizer 130 can quantize the transform coefficients and send them to entropy encoder 190. Entropy encoder 190 can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 130 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form of the quantized transform coefficients.
[0081] The entropy encoder 190 can perform various encoding methods, such as, for example, exponential Golomb, context-adaptive variable-length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoder 190 can encode, together or separately, the information necessary for video / image reconstruction, excluding quantized transform coefficients (e.g., values of syntax elements, etc.). The encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream at the Network Abstraction Layer (NAL) level. The video / image information may further include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). Additionally, the video / image information may further include general constraint information. The information transmitted by signaling, the transmitted information, and / or syntax elements described in this disclosure can be encoded and included in the bitstream through the above encoding process.
[0082] The bitstream can be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) for transmitting the signal output from the entropy encoder 190 and / or a storage unit (not shown) for storing the signal may be included as internal / external components of the image encoding apparatus 100. Alternatively, a transmitter may be provided as a component of the entropy encoder 190.
[0083] The quantized transform coefficients output from quantizer 130 can be used to generate residual signals. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via dequantizer 140 and inverse transformer 150.
[0084] Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame prediction unit 180 or intra-frame prediction unit 185 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If there is no residual for the block to be processed, such as in the case of applying a skip mode, the prediction block can be used as the reconstructed block. Adder 155 may be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be used for inter-frame prediction of the next image by filtering as described below.
[0085] Luminance mapping with chroma scaling (LMCS) can be applied during image encoding and / or reconstruction.
[0086] Filter 160 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 160 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 170, specifically in the DPB of memory 170. Various filtering methods may include, for example, deblocking filtering, sample adaptive shifting, adaptive loop filtering, bilateral filtering, etc. Filter 160 can generate various filtering-related information and send the generated information to entropy encoder 190, as described later in the description of each filtering method. The filtering-related information can be encoded by entropy encoder 190 and output as a bitstream.
[0087] The modified reconstructed image sent to memory 170 can be used as a reference image in inter-frame prediction unit 180. When inter-frame prediction is applied by image coding device 100, prediction mismatch between image coding device 100 and image decoding device can be avoided and coding efficiency can be improved.
[0088] The DPB of memory 170 can store modified reconstructed images for use as reference images in inter-frame prediction unit 180. Memory 170 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of already reconstructed blocks in the image. The stored motion information can be sent to inter-frame prediction unit 180 and used as motion information for spatially or temporally neighboring blocks. Memory 170 can store reconstructed samples of reconstructed blocks in the current image and can transmit the reconstructed samples to intra-frame prediction unit 185.
[0089] Overview of an image decoding device
[0090] Figure 3 This is a schematic view illustrating an image decoding apparatus to which embodiments of the present disclosure may be applied.
[0091] like Figure 3 As shown, the image decoding device 200 may include an entropy decoder 210, a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame prediction unit 260, and an intra-frame prediction unit 265. The inter-frame predictor (inter-frame prediction unit) 260 and the intra-frame predictor (intra-frame prediction unit) 265 may be collectively referred to as "prediction unit (predictor)". The dequantizer 220 and the inverse transformer 230 may be included in a residual processor.
[0092] According to an embodiment, all or at least some of the components of the image decoding device 200 may be configured by hardware components (e.g., a decoder or a processor). Additionally, the memory 250 may include a decoded image buffer (DPB) or may be configured by a digital storage medium.
[0093] The image decoding device 200, having received a bitstream including video / image information, can perform operations related to... Figure 2 The image is reconstructed using a process corresponding to the process executed by the image encoding apparatus 100. For example, the image decoding apparatus 200 can perform decoding using a processing unit applied in the image encoding apparatus. Therefore, the decoding processing unit can be, for example, an encoding unit. The encoding unit can be obtained through a partitioned coding tree unit or a maximum coding unit. The reconstructed image signal decoded and output by the image decoding apparatus 200 can be reproduced by a reproduction apparatus (not shown).
[0094] Image decoding device 200 can receive data in bitstream form from... Figure 2The signal output by the image encoding apparatus. The received signal can be decoded by the entropy decoder 210. For example, the entropy decoder 210 can parse the bitstream to derive information (e.g., video / image information) necessary for image reconstruction (or picture reconstruction). The video / image information may further include information about various parameter sets, such as adaptive parameter sets (APS), picture parameter sets (PPS), sequence parameter sets (SPS), or video parameter sets (VPS). In addition, the video / image information may further include general constraint information. The image decoding apparatus can further decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements transmitted / received by the signal described in this disclosure can be decoded and obtained from the bitstream through a decoding process. For example, the entropy decoder 210 decodes the information in the bitstream based on encoding methods such as exponential Golomb coding, CAVLC, or CABAC, and outputs the values of the syntax elements necessary for image reconstruction and the quantized values of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the target syntax element, decoding information of neighboring blocks and the target block, or information about symbols / bins decoded in previous stages, and perform arithmetic decoding on the bins based on the determined context model by predicting the occurrence probability of the bins, generating symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model after determining the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin. The prediction-related information in the information decoded by the entropy decoder 210 can be provided to the prediction units (inter-frame predictor 260 and intra-frame prediction unit 265), and the residual values of the entropy decoding performed in the entropy decoder 210, i.e., the quantized transform coefficients and related parameter information, can be input to the dequantizer 220. Additionally, the filtering information in the information decoded by the entropy decoder 210 can be provided to the filter 240. Meanwhile, the receiver (not shown) for receiving the signal output from the image encoding device can be further configured as an internal / external element of the image decoding device 200, or the receiver can be a component of the entropy decoder 210.
[0095] Furthermore, the image decoding apparatus according to this disclosure can be referred to as a video / image / picture decoding apparatus. The image decoding apparatus can be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include an entropy decoder 210. The sample decoder may include at least one of a dequantizer 220, an inverse transformer 230, an adder 235, a filter 240, a memory 250, an inter-frame prediction unit 260, or an intra-frame prediction unit 265.
[0096] The dequantizer 220 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 220 can rearrange the quantized transform coefficients in the form of two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the image encoding device. The dequantizer 220 can obtain the transform coefficients by performing dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information).
[0097] The inverse transformer 230 can perform inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0098] The prediction unit can perform prediction on the current block and generate a prediction block that includes prediction samples for the current block. The prediction unit can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the predictions output from the entropy decoder 210, and can determine a specific intra-frame / inter-frame prediction mode (prediction technique).
[0099] As described in the prediction unit of the image coding apparatus 100, the prediction unit can generate a prediction signal based on various prediction methods (techniques) described later.
[0100] Intra-prediction unit 265 can predict the current block by referring to samples in the current image. The description of intra-prediction unit 185 is also applied to intra-prediction unit 265.
[0101] Inter-frame prediction unit 260 can derive a prediction block for the current block based on a reference block (reference sample array) specified by motion vectors on a reference image. In this case, to reduce the amount of motion information transmitted in the inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may further include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame prediction unit 260 can configure a motion information candidate list based on neighboring blocks and derive the motion vector and / or reference image index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode used for the current block.
[0102] Adder 235 generates a reconstructed block by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including inter-frame prediction unit 260 and / or intra-frame prediction unit 265). If no residual exists for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the reconstructed block. The description of adder 155 also applies to adder 235. Adder 235 may be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and can be used for inter-frame prediction of the next image by filtering as described below.
[0103] Filter 240 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 240 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and store the modified reconstructed image in memory 250, specifically in the DPB of memory 250. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0104] The (modified) reconstructed image stored in the DPB of memory 250 can be used as a reference image in inter-frame prediction unit 260. Memory 250 can store motion information of blocks from which motion information in the current image is derived (or decoded) and / or motion information of already reconstructed blocks in the image. The stored motion information can be sent to inter-frame prediction unit 260 so that it can be utilized as motion information of spatially or temporally neighboring blocks. Memory 250 can store reconstructed samples of reconstructed blocks in the current image and transmit the reconstructed samples to intra-frame prediction unit 265.
[0105] In this disclosure, the embodiments described in the filter 160, inter-frame prediction unit 180 and intra-frame prediction unit 185 of the image encoding apparatus 100 can be applied equally or correspondingly to the filter 240, inter-frame predictor 260 and intra-frame predictor 265 of the image decoding apparatus 200.
[0106] Coding layers and structures
[0107] Encoded videos / images according to this document can be processed according to, for example, the encoding layers and structures described below.
[0108] Figure 4 This is a diagram illustrating the layer structure of an image used for encoding.
[0109] The encoded image is divided into the decoding process that handles the image and its own Video Coding Layer (VCL), the subsystem that transmits and stores the encoded information, and the Network Abstraction Layer (NAL) that is responsible for network adaptation functions and exists between the VCL and the subsystems.
[0110] In VCL, VCL data that includes compressed image data (slice data) can be generated, or parameter sets or supplementary enhancement information (SEI) messages that are additionally required for the image decoding process, such as picture parameter sets (PPS), sequence parameter sets (SPS), and video parameter sets (VPS), can be generated.
[0111] In NAL, NAL units can be generated by adding header information (NAL unit header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, RBSP refers to slice data, parameter sets, SEI messages, etc., generated in VCL. The NAL unit header can include NAL unit type information specified based on the RBSP data included in the corresponding NAL unit.
[0112] like Figure 4 As shown, NAL units can be classified into VCL NAL units and non-VCL NAL units based on the RBSP generated in the VCL. A VCL NAL unit can refer to a NAL unit that includes information about the image (slice data), while a non-VCL NAL unit can refer to a NAL unit that includes information required to decode the image (parameter set or SEI message).
[0113] The aforementioned VCL NAL units and non-VCL NAL units can be sent over the network with additional header information according to the subsystem's data standard. For example, NAL units can be converted into predetermined standard data formats such as H.266 / VVC file format, Real-time Transport Protocol (RTP), Transport Streaming (TS), etc., and sent over various networks.
[0114] As described above, NAL units can be specified using the NAL unit type based on the RBSP data structure included in the corresponding NAL unit, and information about the NAL unit type can be stored in the NAL unit header and sent by signal.
[0115] For example, NAL units can be classified into VCL NAL unit types and non-VCL NAL unit types based on whether they include information about images (slice data). VCL NAL unit types can be classified based on the nature and type of the images included in the VCL NAL unit, and non-VCL NAL unit types can be classified based on the type of parameter set.
[0116] The following is an example of a NAL cell type specified based on the type of the parameter set included in a non-VCL NAL cell type.
[0117] - APS (Adaptive Parameter Set) NAL Unit: The type of NAL unit that includes APS.
[0118] - DPS (Decoding Parameter Set) NAL Unit: The type of NAL unit that includes DPS.
[0119] - VPS (Video Parameter Set) NAL Unit: The type used for NAL units that include the VPS.
[0120] - SPS (Sequence Parameter Set) NAL Unit: The type of NAL unit that includes SPS.
[0121] - PPS (Picture Parameter Set) NAL Unit: The type of NAL unit that includes PPS.
[0122] The aforementioned NAL unit type can have syntax information for the NAL unit type, and this syntax information can be stored in the NAL unit header and sent by signals. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified by the nal_unit_type value.
[0123] A slice header (slice header syntax) may include information / parameters that can be commonly applied to slices. APS (APS syntax) or PPS (PPS syntax) may include information / parameters that can be commonly applied to one or more slices or images. SPS (SPS syntax) may include information / parameters that can be commonly applied to one or more sequences. VPS (VPS syntax) may include information / parameters that can be commonly applied to multiple layers. DPS (DPS syntax) may include information / parameters that can be commonly applied to the entire video. DPS may include information / parameters related to the concatenation of encoded video sequences (CVS). In this document, the High-Level Syntax (HLS) may include at least one of the following: APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, and slice header syntax.
[0124] In this disclosure, the image / image information encoded by the image encoding device 100 and transmitted to the image decoding device in the form of a bitstream may include not only intra-image partition information, intra / inter-frame prediction information, residual information, and loop filtering information, but also information included in the slice header, information included in the APS, information included in the PPS, information included in the SPS, and / or information included in the VPS.
[0125] SEI Messages – Source Image Timing Information (SPTI)
[0126] The SPTI SEI message indicates the temporal distance between the source images associated with the corresponding decoded output image before encoding. For example, for content captured by a camera, the temporal distance between source images is the difference between the time it takes for the image sensor to be exposed to produce the source image associated with the current decoded image and the time it takes for the image sensor to be exposed to produce the source images associated with the previously decoded images in output order.
[0127] Table 1 below shows an example of the SPTI SEI message syntax structure.
[0128] [Table 1]
[0129] A `spti_cancel_flag` value of 1 indicates that the SPTI SEI message cancels the persistence of any previous SPTI SEI messages applied to the current layer in the output order. A `spti_cancel_flag` value of 0 indicates that the source image timing information persists.
[0130] `spti_persistence_flag` indicates the persistence of SPTI SEI messages used for the current layer. `spti_persistence_flag` equal to 0 indicates that SPTI SEI messages are applied only to the currently decoded picture. `spti_persistence_flag` equal to 1 indicates that SPTI SEI messages are applied to the currently decoded picture and continue to be applied to all subsequent pictures of the current layer in output order until one or more of the following conditions (conditions 1 to 3) are true.
[0131] - Condition 1: A new CLVS begins in the current layer.
[0132] - Condition 2: End of bit stream.
[0133] - Condition 3: The image in the current layer of the AU associated with the SPTI SEI message is output, following the current image in the output order.
[0134] `spti_source_picture_timing_type` indicates the timing relationship between the source picture and the corresponding decoded output picture, as specified in Table 2. When `(spti_source_picture_timing_type & bitmask)` is not equal to 0, the interpretation related to the `bitMask` value in Table 2 does not apply to SPTI SEI messages. When `spti_source_picture_timing_type` is equal to 0, the timing relationship can be specified by the application.
[0135] [Table 2]
[0136] A value of 1 for `spti_source_timing_equals_output_timing_flag` indicates that the timing of the source image is the same as the timing of the corresponding decoded output image. A value of 0 for `spti_source_timing_equals_output_timing_flag` indicates that the timing of the source image is different from the timing of the corresponding decoded output image.
[0137] When spti_source_timing_qu_equals_output_timing_flag equals 1 and there is an image timing SEI message for the current image, the source image timing can be determined from the information carried in the image timing SEI message.
[0138] spti_time_scale indicates the number of time units that elapse in one second. The value of spti_time_scale does not have to be equal to 0. For example, spti_time_scale for a time coordinate system that measures time using a 27MHz clock could be 27,000,000.
[0139] `spti_num_units_in_elementary_source_picture_interval` indicates the number of time units of a clock operating at a frequency of `spti_time_scale Hz`, which corresponds to the specified base source picture interval for consecutive pictures in output order in CLVS. `spti_max_sublayers_minus_1` plus 1 indicates the maximum number of time sublayers that may exist in CLVS.
[0140] `spti_sublayer_source_picture_interval_scale_factor[i]` (if present) can indicate the scaling factor used to determine the source picture intervals of corresponding consecutive pictures in CLVS with TemporalId less than or equal to i, in output order. `spti_sublayer_source_picture_interval_scale_factor[i]` equal to 0 can be used to indicate that the source picture corresponding to the current decoded output picture is the same as the source picture corresponding to the previously decoded output picture.
[0141] A value of 1 for spti_sublayer_syntheed_picture_flag[i] (when present) indicates that the decoded output image belonging to the i-th temporal sublayer was synthesized and does not correspond to the unmodified original source image. A value of 0 for spti_sublayer_syntheed_picture_flag[i] does not provide such an indication. When spti_sublayer_syntheed_picture_flag[i] does not exist, its value can be inferred to be 0.
[0142] The image encoding / decoding methods according to various embodiments of the present disclosure will be described in detail below.
[0143] In the existing design of SPTI SEI messages, the syntax elements spti_max_sublayers_minus1, spti_sublayer_source_picture_interval_scale_factor[i], and spti_sublayer_syntheed_picture_flag[i] can be signaled regardless of the persistence of the SEI message. These syntax elements may not be necessary when the SEI message is only used for a single picture (i.e., for a picture in the same AU containing the SEI message). Furthermore, the computation of SourcePictureInterval[] and SourcePictureTime[] may not be necessary when the persistence of the SPTI SEI only applies to the current picture. To address these issues, this disclosure proposes the following four solutions.
[0144] 1. Signal the syntax elements spti_sublayer_source_picture_interval_scale_factor[i] and spti_sublayer_syntheed_picture_flag[i] only when spti_persistence_flag is equal to 1.
[0145] 2. Update the descriptions associated with SourcePictureInterval[i] and SourcePictureTime[n] so that they are applied only when spti_persistence_flag is equal to 1.
[0146] 3. When the persistence of SPTI SEI applies only to the current image, set SourcePictureInterval to equal ElementalSourcePictureInterval.
[0147] 4. Specify a constraint such that the value is inferred to be equal to 1 when spti_sublayer_source_picture_interval_scale_factor[i] is not present.
[0148] Example 1
[0149] To present solutions 1-3 of the four solutions described above, Example 1 proposes an updated SPTI SEI message syntax structure as shown in Table 3 below. The SPTI SEI message syntax structure proposed in this example is shown in Table 3 below.
[0150] [Table 3]
[0151] Referring to Table 3, spti_max_sublayers_minus1 can only be signaled if spti_persistence_flag is equal to 1. Increasing spti_max_sublayers_minus_1 by 1 indicates the maximum number of time sublayers that may exist in CLVS. When spti_max_sublayers_minus_1 does not exist and spti_persistence_flag is equal to 0, spti_max_sublayers_minus1 can be inferred to be equal to 0.
[0152] spti_sublayer_source_picture_interval_scale_factor[i] can indicate the scaling factor used to determine the source picture interval of corresponding consecutive pictures in CLVS with TemporalId less than or equal to i and in output order. spti_sublayer_source_picture_interval_scale_factor[i] equal to 0 can indicate that the source picture corresponding to the current decoded output picture is the same as the source picture corresponding to the previous decoded output picture.
[0153] When spti_persistence_flag equals 1, the procedures in Table 4 below can be executed.
[0154] [Table 4]
[0155] A value of 1 for `spti_sublayer_syntheed_picture_flag[i]` indicates that the decoded output image belonging to the i-th temporal sublayer is synthesized and does not correspond to the unmodified original source image. A value of 0 for `spti_sublayer_syntheed_picture_flag[i]` may not provide such an indication. That is, a value of 0 for `spti_sublayer_syntheed_picture_flag[i]` indicates that such a restriction does not exist. When it does not exist, `spti_sublayer_syntheed_picture_flag[i]` can be inferred to be equal to 0.
[0156] Example 2
[0157] Example 2 is an embodiment of Solution 4, one of the four solutions described above. Example 2 proposes updating the SPTI SEI message, as shown in Table 5 below.
[0158] [Table 5]
[0159] `spti_sublayer_source_picture_interval_scale_factor[i]` indicates the scaling factor used to determine the source image intervals of corresponding consecutive images in CLVS with TemporalId less than or equal to i, in output order. `spti_sublayer_source_picture_interval_scale_factor[i]` equal to 0 indicates that the source image corresponding to the current decoded output image is the same as the source image corresponding to the previously decoded output image. When it does not exist, the value of `spti_sublayer_source_picture_interval_scale_factor[i]` is inferred to be equal to 1.
[0160] Figure 5 This is a flowchart of the image encoding method according to this disclosure. (Reference) Figure 5 The image encoding device 100 can encode image information related to the current image (S510). Then, the image encoding device 100 can send a bit stream containing the image information (S530). In other words, the image encoding device 100 can send a bit stream containing image information to the image decoding device 200.
[0161] Here, the image information may include an SPTI SEI message related to the time interval between the source image and the corresponding decoded image. The SPTI SEI message may include persistence information specifying whether the SPTI SEI message applies only to the currently decoded image. Here, the persistence information may be spti_persistence_flag.
[0162] Additionally, the SPTI SEI message may include at least one of the following: first information specifying the maximum number of temporal sublayers, second information specifying the scaling factor used in determining the time interval between source images corresponding to the decoded image, or third information specifying whether the currently decoded image included in the temporal sublayer is synthesized. Here, the first information may be spti_max_sublayers_minus1. The second information may be spti_sublayer_source_picture_interval_scale_factor[i]. The third information may be spti_sublayer_syntheed_picture_flag[i].
[0163] According to embodiments of this disclosure, first information, second information, and / or third information can be sent by signaling based on whether the SPTI SEI message is applied only to the currently decoded image. Specifically, when the SPTI SEI message is applied to more than just the currently decoded image, first information, second information, and / or third information can be sent by signaling.
[0164] According to embodiments of this disclosure, second and / or third information can be further signaled based on the maximum number of time sublayers. Signaling for the second information can be skipped when the scaling factor used in determining the time interval between the source images corresponding to the decoded image is 1.
[0165] According to embodiments of this disclosure, when the SPTI SEI message is applied only to the currently decoded image, the source image interval can be set to be equal to the base source image interval.
[0166] Figure 6 This is a flowchart of the image decoding method according to this disclosure. (Reference) Figure 6 The image decoding device 200 can receive a bitstream including image information (S610). Subsequently, the image decoding device 200 can generate a reconstructed image based on the image information by reconstructing the current image (S630).
[0167] Here, the image information may include an SPTI SEI message related to the time interval between the source image and the corresponding decoded image. Additionally, the SPTI SEI message may include persistence information specifying whether the SPTI SEI message applies only to the currently decoded image. Here, the persistence information may be spti_persistence_flag.
[0168] Additionally, the SPTI SEI message may include at least one of the following: first information specifying the maximum number of temporal sublayers, second information specifying the scaling factor used in determining the time interval between source images corresponding to the decoded image, or third information specifying whether the currently decoded image included in the temporal sublayer is synthesized. Here, the first information may be spti_max_sublayers_minus1. The second information may be spti_sublayer_source_picture_interval_scale_factor[i]. The third information may be spti_sublayer_syntheed_picture_flag[i].
[0169] According to embodiments of this disclosure, first information, second information, and / or third information can be obtained based on the value of persistence information. For example, the first information, second information, and / or third information can be obtained based on persistence information specifying that the SPTI SEI message is not only applied to the currently decoded image.
[0170] According to embodiments of this disclosure, the second and / or third information may be further based on the value of the first information. When the second information is not available, the value of the second information may be inferred to be equal to 1.
[0171] According to embodiments of this disclosure, based on the persistence information that a specified SPTI SEI message is applied only to the currently decoded image, the source image interval can be set to be equal to the base source image interval.
[0172] For clarity, the exemplary methods of this disclosure are expressed as a series of operations, but this is not intended to limit the order in which the steps are performed. Steps may be performed simultaneously or in different orders if necessary. To implement the method according to this disclosure, additional steps may be included in addition to those illustrated, or some steps may be excluded while including the remaining steps. Alternatively, some steps may be excluded, and additional steps may be included.
[0173] In this disclosure, the image encoding device 100 or the image decoding device 200 that performs a predetermined operation (step) can perform an operation (step) that checks the execution conditions or circumstances of the operation (step). For example, when describing the execution of a predetermined operation when a predetermined condition is met, the image encoding device 100 or the image decoding device 200 can perform the predetermined operation after checking whether the predetermined condition is met.
[0174] The various embodiments of this disclosure are not intended to enumerate all possible combinations, but rather to describe representative aspects of this disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more thereof.
[0175] Various embodiments of this disclosure can be implemented by hardware, firmware, software, or a combination thereof. When this disclosure is implemented in hardware, it can be implemented using application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, etc.
[0176] Furthermore, the image decoding apparatus 200 and image encoding apparatus 100 applying the embodiments of this disclosure can be included in multimedia broadcast transceivers, mobile communication terminals, home theater video equipment, digital cinema video equipment, surveillance cameras, video chat devices, and real-time communication devices such as video communication, mobile streaming media devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, 3D video devices, video telephony devices, and medical video devices, and can be used to process image signals or data signals. For example, OTT video devices can include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0177] Figure 7 This is a diagram illustrating an exemplary content streaming system to which embodiments of this disclosure are applicable.
[0178] like Figure 7 As shown, the content streaming system using embodiments of this disclosure may mainly include an encoding server, a streaming server, a web server, media storage, user equipment, and multimedia input devices.
[0179] An encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then sends the bitstream to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders generate bitstreams directly, the encoding server can be omitted.
[0180] The bitstream can be generated by the image encoding method or image encoding apparatus 100 applying the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0181] A streaming server sends multimedia data to a user's device based on a user's request via a web server, and the web server acts as a medium for notifying the user of services. When a user requests a desired service from the web server, the web server can deliver it to the streaming server, and the streaming server can send the multimedia data to the user. In this scenario, the content streaming system may include a separate control server. In this case, the control server acts as a command / response controller between devices in the content streaming system.
[0182] A streaming server can receive content from media storage and / or encoding servers. For example, when content is received from an encoding server, it can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0183] Examples of user equipment may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, tablet PCs, tablet computers, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays), digital televisions, desktop computers, digital signage, etc.
[0184] In a content streaming system, each server can operate as a distributed server, in which case the data received from each server can be distributed.
[0185] The scope of this disclosure includes software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) for enabling the operation of methods according to various embodiments to be executed on a device or computer, and non-transitory computer-readable media having such software or instructions stored thereon and executable on a device or computer.
[0186] Industrial applicability
[0187] The embodiments of this disclosure can be used to encode / decode images.
Claims
1. An image decoding method performed by an image decoding device, the image decoding method comprising: Receive a bitstream including image information; as well as Based on the image information, a reconstructed image is generated by reconstructing the current image. The image information includes Source Image Timing Information (SPTI) SEI messages related to the time interval between source images, and the source images are associated with the corresponding decoded images.
2. The method according to claim 1, wherein, The SPTI SEI message includes persistent information specifying whether the SPTI SEI message applies only to the currently decoded image.
3. The method according to claim 2, wherein, The SPTI SEI message includes at least one of the following: First information, which specifies the maximum number of time sub-layers; The second information specifies the scaling factor used in determining the time interval between the source images corresponding to the decoded image; or The third information specifies whether the currently decoded image included in the temporal sublayer has been synthesized.
4. The method according to claim 3, wherein, The first information, the second information, and the third information are obtained based on the persistent information.
5. The method according to claim 4, wherein, Based on the persistence information, the SPTI SEI message is specified to be applied not only to the current decoded image, to obtain the first information, the second information, and the third information.
6. The method according to claim 4, wherein, The second and third information are further obtained based on the value of the first information.
7. The method according to claim 4, wherein, Since the second information was not obtained, the value of the second information was inferred to be equal to 1.
8. The method according to claim 1, wherein, Based on the persistence information, the SPTI SEI message is specified to be applied only to the currently decoded image, and the source image interval is set to be equal to the base source image interval.
9. An image encoding method performed by an image encoding device, the image encoding method comprising: Generate image information related to the current image, and Encode the bitstream including the image information. The image information includes Source Image Timing Information (SPTI) SEI messages related to the time interval between source images, and the source images are associated with the corresponding decoded images.
10. A non-transitory computer-readable medium storing a bitstream generated by the image encoding method according to claim 9.
11. A method for transmitting a bitstream generated by an image encoding method, the image encoding method comprising: Generate image information related to the current image, and Encode the bitstream including the image information. The image information includes Source Image Timing Information (SPTI) SEI messages related to the time interval between source images, and the source images are associated with the corresponding decoded images.