Decoding device, encoding device, and data transmission device
By using image decoding and encoding devices to deduce the number of slices in the current image from the markers and symbols in the image information, the problem of high cost in sending and storing high-resolution, high-quality images is solved, and efficient image compression and segmentation are achieved.
Patent Information
- Application Number
- CN202511337836.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-27
- Filing Date
- 2020-11-26
- Publication Date
- 2025-11-04
AI Technical Summary
The transmission and storage of high-resolution, high-quality images are costly, necessitating efficient image compression techniques to reduce the amount of information.
By using image decoding and encoding devices, the number of slices in the current image is deduced using the markers and symbols in the image information, thereby achieving image segmentation and encoding and improving image encoding efficiency.
It improves overall image/video compression efficiency and enhances segmentation efficiency by using segmentation information from the current image to improve segmentation efficiency.
Smart Images

Figure CN120897066A_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application No. 202080093704.1 (International Application No.: PCT / KR2020 / 016944, Application Date: November 26, 2020, Invention Title: Method and Apparatus for Sending Signals to Notify Image Segmentation Information). Technical Field
[0002] This disclosure relates to image coding technology, and most specifically, to a method and apparatus for signaling image segmentation information in an image coding system. Background Technology
[0003] Recently, there has been a growing demand for high-resolution, high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, across various fields. Because image data is high-resolution and high-quality, the amount of information or bits to be transmitted increases compared to traditional image data. Therefore, transmission and storage costs increase when using media such as traditional wired / wireless broadband lines to transmit image data or when storing image data using existing storage media.
[0004] Therefore, there is a need for efficient image compression technology that can effectively transmit, store, and reproduce information from high-resolution, high-quality images. Summary of the Invention
[0005] Technical issues
[0006] This disclosure provides methods and apparatus for improving image coding efficiency.
[0007] This disclosure also provides a method and apparatus for signaling segmentation information of an image.
[0008] This disclosure also provides a method and apparatus for decoding a current image based on segmentation information of the current image.
[0009] Technical solution
[0010] In one aspect, an image decoding method performed by a decoding device is provided. The method includes the following steps: obtaining image information about a current image from a bitstream; and decoding the current image based on the image information, wherein the image information includes a first flag related to the presence of sub-image information and a second flag related to whether each sub-image includes only one slice, and wherein, based on the first flag and the second flag, the number of slices included in the current image is derived to be equal to 1.
[0011] On the other hand, an image encoding method performed by an encoding device is provided. The method includes the steps of: deriving at least one slice by segmenting a current image; and encoding image information of the current image based on the at least one slice, wherein the image information includes the first flag related to the presence of sub-image information and the second flag related to whether each sub-image in the current image includes only one slice, and wherein, based on the first flag and the second flag, the number of slices included in the current image is derived to be equal to 1.
[0012] In another aspect, a non-transitory computer-readable storage medium is provided for storing a bitstream, the bitstream including image information that causes an image decoding method to be executed. The image decoding method includes the steps of: obtaining image information about a current image from the bitstream; and decoding the current image based on the image information, wherein the image information includes the first flag related to the presence of sub-image information and the second flag related to whether each sub-image includes only one slice, and wherein, based on the first flag and the second flag, the number of slices included in the current image is derived to be equal to 1.
[0013] Beneficial effects
[0014] According to this disclosure, the overall image / video compression efficiency can be improved.
[0015] According to this disclosure, the efficiency of segmentation can be improved.
[0016] According to this disclosure, segmentation efficiency can be improved based on the segmentation information of the current image. Attached Figure Description
[0017] Figure 1 Examples of video / image coding systems to which the embodiments of this document can be applied are illustrated.
[0018] Figure 2 This is a schematic diagram illustrating the configuration of a video / image encoding device to which the embodiments described in this document can be applied.
[0019] Figure 3 This is a schematic diagram illustrating the configuration of a video / image decoding device to which the embodiments described in this document can be applied.
[0020] Figure 4 An exemplary hierarchy of encoded data is illustrated.
[0021] Figure 5 This is an example illustration of image segmentation.
[0022] Figure 6 This is a flowchart illustrating the image encoding process according to an implementation method.
[0023] Figure 7 This is a flowchart illustrating the image decoding process according to an implementation method.
[0024] Figure 8 This is a flowchart illustrating the operation of an encoding device according to an embodiment.
[0025] Figure 9 This is a block diagram illustrating the configuration of an encoding device according to an embodiment.
[0026] Figure 10 This is a flowchart illustrating the operation of a decoding device according to an embodiment.
[0027] Figure 11 This is a block diagram illustrating the configuration of a decoding device according to an embodiment.
[0028] Figure 12 Examples of content streaming systems to which the implementation methods disclosed in this document can be applied are illustrated. Detailed Implementation
[0029] This document can be modified in various ways and has various embodiments, which will be described in detail and illustrated in the accompanying drawings. However, this is not intended to limit this document to the specific embodiments. The general terminology used in this specification is used only to describe specific embodiments and is not intended to limit the technical spirit of this document. Singular expressions include plural expressions unless the context clearly indicates otherwise. Terms such as "comprising" and "having" in this specification should be understood to indicate the presence of the features, numbers, steps, operations, elements, parts or combinations thereof described in the specification, and do not exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, elements, parts or combinations thereof.
[0030] Furthermore, the elements in the accompanying drawings described in this document are illustrated independently for the convenience of describing different features and functions. This does not mean that each element is implemented as different hardware or different software. For example, at least two of the elements can be combined to form a single element, and a single element can also be divided into multiple elements. Embodiments in which elements are combined and / or separated are also included within the scope of this document without departing from its spirit.
[0031] In this disclosure, "A or B" may mean "A only", "B only", or "both A and B". In other words, in this disclosure, "A or B" can be interpreted as "A and / or B". For example, in this disclosure, "A, B or C" may mean "A only", "B only", "C only", or "any combination of A, B, and C".
[0032] The forward slash ( / ) or comma used in this disclosure can mean "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0033] In this disclosure, "at least one of A and B" may mean "only A", "only B" or "both A and B". Furthermore, in this specification, the expression "at least one of A or B" or "at least one of A and / or B" may be interpreted as "at least one of A and B".
[0034] Additionally, in this disclosure, "at least one of A, B, and C" may mean "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0035] Additionally, the parentheses used in this disclosure may mean "for example". Specifically, when indicated as "prediction (intra-frame prediction)", it may mean that "intra-frame prediction" is proposed as an example of "prediction". In other words, "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" may be proposed as an example of "prediction". Furthermore, when indicated as "prediction (i.e., intra-frame prediction)", it may also mean that "intra-frame prediction" is proposed as an example of "prediction".
[0036] In this disclosure, a technical feature described separately in a single drawing may be implemented individually or simultaneously.
[0037] In the following, preferred embodiments of this document are described in more detail with reference to the accompanying drawings. In the drawings, the same reference numerals are used for the same elements, and redundant descriptions of the same elements may be omitted.
[0038] Figure 1 Examples of video / image coding systems to which the embodiments of this document can be applied are illustrated.
[0039] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data in the form of a file or stream to the receiving device via a digital storage medium or network.
[0040] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0041] Video sources can acquire video / images through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. For example, a video / image capture device may include one or more cameras, a video / image archive including previously captured video / images, etc. For example, a video / image generation device may include a computer, tablet computer, and smartphone, and may generate video / images (electronically). For example, virtual video / images may be generated via a computer, etc. In this case, the video / image capture process may be replaced by a process that generates related data.
[0042] Encoding devices can encode input video / images. For compression and encoding efficiency, encoding devices can perform a series of processes such as prediction, transformation, and quantization. The encoded data (encoded video / image information) can be output as a bitstream.
[0043] The transmitter can transmit encoded images / image information or data, output as a bitstream, to the receiver of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter may include elements for generating media files according to a predetermined file format and may include elements for transmission over a broadcast / communication network. The receiver can receive / extract the bitstream and transmit the received bitstream to a decoding device.
[0044] Decoding devices can decode video / images by performing a series of processes such as inverse quantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0045] The renderer can render decoded video / images. The rendered video / images can be displayed on a monitor.
[0046] This document relates to video / image coding. For example, the methods / implementations disclosed in this document can be applied to methods disclosed in Universal Video Coding (VVC), EVC (Essential Video Coding) standard, AOMedia Video 1 (AV1) standard, second-generation audio-visual coding standard (AVS2), or next-generation video / image coding standards (e.g., H.267 or H.268).
[0047] This document presents various implementations of video / image coding, and unless otherwise mentioned, these implementations can be combined with each other.
[0048] In this disclosure, video can refer to a collection of images over time. Generally, an image refers to a unit of image representing a specific time period, and a slice / tile is a unit that constitutes a portion of an image. A slice / tile may include one or more coding tree units (CTUs). An image may include one or more slices / tiles.
[0049] A tile is a rectangular area of CTUs within a specific tile column and a specific tile row in an image. A tile column is a rectangular area of CTUs with a height equal to the height of the image and a width specified by a syntax element in the image parameter set. A tile row is a rectangular area of CTUs with a height specified by a syntax element in the image parameter set and a width equal to the width of the image. A tile scan is a specific ordering of CTUs in a segmented image, where CTUs are ordered consecutively by a CTU raster scan within a tile, and tiles in the image are ordered consecutively by a raster scan of the image's tiles. A slice may include multiple intact tiles or multiple consecutive columns of CTUs within a tile of an image that may be included in a single NAL unit. In this disclosure, tile groups can be used interchangeably with slices. For example, in this disclosure, a tile group / tile group header may be referred to as a slice / slice header.
[0050] Furthermore, an image can be divided into two or more sub-images. A sub-image can be a rectangular region of one or more slices within the image.
[0051] A pixel, or image unit, can refer to the smallest unit that makes up a picture (or image). Additionally, the term "sample" can be used as the counterpart to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or the pixel / pixel value of the chrominance component.
[0052] A unit can represent a basic unit of image processing. A unit may include a specific region of an image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region". In general, an M×N block may include a set (or array) of samples (or sample arrays) or transform coefficients in M columns and N rows. Figure 2 This is a schematic illustration of the configuration of a video / image encoding apparatus to which the embodiments of this document can be applied. In the following, the video encoding apparatus may include an image encoding apparatus.
[0053] Reference Figure 2The encoding device 200 includes an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transform 232, a quantizer 233, an inverse quantizer 234, and an inverse transform 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0054] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into one or more processors. For example, a processor may be referred to as a coding unit (CU). In this case, the coding unit may be recursively segmented from coding tree unit (CTU) or maximum coding unit (LCU) according to a quadtree-binary-trinary tree (QTBTTT) structure. For example, a coding unit may be segmented into multiple deeper coding units based on a quadtree structure, a binary tree structure, and / or a ternary structure. In this case, for example, a quadtree structure may be applied first, followed by a binary tree structure and / or a ternary structure. Alternatively, a binary tree structure may be applied first. The encoding process according to this disclosure may be performed based on the final coding unit that is no longer segmented. In this case, the maximum coding unit may be used as the final coding unit based on image characteristics, coding efficiency, etc., or if necessary, the coding unit may be recursively segmented into deeper coding units, and the coding unit with the optimal size may be used as the final coding unit. Here, the encoding process may include prediction, transformation, and reconstruction processes (described later). As another example, the processor may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may be split or separated from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.
[0055] In some cases, a unit can be used interchangeably with terms such as block or region. Generally, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can typically represent pixels or pixel values, either pixel / pixel values representing only the luminance component or pixel / pixel values representing only the chrominance component. A sample can be used as a term corresponding to a picture (or image) of pixels or pictographs.
[0056] In encoding device 200, a residual signal (residual block, residual sample array) is generated by subtracting the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array), and the generated residual signal is sent to converter 232. In this case, as shown, the portion in encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be referred to as subtractor 231. The predictor can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block that includes the prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As described later in the description of each prediction mode, the predictor can generate various prediction-related information such as prediction mode information and send the generated information to entropy encoder 240. The prediction information can be encoded in entropy encoder 240 and output as a bitstream.
[0057] Intra-predictor 222 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples may be located near or separated from the current block. In intra-prediction, the prediction mode may include multiple non-directional modes and multiple directional modes. For example, non-directional modes may include DC mode and planar mode. For example, depending on the level of detail in the prediction direction, the directional modes may include 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes may be used depending on the settings. Intra-predictor 222 can use the prediction modes applied to neighboring blocks to determine the prediction mode applied to the current block.
[0058] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. Here, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a juxtaposed reference block, juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to deduce the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information of neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0059] Predictor 220 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as Intra-Frame and Inter-Frame Prediction Combination (CIIP). Alternatively, the predictor can predict blocks based on Intra-Block Copy (IBC) prediction mode or Palette mode. IBC prediction mode or Palette mode can be used for content image / video coding such as games, for example, Screen Content Coding (SCC). IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction in terms of deriving reference blocks in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. Palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying Palette mode, sample values within the frame can be signaled based on information about the palette table and palette index.
[0060] The predicted signal generated by the predictor (including inter-frame predictor 221 and / or intra-frame predictor 222) can be used to generate a reconstructed signal or a residual signal. Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform technique can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when representing the relationship information between pixels using a graph. CNT refers to a transform generated based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or to blocks of variable size that are not square.
[0061] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector form based on the coefficient scan order, and generate information about the quantized transform coefficients based on the one-dimensional vector form. Entropy encoder 240 can perform various encoding methods such as (e.g.) exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can encode information necessary for video / image reconstruction, other than the quantized transform coefficients (e.g., values of syntax elements), either together or separately. The encoded information (e.g., encoded video / image information) can be sent or stored in NAL (Network Abstraction Layer) units as a bitstream. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. In this document, information and / or syntax elements sent / signed from the encoding device to the decoding device may be included in the video / image information. The video / image information may be encoded by the encoding process described above and included in the bitstream. The bitstream may be transmitted over a network or stored in a digital storage medium. The network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that transmits the signal output from the entropy encoder 240 or a storage device (not shown) that stores the signal may be included as an internal / external element of the encoding device 200, and alternatively, the transmitter may be included in the entropy encoder 240.
[0062] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying inverse quantization and inverse transform to the quantized transform coefficients via inverse quantizer 234 and inverse transformer 235. Adder 250 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). If there is no residual for the block to be processed (e.g., in the case of applying a skip mode), the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and, as described below, can be used for inter-frame prediction of the next image after filtering.
[0063] In addition, a luminance mapping with chroma scaling (LMCS) can be applied during image encoding and / or reconstruction processing.
[0064] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 270, specifically in the DPB of memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset (SAO), adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various types of filtering-related information and transmit the generated information to entropy encoder 290, as described in the subsequent descriptions of each filtering method. The filtering-related information can be encoded by entropy encoder 290 and output as a bitstream.
[0065] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. When inter-frame prediction is applied by the encoding device, prediction mismatch between the encoding device 200 and the decoding device can be avoided, and encoding efficiency can be improved.
[0066] The DPB of memory 270 can store modified reconstructed images for use as reference images in inter-frame predictor 221. Memory 270 can store motion information of blocks from which motion information in the current image is derived (or encoded) and / or motion information of blocks in already reconstructed images. The stored motion information can be transmitted to inter-frame predictor 221 to be used as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current image and can transmit these reconstructed samples to intra-frame predictor 222.
[0067] Figure 3This is a schematic diagram illustrating the configuration of a video / image decoding device to which the disclosed content of this document can be applied.
[0068] Reference Figure 3 The decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an intra-frame predictor 331 and an inter-frame predictor 332. The residual processor 320 may include an inverse quantizer 321 and an inverse transformer 322. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 may be configured by hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0069] When the input includes a bitstream containing video / image information, the decoding device 300 can reconstruct and... Figure 2 The encoding device processes video / image information corresponding to the image. For example, the decoding device 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can use a processor applied in the encoding device to perform decoding. Therefore, for example, the processor for decoding can be an encoding unit, and the encoding unit can be segmented from encoding tree units or maximum encoding units according to a quadtree structure, binary tree structure, and / or ternary tree structure. One or more transform units can be derived from the encoding units. The reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproduction device.
[0070] Decoding device 300 can receive data from... in the form of a bitstream. Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information necessary for image reconstruction (or picture reconstruction) (e.g., video / image information). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The signaling / receiving information and / or syntax elements described subsequently in this document can be decoded by the decoding process and obtained from the bitstream. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as exponential Golomb coding, CABAC, or CAVLC, and output the syntax elements required for image reconstruction and the quantized values of the transform coefficients for the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine the context model using information about the target syntax element, the decoding information of the target block, or information about the symbols / bins decoded in the previous stage, and perform arithmetic decoding on the bins by predicting the occurrence probability of the bins based on the determined context model, generating symbols corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method can update the context model by using the information of the decoded symbols / bins for the context model of the next symbol / bin after determining the context model. The prediction-related information in the information decoded by the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantized transform coefficients and related parameter information) from the entropy decoding performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive the residual signals (residual blocks, residual samples, residual sample arrays). In addition, the filtering information in the information decoded by the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) for receiving signals output from the encoding device can be additionally configured as an internal / external component of the decoding device 300, or the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this document can be referred to as a video / image / picture decoding device, and the decoding device can be classified as an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0071] The dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients in a two-dimensional block format. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. The dequantizer 321 can use quantization parameters (e.g., quantization step size information) to perform dequantization on the quantized transform coefficients and obtain the transform coefficients.
[0072] The inverse transformer 322 performs inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0073] The predictor can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 310, and can determine a specific intra-frame / inter-frame prediction mode.
[0074] Predictor 330 can generate prediction signals based on various prediction methods described below. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but also apply both intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as Intra-Frame and Inter-Frame Prediction Combination (CIIP). Alternatively, the predictor can predict blocks based on Intra-Block Copy (IBC) prediction mode or Palette mode. IBC prediction mode or Palette mode can be used for content image / video coding such as games, for example, Screen Content Coding (SCC). IBC essentially performs prediction in the current frame, but can be performed similarly to inter-frame prediction in terms of deriving reference blocks in the current frame. That is, IBC can use at least one of the inter-frame prediction techniques described in this document. Palette mode can be considered as an example of intra-frame coding or intra-frame prediction. When applying Palette mode, sample values within the frame can be signaled based on information about the palette table and palette index.
[0075] Intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the referenced samples can be located near the current block or separately. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Intra-predictor 331 can determine the prediction mode applied to the current block by using the prediction modes applied to neighboring blocks.
[0076] Inter-frame predictor 332 can deduce the predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index for the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.
[0077] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, or reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block or prediction sample array) output from predictor 330. If there is no residual to process the target block (e.g., in the case of applying a jump mode), the prediction block can be used as the reconstruction block.
[0078] Adder 340 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current image, and as described later, it can also be output by filtering or used for inter-frame prediction of the next image.
[0079] In addition, Luminance Mapping with Chroma Scaling (LMCS) can also be applied to image decoding processing.
[0080] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and store the modified reconstructed image in memory 360, specifically in the DPB of memory 360. Various filtering methods may include, for example, deblocking filtering, adaptive sample shifting, adaptive loop filtering, bilateral filtering, etc.
[0081] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks from which motion information in the current image is derived (or decoded) and / or motion information of blocks in already reconstructed images. The stored motion information can be transmitted to inter-frame predictor 332 to be used as motion information for spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and transmit the reconstructed samples to intra-frame predictor 331.
[0082] In this disclosure, the embodiments described in the filter 260, inter-frame predictor 221, and intra-frame predictor 222 of the encoding device 200 can be the same as, or applied separately to, the filter 350, inter-frame predictor 332, and intra-frame predictor 331 of the decoding device 300, corresponding to the filter 350, inter-frame predictor 332, and intra-frame predictor 331 of the decoding device 300. This also applies equally to the inter-frame predictor 332 and the intra-frame predictor 331.
[0083] As described above, during video encoding, prediction is performed to enhance compression efficiency. A prediction block, which includes predicted samples of the current block (i.e., the target coding block), can be generated through prediction. In this case, the prediction block includes predicted samples in the spatial domain (or pixel domain). The prediction block is derived identically in both the encoding and decoding devices. The encoding device can enhance image coding efficiency by signaling information (residual information) to the decoding device regarding the residual between the original block (rather than the original sample values of the original block) and the prediction block. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block and the prediction block, and generate a reconstructed image including the reconstructed block.
[0084] Residual information can be generated through transform and quantization processes. For example, the encoding device can derive the residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, derive quantized transform coefficients by performing a quantization process on the transform coefficients, and signal the relevant residual information (via bitstream) to the decoding device. In this case, the residual information may include the values of the quantized transform coefficients, position information, transform scheme, transform kernel, and quantization parameters. The decoding device can perform inverse quantization / inverse transform based on the residual information and derive residual samples (or residual blocks). The decoding device can generate a reconstructed image based on the prediction block and the residual block. Furthermore, for inter-frame prediction reference of subsequent images, the encoding device can derive the residual block by performing inverse quantization / inverse transform on the quantized transform coefficients and generate a reconstructed image.
[0085] Figure 4 An exemplary hierarchy of encoded data is illustrated.
[0086] Reference Figure 4 Encoded data can be divided into the Video Coding Layer (VCL), which manipulates the encoding processing of video / images and the video / images themselves, and the Network Abstraction Layer (NAL), which stores and transmits the encoded video / image data and is located between the Video Coding Layer (VCL) and the lower system.
[0087] VCL can generate Supplemental Enhancement Information (SEI) messages, which are supplemented during the encoding process of the headers corresponding to sequences and images, as well as the parameter sets (Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc.) for videos / images. The SEI message is separate from the information for video / images (slice data). VCL including information for video / images includes slice data and slice headers. Furthermore, the slice header can be referred to as the tile group header, and the slice data can be referred to as the tile group data.
[0088] In NAL, NAL cells can be generated by adding header information (NAL cell header) to the raw byte sequence payload (RBSP) generated in VCL. In this case, the RBSP is referred to as slice data, parameter set, SEI message, etc., generated in VCL. The NAL cell header can include NAL cell type information specified according to the RBSP data included in the NAL cell.
[0089] As the basic unit of NAL, the NAL unit performs the function of mapping the encoded image to bit sequences of lower systems such as file formats, real-time transport protocols (RTP), and transport streams (TS) according to predetermined specifications.
[0090] As shown in the figure, based on the RBSP generated in the VCL, the NAL unit can be divided into VCL NAL units and non-VCL NAL units. A VCL NAL unit can refer to a NAL unit that includes information for the image (slice data), while a non-VCL NAL unit can refer to a NAL unit that includes information required for decoding the image (parameter set or SEI message).
[0091] The aforementioned VCL NAL units and non-VCL NAL units can be transmitted over the network by attaching header information according to the data specifications of the lower system. For example, NAL units can be transformed into predetermined data formats such as H.266 / VVC file format, RTP (Real-Time Transport Protocol), TS (Transport Stream), etc.
[0092] As described above, for a NAL cell, the NAL cell type can be specified according to the RBSP data structure included in the NAL cell, and the information of the NAL cell type can be stored in the NAL cell header and signaled.
[0093] For example, NAL units can be classified into VCLNAL unit types and non-VCL NAL unit types based on whether they include information for images (slice data). VCL NAL unit types can be classified according to the nature and type of the images included in the VCL NAL unit, while non-VCL NAL unit types can be classified according to the type of parameter set.
[0094] The following describes an example of specifying a NAL unit type based on the type of the parameter set included in a non-VCL NAL unit type. A NAL unit type can be specified based on the type of the parameter set. For example, the NAL unit type can be specified as one of the following: APS (Adaptive Parameter Set) NAL unit (the type of NAL unit including APS), DPS (Decoding Parameter Set) NAL unit (the type of NAL unit including DPS), VPS (Video Parameter Set) NAL unit (the type of NAL unit including VPS), SPS (Sequence Parameter Set) NAL unit (the type of NAL unit including SPS), and PPS (Picture Parameter Set) NAL unit (the type of NAL unit including PPS).
[0095] The aforementioned NAL unit type can have syntax information specific to the NAL unit type, and this syntax information can be stored in the NAL unit header and signaled. For example, the syntax information can be nal_unit_type, and the NAL unit type can be specified as the nal_unit_type value.
[0096] Furthermore, as described above, an image can have multiple slices, and a slice can include a slice header and slice data. In this case, in addition to multiple slices (a set of slice headers and slice data), an image header can also be added to an image. The image header (image header syntax) can include information / parameters that can be applied collectively to the image. The slice header (slice header syntax) can include information / parameters that can be applied collectively to the slice. APS (APS syntax) or PPS (PPS syntax) can include information / parameters that can be applied collectively to one or more slices or images. SPS (SPS syntax) can include information / parameters that can be applied collectively to one or more sequences. VPS (VPS syntax) can include information / parameters that can be applied collectively to multiple layers. DPS (DPS syntax) can include information / parameters that can be applied collectively to the entire video. DPS can include information / parameters related to the concatenation of CVS (Coded Video Sequence). In this disclosure, the High-Level Syntax (HLS) can include at least one of APS syntax, PPS syntax, SPS syntax, VPS syntax, DPS syntax, image header syntax, or slice header syntax.
[0097] In this disclosure, image / video information encoded by an encoding device and transmitted to a decoding device in bitstream format may include information contained in the slice header, information contained in the image header, information contained in the APS, information contained in the PPS, information contained in the SPS, information contained in the VPS, and / or information contained in the DPS, as well as information related to segmentation in the image, intra / inter-frame prediction information, residual information, loop filtering information, etc. Additionally, the image / video information may also include information from the NAL unit header.
[0098] Figure 5 This is an example illustration of image segmentation.
[0099] An image can be segmented into Coding Tree Units (CTUs), and each CTU can correspond to a Coding Tree Block (CTB). A CTU may include a CTU for luminance samples and two corresponding CTUs for chrominance samples. Furthermore, the maximum available size of a CTU used for encoding and prediction may differ from the maximum available size of a CTU used for transformation.
[0100] A tile can correspond to a series of CTUs covering a rectangular area, and an image can be divided into one or more tile rows and one or more tile columns.
[0101] Furthermore, a slice can include an integer number of intact tiles or an integer number of consecutive intact CTU columns. In this case, two types of slicing modes can be supported, including raster scan slicing mode and rectangular slicing mode.
[0102] In raster scan slicing mode, a slice can include a series of intact tiles in a tile raster scan of an image. In rectangular slicing mode, a slice can include multiple intact tiles that together form a rectangular area of the image. Alternatively, in rectangular slicing mode, a slice can include multiple consecutive CTU columns of tiles that together form a rectangular area of the image. Tiles in a rectangular slice can be scanned in tile raster scan order within the rectangular area corresponding to the respective slice.
[0103] In addition, a sub-image may include one or more slices that cover a rectangular area of the image.
[0104] Figure 5 (a) is an illustration of an example of an image being segmented into raster scan slices. For example, the image can be segmented into 12 tiles and 3 raster scan slices.
[0105] in addition, Figure 5 (b) is an illustration of an example of an image being divided into rectangular slices. For example, the image can be divided into 24 tiles (6 tile rows and 4 tile columns) and 9 rectangular slices.
[0106] also, Figure 5 (c) is an illustration of an example of an image being divided into tiles and rectangular slices. For example, the image can be divided into 24 tiles (2 tile rows and 2 tile columns) and 4 rectangular slices.
[0107] Figure 6 This is a flowchart illustrating the image encoding process according to an implementation method.
[0108] In one implementation, image segmentation can be performed by the image segmenter 210 of the encoding device (step S600), and image encoding can be performed by the entropy encoder 240 of the encoding device (step S610).
[0109] The encoding device according to the embodiment can deduce the slices and / or tiles included in the current image (step S600). For example, the encoding device can perform image segmentation to encode the input current image. For example, the encoding device can deduce the slices and / or tiles included in the current image. The encoding device can segment the image in various formats by considering the image properties and encoding efficiency of the current image, and generate information representing the segmentation format with the best encoding efficiency, and then signal this information to the decoding device.
[0110] The encoding device according to the embodiment can encode the current image based on derived slices and / or tiles (step S610). For example, the encoding device can encode video / image information including slice and / or tile information and output the information in bitstream format. The output bitstream can be forwarded to the decoding device via a digital storage medium or network.
[0111] Figure 7 This is a flowchart illustrating the image decoding process according to an implementation method.
[0112] In the implementation, the entropy decoder 310 of the decoding device may perform the step of obtaining video / image information from the bitstream (step S710) and the step of deriving slices and / or tiles in the current image (step S720), and the adder 340 of the decoding device may perform the step of reconstructing the current image based on slices and / or tiles.
[0113] The decoding device according to the embodiment can obtain video / image information from the received bitstream (step S710). The video / image information may include HLS, and the HLS may include slice information or tile information. The slice information may include information for specifying one or more slices in the current image, and the tile information may include information for specifying one or more tiles in the current image. The slice information or tile information can be obtained through various parameter sets, image headers, and / or slice headers.
[0114] In addition, the current image may include a tile containing one or more slices or a slice containing one or more tiles.
[0115] According to the implementation, the decoding device can deduce slices and / or tiles based on video / image information including information on slices and / or tiles in the current image (step S720).
[0116] The decoding device according to the implementation can reconstruct (decode) the current image based on slices and / or tiles (step S730).
[0117] Furthermore, as mentioned above, the image can be segmented into sub-images, tiles, and slices. Information about sub-images can be signaled via SPS, and information about tiles and rectangular slices can be signaled via PPS. Additionally, information about raster scan slices can be signaled via slice hair signals.
[0118] For example, the SPS syntax that includes information about sub-images can be represented as shown in the table below.
[0119] [Table 1]
[0120]
[0121] For example, the PPS syntax, which includes information about tiles and rectangular slices, can be represented as shown in the table below.
[0122] [Table 2]
[0123]
[0124] Additionally, for example, the slice header syntax, which includes information about raster scan slices, can be represented as shown in the table below.
[0125] [Table 3]
[0126]
[0127] Furthermore, information about slices and tiles in the current image can include flags related to whether each subpick in the current image contains a single slice. This flag can be called `single_slice_per_subpic_flag` or `pps_single_slice_per_subpic_flag`, but is not limited to these. Additionally, information about subpicks can include flags related to the presence of subpick information, and this flag can be called `subpics_present_flag` or `sps_subpic_info_present_flag`, but is not limited to these. For example, subpick information can be included in a parameter set. For example, subpick information can be included in SPS.
[0128] Conventionally, if the value of the flag related to the existence of sub-image information is zero, the value of the flag is restricted such that the value of the flag related to whether the sub-image includes only one slice is also zero. That is, if the value of the flag related to the existence of sub-image information is zero, the sub-image is determined to be unavailable, and the value of the flag related to whether the sub-image includes only one slice is restricted to zero. However, this condition is very strict. For example, even without sub-image information, the current image can be divided into two or more tiles, and all tiles can be included in a single slice. In this case, the current image includes only one slice.
[0129] Therefore, embodiments of this disclosure propose a method to eliminate the limitation that the value of a flag related to whether a sub-image includes only one slice becomes zero when the value of a flag related to the presence of sub-image information is zero. In this case, the flag related to whether a sub-image includes only one slice can indicate that the current image includes only one slice even when sub-image information is absent.
[0130] For example, according to the implementation, even when sub-picture information is absent in the coded layer video sequence (CLVS), a flag related to whether a sub-picture includes only one slice can exist. That is, even when sub-picture information is absent in the CLVS, the flag related to whether a sub-picture includes only one slice can have a value of zero or 1.
[0131] For example, if the value of the flag related to the existence of sub-image information is zero and the value of the flag related to whether the sub-image contains only one slice is 1, the current image may contain only one slice. That is, if there is no sub-image that sends a signal and the value of the flag related to whether the sub-image contains only one slice is 1, the number of slices in the image can be inferred to be 1.
[0132] Furthermore, if the value of the flag related to the existence of sub-image information is zero, the number of sub-images in the current image can be 1. For example, if the value of the flag related to the existence of sub-image information is zero, the number of sub-images existing in each of all images in the SPS referring to the image information can be 1.
[0133] Furthermore, a flag related to the number of slices included in the current image can be included in the PPS of the image information. This flag can be called `num_slices_in_pic_minus1` or `pps_num_slices_in_pic_minus1`, but is not limited to these names. Additionally, a flag related to the number of subpics included in the current image can be included in the SPS of the image information. This flag can be called `sps_num_subpics_minus1`, but is not limited to these names.
[0134] If no sub-image information exists and the flag related to whether a sub-image includes only one slice is valued at 1, the flag related to the number of slices included in the current image can be inferred to have a value of zero. Furthermore, if no sub-image information exists and the flag related to whether a sub-image includes only one slice is valued at 1, the flags related to the number of slices included in the current image and the flags related to the number of sub-images included in the current image can be inferred to have the same value.
[0135] Additionally, if the value of the flag related to whether a sub-image includes only one slice is 1, all CTUs in the image can belong to only one slice included in the image.
[0136] The semantics of the syntax elements, including flags related to whether a sub-image contains only one slice and flags related to the number of slices included in the current image, can be represented as shown in the table below.
[0137] [Table 4]
[0138]
[0139] Referring to the table above, when the value of `single_slice_per_subpic_flag` corresponding to the flag related to whether a subpicture includes only one slice is 1, each subpicture can include a single rectangular slice. Conversely, when the value of `single_slice_per_subpic_flag` is zero, each subpicture can include one or more rectangular slices. When the value of `single_slice_per_subpic_flag` is 1, the value of `num_slices_in_pic_minus1` corresponding to the flag related to the number of slices included in the current picture can be inferred to have the value of `sps_num_subpics_minus1` corresponding to the flag related to the number of subpictures included in the current picture.
[0140] Furthermore, when the value of single_slice_per_subpic_flag is 1 and the value of subpics_present_flag, which corresponds to the flag related to the existence of subpic information, is zero, the images referenced by PPS can each have a single slice.
[0141] In addition, the scanning process can be determined according to the following table as the order in which the tiles in the image are decoded.
[0142] [Table 5]
[0143]
[0144]
[0145] Figure 8 This is a flowchart illustrating the operation of an encoding device according to an embodiment, and Figure 9 This is a block diagram illustrating the configuration of an encoding device according to an embodiment.
[0146] Figure 8 The method shown can be derived from Figure 2 or Figure 9 The encoding device shown in the figure performs the operation. Figure 9 Step S810 shown can be performed by Figure 2 The image segmenter 210 shown in the figure is executed, and step S820 can be performed by... Figure 2The entropy encoder 240 shown in the figure is executed. Furthermore, the operation according to steps S810 and S820 is based on a reference. Figures 1 to 7 This is part of the description of the instruction manual. Therefore, descriptions related to the reference are omitted or briefly described. Figures 1 to 7 The description overlaps with the detailed description in the instruction manual.
[0147] Reference Figure 8 According to the embodiment, the encoding device can segment the current image and derive at least one slice (step S810). For example, the image segmenter 210 of the encoding device can generate segmentation information of the current image based on at least one slice.
[0148] The encoding device according to the embodiment can encode image information of the current image based on at least one slice (step S810). The image information may include segmentation information generated based on at least one slice.
[0149] For example, image information may include a first flag related to the existence of sub-image information and a second flag related to whether the sub-image includes only one slice. For example, based on the first and second flags, the number of slices included in the current image can be deduced to be equal to 1.
[0150] For example, if the value of the first flag related to the existence of sub-image information is zero and the value of the second flag is 1, the number of slices included in the current image can be deduced to be equal to 1.
[0151] For example, if the value of the first flag related to the existence of sub-image information is zero, the number of sub-images existing in the current image can be 1.
[0152] For example, a first flag related to the existence of sub-image information can be included in the SPS (Sequence Parameter Set) of the image information.
[0153] For example, a second flag related to whether a sub-image includes only one slice can be included in the image information's PPS (Picture Parameter Set).
[0154] For example, the image information may include a third flag related to the number of slices included in the current image, and the third flag may be included in the PPS of the image information.
[0155] Additionally, for example, the image information may include a fourth flag related to the number of sub-images included in the current image, and the fourth flag may be included in the SPS of the image information.
[0156] Additionally, for example, when the value of the first flag is zero, the number of sub-pictures present in each of all pictures in the SPS referring to the image information can be derived to be equal to 1.
[0157] In addition, the image information may include prediction information for the current image. The prediction information may include information about the inter-frame prediction mode or intra-frame prediction mode performed in the current image. The encoding device can generate the prediction information for the current image and encode it.
[0158] In addition, the bitstream can be sent to the decoding device via a network or (digital) storage medium. Here, the network can include broadcast networks and / or communication networks. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc.
[0159] Figure 10 This is a flowchart illustrating the operation of an encoding device according to an embodiment, and Figure 11 This is a block diagram illustrating the configuration of a decoding device according to an embodiment.
[0160] Figure 10 The method disclosed in the article can be derived from Figure 3 or Figure 11 The decoding device shown in the diagram performs the operation. Specifically, steps S1010 and S1020 can be performed by... Figure 3 The entropy decoder 310 shown in the figure is executed. Furthermore, the operations according to steps S1010 to S1020 are based on a reference. Figures 1 to 7 This is part of the description of the instruction manual. Therefore, descriptions related to the reference are omitted or briefly described. Figures 1 to 7 The description overlaps with the detailed description in the instruction manual.
[0161] According to the implementation, the decoding device can obtain image information of the current image from the bitstream (step S1010). For example, the entropy decoder 310 of the decoding device can obtain image information including segmentation information of the current image from the bitstream. For example, the segmentation information may include slice information of the current image. In addition, the image information may include at least a portion of prediction-related information or residual-related information. For example, the prediction-related information may include inter-frame prediction mode information or inter-frame prediction type information.
[0162] The decoding device according to the embodiment can decode the current image based on at least image information (step S1020). For example, the entropy decoder 310 of the decoding device can deduce the segmentation structure of the current image based on the slice information of the current image.
[0163] For example, image information may include a first flag related to the existence of sub-image information and a second flag related to whether the sub-image includes only one slice. For example, based on the first and second flags, the number of slices included in the current image can be deduced to be equal to 1.
[0164] For example, if the value of the first flag is zero and the value of the second flag is 1, the number of slices included in the current image can be deduced to be equal to 1.
[0165] For example, if the value of the flag related to the existence of sub-image information is zero, the number of sub-images existing in the current image can be 1.
[0166] For example, a first flag related to the existence of sub-image information can be included in the SPS (Sequence Parameter Set) of the image information.
[0167] For example, a second flag related to whether a sub-image includes only one slice can be included in the image information's PPS (Picture Parameter Set).
[0168] For example, the image information may include a third flag related to the number of slices included in the current image, and the third flag may be included in the PPS of the image information.
[0169] Additionally, for example, the image information may include a fourth flag related to the number of sub-images included in the current image, and the flag related to the number of sub-images included in the current image may be included in the SPS of the image information.
[0170] Additionally, for example, when the value of the first flag is zero, the number of sub-pictures present in each of all pictures in the SPS referring to the image information can be derived to be equal to 1.
[0171] Although the method has been described based on a flowchart listing the steps or blocks in the above embodiments, the steps in this document are not limited to a specific order, and specific steps may be performed in different steps or in different orders or simultaneously relative to the steps described above. Furthermore, those skilled in the art will understand that the steps in the flowchart are not exclusive, and one or more steps may be included or removed from the flowchart without affecting the scope of this disclosure.
[0172] The methods mentioned above according to this disclosure can be in the form of software, and the encoding and / or decoding devices according to this disclosure can be included in an apparatus for performing image processing (e.g., TV, computer, smartphone, set-top box, display device, etc.).
[0173] When the embodiments of this disclosure are implemented in software, the methods described above can be implemented using modules (processes or functions) that perform the functions mentioned above. Modules can be stored in memory and executed by a processor. Memory can be installed internally or externally to the processor and can be connected to the processor via various known means. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. Memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, embodiments of this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in the corresponding figures can be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information about the implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0174] Furthermore, the decoding and encoding devices using the embodiments described in this document can be included in multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices such as video communication, mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, vehicle-mounted terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, or ship terminals), and medical video devices; and can be used to process image signals or data. For example, OTT video devices can include game consoles, Blu-ray players, networked TVs, home theater systems, smartphones, tablet PCs, and digital video recorders (DVRs).
[0175] Furthermore, the processing methods described in this document can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data with data structures according to this document can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. Computer-readable recording media also include media implemented in the form of carrier waves (e.g., transmission over the Internet). Additionally, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks.
[0176] Furthermore, the embodiments described in this document can be implemented as a computer program product using program code, and the program code can be executed by a computer according to the embodiments described in this document. The program code can be stored on a computer-readable medium.
[0177] Figure 12 Examples of content streaming systems to which the implementation methods disclosed in this document can be applied are illustrated.
[0178] Reference Figure 12 The content streaming system implemented in this document can generally include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0179] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generate a bitstream, and then transmit it to a streaming server. As another example, if the multimedia input device, such as a smartphone, camera, or camcorder, directly generates the bitstream, the encoding server can be omitted.
[0180] The encoding method or bitstream generation method applied to the embodiments described in this document can be used to generate bitstreams. Furthermore, the streaming server can temporarily store the bitstream during the sending or receiving process.
[0181] A streaming server transmits multimedia data to a user's device via a web server based on a user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server forwards the request to the streaming server, which then delivers the multimedia data to the user. In this respect, the content streaming system may include a separate control server, which in this case controls the commands / responses between the various devices within the content streaming system.
[0182] A streaming server can receive content from media storage and / or encoding servers. For example, if content is received from an encoding server, it can be received in real time. In this case, the streaming server can store the bitstream for a predetermined period of time to provide a smooth streaming service.
[0183] For example, user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation systems, board PCs, tablet PCs, ultrabooks, wearable devices (e.g., watch-type terminals (smartwatches), glasses-type terminals (smart glasses), head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0184] Each server in the content streaming system can be operated as a distributed server, and in this case, the data received by each server can be processed in a distributed manner.
[0185] The claims of this disclosure can be combined in various ways. For example, the technical features in the method claims of this disclosure can be combined to be implemented or performed in a device, and the technical features in the device claims can be combined to be implemented or performed in a method. Furthermore, the technical features in the method claims and the device claims can be combined to be implemented or performed in a device.
Claims
1. A decoding device for image decoding, the decoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: Obtain image information about the current image from the bitstream; and The current image is decoded based on the image information. The image information includes a first flag related to the existence of sub-image information and a second flag related to whether each sub-image includes only one slice. Based on the first and second flags, the number of slices included in the current image is deduced to be equal to 1. Specifically, based on the fact that the value of the first flag is equal to 0 and the value of the second flag is equal to 1, the number of slices included in the current image is deduced to be equal to 1. Wherein, a value of 1 for the second flag indicates that each sub-image consists of one and only one slice, and The first flag is included in the sequence parameter set.
2. An encoding device for image encoding, the encoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: At least one slice is derived by segmenting the current image; and Image information for the current image is encoded based on the at least one slice. The image information includes a first flag related to the existence of sub-image information and a second flag related to whether each sub-image includes only one slice. Based on the first and second flags, the number of slices included in the current image is deduced to be equal to 1. Specifically, based on the fact that the value of the first flag is equal to 0 and the value of the second flag is equal to 1, the number of slices included in the current image is deduced to be equal to 1. Wherein, a value of 1 for the second flag indicates that each sub-image consists of one and only one slice, and The first flag is included in the sequence parameter set.
3. An apparatus for transmitting data for an image, the apparatus comprising: At least one processor, configured to obtain a bitstream generated by a method comprising: performing: deriving at least one slice by segmenting a current image and generating the bitstream by encoding image information for the current image based on the at least one slice; and A transmitter configured to transmit the data containing the bit stream. The image information includes a first flag related to the existence of sub-image information and a second flag related to whether each sub-image includes only one slice. Based on the first and second flags, the number of slices included in the current image is deduced to be equal to 1. Specifically, based on the fact that the value of the first flag is equal to 0 and the value of the second flag is equal to 1, the number of slices included in the current image is deduced to be equal to 1. Wherein, the value of the second flag is equal to 1, indicating that each sub-image consists of one and only one slice, and wherein the first flag is included in the sequence parameter set.