Image encoding / decoding method and device, and recording medium storing bit stream
By employing a tree-based block segmentation and adaptive scanning order image coding method, the problem of low efficiency in high-resolution image coding is solved, achieving a more efficient encoding and decoding process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2024-09-12
- Publication Date
- 2026-04-21
AI Technical Summary
Existing image compression techniques are inefficient in encoding high-resolution and high-quality images, especially in terms of information notification for adaptive decoding/encoding order.
A tree-based block partitioning method is adopted to divide the current block into multiple sub-blocks. By using index information and high-level syntax signals to notify the adaptive scanning order through predefined decoding and encoding orders, the encoding efficiency is improved.
By applying adaptive decoding/encoding order, the efficiency of image encoding is improved, especially in the configuration of top-level block units, which enables a more efficient encoding and decoding process.
Smart Images

Figure CN121909645A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an image encoding / decoding method and apparatus, as well as a recording medium for storing bit streams. Background Technology
[0002] Recently, the demand for high-resolution and high-quality images, such as HD (high-definition) and UHD (ultra-high-definition) images, has been increasing in various application areas, and therefore, efficient image compression technologies are being discussed.
[0003] Various techniques exist, such as inter-frame prediction techniques that use video compression technology to predict pixel values included in the current frame from frames before or after the current frame, intra-frame prediction techniques that use pixel information in the current frame to predict pixel values included in the current frame, and entropy coding techniques that assign short symbols to values that occur frequently and long symbols to values that occur infrequently. These image compression techniques can be used to effectively compress image data and send or store it. Summary of the Invention
[0004] Technical issues
[0005] This disclosure aims to provide a method and apparatus for using an adaptive decoding / encoding order for sub-blocks based on tree-based block segmentation.
[0006] This disclosure aims to provide a method and apparatus for efficiently signaling information / grammar using adaptive decoding / encoding order.
[0007] According to this disclosure, a method and apparatus are provided for using an adaptive decoding / encoding order for a top-level block unit that configures a picture.
[0008] Technical solution
[0009] The image decoding method and apparatus according to this disclosure can divide a current block into multiple sub-blocks based on one of a predefined block segmentation type, and decode the multiple sub-blocks according to a predetermined decoding order. Here, the decoding order can be determined based on one or more scan order candidates. The one or more scan order candidates may include at least one of a horizontal Z-scan order or a vertical Z-scan order.
[0010] In the image decoding method and apparatus according to the present invention, the decoding order can be determined based on index information indicating one of the one or more scan order candidates.
[0011] In the image decoding method and apparatus according to this disclosure, index information can be signaled based on information indicating whether the current block is divided into four sub-blocks.
[0012] The image decoding method and apparatus disclosed herein can obtain high-level syntax regarding whether to use adaptive scan order from the bitstream.
[0013] In the image decoding method and apparatus according to the present disclosure, a high-level syntax can be notified by a signal in at least one of a sequence parameter set, a picture parameter set, or a picture header.
[0014] In the image decoding method and apparatus according to this disclosure, index information can be signaled based on high-level syntax.
[0015] In the image decoding method and apparatus according to the present disclosure, a current block including the current block can be divided into multiple top-level block units, and the multiple top-level block units can be decoded based on a decoding order that combines the raster scan order and the Z scan order.
[0016] The image encoding method and apparatus according to this disclosure can divide a current block into multiple sub-blocks based on one of a predefined block segmentation type, and encode the multiple sub-blocks according to a predetermined encoding order. Here, the encoding order can be determined based on one or more scan order candidates. The one or more scan order candidates may include at least one of a horizontal Z-scan order or a vertical Z-scan order.
[0017] A computer-readable digital storage medium is provided for storing encoded video / image information, which enables a decoding device according to the present disclosure to perform an image decoding method.
[0018] A computer-readable digital storage medium is provided for storing video / image information generated based on the image encoding method of this disclosure.
[0019] A method and apparatus are provided for transmitting video / image information generated according to the image encoding method of this disclosure.
[0020] Beneficial effects
[0021] According to this disclosure, encoding efficiency can be improved by using an adaptive decoding / encoding order for sub-blocks obtained through tree-based block partitioning.
[0022] According to this disclosure, coding efficiency can be improved by efficiently signaling the syntax for using adaptive decoding / encoding order.
[0023] According to this disclosure, encoding efficiency can be improved by using an adaptive decoding / encoding order for the top-level block unit configured for a screen. Attached Figure Description
[0024] Figure 1 A video / image encoding system according to this disclosure is shown.
[0025] Figure 2 A schematic block diagram of an encoding apparatus to which embodiments of the present disclosure are applicable and which performs encoding of video / image signals is shown.
[0026] Figure 3 A schematic block diagram of a decoding apparatus to which embodiments of the present disclosure are applicable and which performs decoding of video / image signals is shown.
[0027] Figure 4 An image decoding method performed by a decoding device 300 as an embodiment of the present disclosure is illustrated.
[0028] Figure 5 This example illustrates the decoding order of multiple CTUs within a single frame, based on the raster scan order.
[0029] Figure 6 and Figure 7 An example is shown where adaptive scan order is used for multiple CTUs within a single frame.
[0030] Figure 8 An illustrative configuration of a decoding device 300 performing an image decoding method according to the present disclosure is shown.
[0031] Figure 9 An image encoding method performed by an encoding device 200 as an embodiment of the present disclosure is illustrated.
[0032] Figure 10 An illustrative configuration of an encoding device 200 performing an image encoding method according to the present disclosure is shown.
[0033] Figure 11 Examples of content streaming systems to which embodiments of the present disclosure can be applied are shown. Detailed Implementation
[0034] Because this disclosure can be modified in various ways and has multiple embodiments, specific embodiments will be shown in the accompanying drawings and described in detail in the specific embodiments. However, this disclosure is not intended to be limited to the specific embodiments, but should be understood to include all changes, equivalents, and substitutions included within the spirit and scope of this disclosure. Similar reference numerals are used for similar components in the description of the various figures.
[0035] Terms such as "first," "second," etc., may be used to describe various components, but components should not be limited by these terms. Terms are used only to distinguish one component from others. For example, without departing from the scope of this disclosure, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component. Terms include any combination of one or more of the associated terms.
[0036] When a component is described as "connected" or "linked" to another component, it should be understood that it can be directly connected or linked to the other component, but the other component can exist in between. On the other hand, when a component is described as "directly connected" or "directly linked" to another component, it should be understood that there is no other component in between.
[0037] The terminology used in this application is for describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, it should be understood that terms such as “comprising” or “having” are intended to specify the presence of the features, quantities, steps, operations, components, portions, or combinations thereof described in this specification, but do not preclude the possibility of the presence or addition of one or more other features, quantities, steps, operations, components, portions, or combinations thereof.
[0038] This disclosure relates to video / image coding. For example, the methods / implementations disclosed herein can be applied to methods disclosed in the Multifunctional Video Coding (VVC) standard. Additionally, the methods / implementations disclosed herein can be applied to methods disclosed in the Basic Video Coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the Audio Video Coding 2 (AVS2) standard, or next-generation video / image coding standards (e.g., H.267 or H.268).
[0039] This specification sets forth various implementations of video / image encoding, and unless otherwise specified, these implementations may be combined with each other.
[0040] In this article, video can refer to a collection of images over time. A frame typically refers to a unit representing an image within a specific time period, and a slice / tile is a unit that forms part of a frame during encoding. A slice / tile can include at least one Code Tree Unit (CTU). A frame can consist of at least one slice / tile. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a frame. A tile column is a rectangular area of CTUs with a height equal to the height of the frame and a width specified by the syntax requirements of the frame parameter set. A tile row is a rectangular area of CTUs with a height specified by the frame parameter set and a width equal to the width of the frame. CTUs within a tile can be arranged continuously according to CTU raster scans, and tiles within a frame can be arranged continuously according to tile raster scans. A slice can include an integer number of complete tiles of a frame that can be exclusively included in a single NAL unit, or an integer number of consecutive complete CTU rows within a tile. Furthermore, a frame can be divided into at least two sub-frames. A sub-frame can be a rectangular area of at least one slice within a frame.
[0041] A pixel, or pelin, can represent the smallest unit that makes up a frame (or image). Additionally, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component.
[0042] A unit can represent the basic unit of image processing. A unit may include a specific region of the image and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "region". In general, an M×N block may include a set (or array) of transform coefficients or samples (or sample arrays) consisting of M columns and N rows.
[0043] In this document, “A or B” can mean “A only”, “B only”, or “both A and B”. In other words, “A or B” can be interpreted as “A and / or B”. For example, “A, B or C” can mean “A only”, “B only”, “C only”, or “any combination of A, B and C”.
[0044] The forward slash ( / ) or comma used in this article can indicate "and / or". For example, "A / B" can mean "A and / or B". Therefore, "A / B" can mean "A only", "B only", or "both A and B". For example, "A, B, C" can mean "A, B, or C".
[0045] In this document, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, in this document, expressions such as "at least one of A or B" or "at least one of A and / or B" can be interpreted in the same way as "at least one of A and B".
[0046] Additionally, in this document, "at least one of A, B, and C" can mean "A only", "B only", "C only" or "any combination of A, B, and C". Furthermore, "at least one of A, B, or C" or "at least one of A, B, and / or C" can mean "at least one of A, B, and C".
[0047] Additionally, the parentheses used in this document can indicate "for example". Specifically, when the indication is "prediction (intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction". In other words, "prediction" in this document is not limited to "intra-frame prediction", and "intra-frame prediction" can be cited as an example of "prediction". Furthermore, even when the indication is "prediction (i.e., intra-frame prediction)", "intra-frame prediction" can be cited as an example of "prediction".
[0048] In this article, the technical features described individually in a single diagram can be implemented individually or simultaneously.
[0049] Figure 1 A video / image encoding system according to this disclosure is shown.
[0050] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device).
[0051] A source device can transmit encoded video / image information or data to a receiving device in the form of a file or stream via a digital storage medium or network. The source device may include a video source, an encoding device, and a transmitting unit. The receiving device may include a receiving unit, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may consist of a separate device or external components.
[0052] A video source can acquire video / images through processes that capture, synthesize, or generate video / images. A video source may include means for capturing video / images and means for generating video / images. Means for capturing video / images may include at least one camera, a video / image archive containing previously captured video / images, etc. Means for generating video / images may include a computer, tablet computer, smartphone, etc., and can generate video / images (electronically). For example, virtual video / images can be generated by a computer, etc., and in this case, the process of capturing video / images can be replaced by a process of generating related data.
[0053] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and encoding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0054] The transmitting unit can send encoded video / image information or data, output in bitstream form, to the receiving unit of the receiving device in the form of a file or stream via a digital storage medium or network. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitting unit can include elements for generating media files according to a predetermined file format, and may include elements for transmission via a broadcast / communication network. The receiving unit can receive / extract the bitstream and send it to a decoding device.
[0055] Decoding devices can decode video / images by performing a series of processes, such as dequantization, inverse transform, and prediction, that correspond to the operations of encoding devices.
[0056] The renderer can render decoded video / images. The rendered video / images can be displayed through a display unit.
[0057] Figure 2 A rough block diagram of an encoding apparatus that can be applied to embodiments of the present disclosure and perform encoding of video / image signals is shown.
[0058] Reference Figure 2 The encoding device 200 may consist of an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may also include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 may be configured by at least one hardware component (e.g., an encoder chipset or processor). Additionally, the memory 270 may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may also include the memory 270 as an internal / external component.
[0059] Image segmenter 210 can segment an input image (or picture or frame) input to encoding device 200 into at least one processing unit. As an example, a processing unit can be referred to as a coding unit (CU). In this case, the coding unit can be recursively segmented from coding tree unit (CTU) or maximum coding unit (LCU) according to a quadtree-binary-tritree (QTBTTT) structure.
[0060] For example, a coding unit can be segmented into multiple deeper coding units based on a quadtree, binary tree, and / or ternary tree structure. In this case, for example, a quadtree structure can be applied first, followed by a binary tree and / or ternary tree structure. Alternatively, a binary tree structure can be applied before the quadtree structure. The coding process according to this specification can be performed based on the final coding unit that is no longer segmented. In this case, based on image characteristics, coding efficiency, etc., the largest coding unit can be directly used as the final coding unit, or, if necessary, the coding unit can be recursively segmented into deeper coding units, and the coding unit with the optimal size can be used as the final coding unit. Here, the coding process can include processes such as prediction, transformation, and reconstruction, as described later.
[0061] As another example, the processing unit may also include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit can be divided or segmented from the aforementioned final encoding unit, respectively. The prediction unit may be a unit for predicting samples, and the transform unit may be a unit for deriving transform coefficients and / or a unit for deriving residual signals from transform coefficients.
[0062] In some cases, a unit can be used interchangeably with terms such as block or region. Generally, an M×N block can represent a set of transform coefficients or samples consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Samples can be used as a term to form a frame (or image) corresponding to a pixel or cell.
[0063] Encoding device 200 can subtract the prediction signal (prediction block, prediction sample array) output from inter-frame predictor 221 or intra-frame predictor 222 from the input image signal (original block, original sample array) to generate a residual signal (residual signal, residual sample array), and the generated residual signal is sent to converter 232. In this case, the unit in encoding device 200 that subtracts the prediction signal (prediction block, prediction sample array) from the input image signal (original block, original sample array) can be called subtractor 231.
[0064] Predictor 220 can perform prediction on the block to be processed (hereinafter referred to as the current block) and generate a prediction block that includes prediction samples of the current block. Predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction on a per-block or per-unit basis. Predictor 220 can generate various information about the prediction (e.g., prediction mode information) and send it to entropy encoder 240, as described later in the description of the various prediction modes. The information about the prediction can be encoded in entropy encoder 240 and output as a bitstream.
[0065] Intra-predictor 222 can predict the current block by referencing samples within the current frame. Depending on the prediction mode, the referenced samples can be located near the current block or positioned at a specific distance away from the current block. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. The non-directional mode can include at least one of a DC mode or a planar mode. Depending on the level of detail of the prediction direction, the directional modes can include 33 or 65 directional modes. However, this is just an example; more or fewer directional modes can be used depending on the configuration. Intra-predictor 222 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0066] Inter-frame predictor 221 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. The reference frame including the reference block and the reference frame including the temporally neighboring block may be the same or different. The temporally neighboring block may be referred to as a co-located reference block, a co-located CU (colCU), etc., and the reference frame including the temporally neighboring block may be referred to as a co-located frame (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate to use to deduce the motion vector and / or reference frame index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in skip mode and merge mode, the inter-frame predictor 221 can use motion information from neighboring blocks as motion information for the current block. In skip mode, unlike merge mode, residual signals may not be sent. In motion vector prediction (MVP) mode, motion vectors from neighboring blocks are used as motion vector predictors, and the motion vector difference is signaled to indicate the motion vector of the current block.
[0067] Predictor 220 can generate a prediction signal based on various prediction methods described later. For example, the predictor can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply both intra-frame and inter-frame prediction simultaneously. This can be referred to as the Inter-intra-frame Combined Prediction (CIIP) mode. Alternatively, the predictor can predict blocks based on the Intra-Block Copy (IBC) prediction mode, or it can predict blocks based on a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as Screen Content Coding (SCC) in games, etc. IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When applying a palette mode, sample values within the frame can be signaled based on information about the palette table and palette index. The prediction signal generated by predictor 220 can be used to generate a reconstructed signal or a residual signal.
[0068] Transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques may include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen–Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT represents the transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT represents the transform obtained based on generating a prediction signal using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or it can be applied to non-square blocks of variable size.
[0069] Quantizer 233 can quantize the transform coefficients and send them to entropy encoder 240, which can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. This information about the quantized transform coefficients can be called residual information. Quantizer 233 can rearrange the block-form quantized transform coefficients into a one-dimensional vector based on the coefficient scan order, and can generate information about the quantized transform coefficients based on this one-dimensional vector form.
[0070] The entropy encoder 240 can perform various encoding methods such as Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). The entropy encoder 240 can encode information required for video / image reconstruction other than quantization transform coefficients (e.g., values of syntax elements, etc.) together or separately.
[0071] Encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form at the Network Abstraction Layer (NAL) unit level. The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. Information and / or syntax elements transmitted from the encoding device to / signaled to the decoding device may be included in the video / image information. The video / image information can be encoded and included in the bitstream through the encoding process described above. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks and / or communication networks, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitting unit (not shown) for transmitting the signal output from the entropy encoder 240 and / or a storage unit (not shown) for storing the signal may be configured as internal / external components of the encoding device 200, or the transmitting unit may also be included in the entropy encoder 240.
[0072] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, the residual signal (residual block or residual sample) can be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients via dequantizer 234 and inverse transformer 235. Adder 250 can add the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222 to generate a reconstructed signal (reconstructed frame, reconstructed block, reconstructed sample array). When there is no residual for the block to be processed (similar to when a skip mode is applied), the prediction block can be used as a reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, and can also be used for inter-frame prediction of the next frame by filtering, as described later. Furthermore, luminance mapping and chroma scaling (LMCS) can be applied in frame encoding and / or reconstruction processing.
[0073] Filter 260 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270 (specifically, the DPB of memory 270). Various filtering methods can include deblocking filtering, sample adaptive offsetting, adaptive loop filtering, bilateral filtering, etc. Filter 260 can generate various information about the filtering and send it to entropy encoder 240. The information about the filtering can be encoded in entropy encoder 240 and output as a bitstream.
[0074] The modified reconstructed frame sent to memory 270 can be used as a reference frame in inter-frame predictor 221. When inter-frame prediction is applied through it, the encoding device can avoid prediction mismatch between the encoding device 200 and the decoding device, and can also improve encoding efficiency.
[0075] The DPB of memory 270 can store modified reconstructed frames for use as reference frames in inter-frame predictor 221. Memory 270 can store motion information of blocks in the current frame from which motion information is derived (or encoded) and / or of blocks in previously reconstructed frames. The stored motion information can be sent to inter-frame predictor 221 as motion information for spatially or temporally neighboring blocks. Memory 270 can store reconstructed samples of reconstructed blocks in the current frame and send them to intra-frame predictor 222.
[0076] Figure 3 A rough block diagram of a decoding device that can be implemented using embodiments of the present disclosure and perform decoding of video / image signals is shown.
[0077] Reference Figure 3 The decoding device 300 can be configured to include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 332 and an intra-frame predictor 331. The residual processor 320 may include a dequantizer 321 and an inverse transformer 322.
[0078] According to the implementation, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above can be configured by a single hardware component (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded screen buffer (DPB) and can be configured by a digital storage medium. The hardware component may also include the memory 360 as an internal / external component.
[0079] When the input includes a bitstream containing video / image information, the decoding device 300 can respond to... Figure 2 The encoding device processes video / image information to reconstruct the image. For example, the decoding device 300 can deduce units / blocks based on block segmentation information obtained from the bitstream. The decoding device 300 can perform decoding using processing units applied in the encoding device. Therefore, the decoding processing unit can be an encoding unit, and the encoding unit can be segmented from the encoding tree unit or a larger encoding unit according to a quadtree structure, binary tree structure, and / or ternary tree structure. At least one transform unit can be derived from the encoding unit. Furthermore, the reconstructed image signal decoded and output by the decoding device 300 can be played back by a playback device.
[0080] Decoding device 300 can receive data in bitstream form from... Figure 2 The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive the information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), or Video Parameter Set (VPS). In addition, the video / image information may also include general constraint information. The decoding device can also decode the picture based on the information about the parameter sets and / or general constraint information. The information and / or syntax elements that are signaled / received, as described later herein, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element from the bitstream, determine a context model using information about the syntax element to be decoded, decoding information of neighboring blocks and the block to be decoded, or information about symbols / bins decoded in previous steps, perform arithmetic decoding of bins by predicting the occurrence probability of bins based on the determined context model, and generate symbols corresponding to the values of each syntax element. In this case, after determining the context model, the CABAC entropy decoding method can update the context model by using information about the decoded symbols / bins for the context model of the next symbol / bin. Among the information decoded in the entropy decoder 310, prediction information is provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantization transform coefficients and related parameter information) from which entropy decoding has been performed in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). In addition, filtering information among the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, the receiving unit (not shown) that receives the signal output from the encoding device can be further configured as an internal / external component of the decoding device 300, or the receiving unit can be a component of the entropy decoder 310.
[0081] Furthermore, the decoding device according to this specification may be referred to as a video / image / screen decoding device, and the decoding device may be divided into an information decoder (video / image / screen information decoder) and a sample decoder (video / image / screen sample decoder). The information decoder may include an entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0082] Dequantizer 321 can dequantize the quantized transform coefficients and output the transform coefficients. Dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the coefficient scan order performed in the encoding device. Dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0083] The inverse transformer 322 performs an inverse transformation on the transformation coefficients to obtain the residual signal (residual block, residual sample array).
[0084] Predictor 320 can perform prediction on the current block and generate a prediction block that includes prediction samples of the current block. Predictor 320 can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on the prediction information output from entropy decoder 310, and determine the specific intra-frame / inter-frame prediction mode.
[0085] Predictor 320 can generate prediction signals based on various prediction methods described later. For example, predictor 320 can not only apply intra-frame prediction or inter-frame prediction to predict a block, but can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as the Inter-Frame Intra-Frame Combined Prediction (CIIP) mode. Alternatively, the predictor can predict blocks based on the Intra-Frame Block Copy (IBC) prediction mode, or it can predict blocks based on a palette mode. The IBC prediction mode or palette mode can be used for content image / video coding such as Screen Content Coding (SCC) in games, etc. IBC essentially performs prediction within the current frame, but it can be performed similarly to inter-frame prediction because it derives a reference block within the current frame. In other words, IBC can use at least one of the inter-frame prediction techniques described herein. The palette mode can be considered an example of intra-frame coding or intra-frame prediction. When the palette mode is applied, information about the palette table and palette index can be included in the video / image information and signaled.
[0086] Intra-predictor 331 can predict the current block by referencing samples within the current frame. Depending on the prediction mode, the referenced samples can be located near the current block or at a specific distance away. In intra-prediction, the prediction mode can include at least one non-directional mode and multiple directional modes. Intra-predictor 331 can determine the prediction mode applied to the current block by using prediction modes applied to neighboring blocks.
[0087] Inter-frame predictor 332 can deduce the predicted block of the current block based on a reference block (reference sample array) specified by motion vectors on a reference frame. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted on a block, sub-block, or sample basis based on the correlation of motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference frame indices. Motion information may also include inter-frame prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). For inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current frame and temporally neighboring blocks existing in the reference frame. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference frame index of the current block based on the received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the inter-frame prediction mode of the current block.
[0088] Adder 340 can add the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from the predictor (including inter-frame predictor 332 and / or intra-frame predictor 331) to generate a reconstruction signal (reconstructed frame, reconstruction block, reconstruction sample array). When there is no residual for the block to be processed (similar to when a skip mode is applied), the prediction block can be used as a reconstruction block.
[0089] Adder 340 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next block to be processed in the current frame, can be output through filtering as described later, or can be used for inter-frame prediction of the next frame. In addition, luminance mapping and chroma scaling (LMCS) can be applied in the frame decoding process.
[0090] Filter 350 can improve subjective / objective image quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image and send the modified reconstructed image to memory 360 (specifically, the DPB of memory 360). Various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0091] The (modified) reconstructed frame stored in the DPB of memory 360 can be used as a reference frame in inter-frame predictor 332. Memory 360 can store motion information of blocks in the current frame from which motion information is derived (or decoded) and / or motion information of blocks in previously reconstructed frames. The stored motion information can be sent to inter-frame predictor 332 as motion information of spatially or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current frame and send them to intra-frame predictor 331.
[0092] The embodiments described in this document in the filter 260, inter-frame predictor 221 and intra-frame predictor 222 of the encoding device 200 can also be applied equivalently or correspondingly to the filter 350, inter-frame predictor 332 and intra-frame predictor 331 of the decoding device 300, respectively.
[0093] Figure 4 An image decoding method performed by a decoding device 300 as an embodiment of the present disclosure is illustrated.
[0094] Reference Figure 4 This allows the current block to be divided into multiple sub-blocks (S400).
[0095] According to this disclosure, the current block can be an encoded block or a transform block. Each of the plurality of sub-blocks can be an encoded block or a transform block whose width or height is at least one of the width or height of the current block. The segmentation depth of each sub-block can be greater than the segmentation depth of the current block. As an example, when the segmentation depth of the current block is N, the segmentation depth of each sub-block can be (N+1). The plurality of sub-blocks can be squares or rectangles. At least one of the plurality of sub-blocks can have a geometric shape. In this disclosure, the term sub-block can be replaced with terms such as sub-block or partition.
[0096] The current block can be divided based on one or more dividing lines. Dividing lines can include at least one of vertical lines, horizontal lines, or dividing lines with a predetermined angle. At least one of the dividing lines passes through the center of the current block. Alternatively, the dividing lines may not pass through the center of the current block.
[0097] Segmentation can be performed based on one of the predefined block segmentation types. The block segmentation type can include at least one of quad split, binary split, or ternary split. Quad split can be a type that divides a coded block into four sub-blocks. Here, the width and height of the four sub-blocks can be half the width and half the height of the coded block, respectively. Binary split can be a type that divides a coded block into two sub-blocks. Ternary split can be a type that divides a coded block into three sub-blocks. Tree-based block segmentation can be used in the image decoding method according to this disclosure.
[0098] The segmentation information used for block partitioning can be signaled. The segmentation information may include at least one of the following: information indicating whether partitioning is performed, information indicating whether a block is partitioned into four sub-blocks, information indicating whether a block is partitioned into two sub-blocks, or information indicating the partitioning direction.
[0099] As an example, a signal (split_cu_flag) can be used to indicate whether the current block has been split. When split_cu_flag is false (or 0), it indicates that the current block has not been split. In this case, decoding can be performed based on the shape and / or size of the current block without additional block splitting. On the other hand, when split_cu_flag is true (or 1), it indicates that the current block has been split into sub-blocks using additional block splitting. Additional block splitting can be performed based on quad splitting, binary splitting, or triangular splitting. Alternatively, additional block splitting can be performed based on asymmetric binary splitting.
[0100] When `split_cu_flag` is true, `split_qt_flag` can be signaled during the signaling process for the type of additional block splitting for the current block. `split_qt_flag` indicates whether the current block is split based on a quad split. When `split_qt_flag` is true (or 1), the current block can be split into four sub-blocks. On the other hand, when `split_qt_flag` is false (or 0), the current block may not be split into four sub-blocks and can be split into sub-blocks using another block split. However, `split_qt_flag` can only be signaled if the current block satisfies all the conditions for allowing quad splitting; otherwise, `split_qt_flag` can be set to false.
[0101] When `split_qt_flag` is false, information about block splitting (excluding quadruple splitting) can be signaled for the current block. At least one of `mtt_split_cu_vertical_flag` or `mtt_split_cu_binary_flag` can be signaled. `mtt_split_cu_vertical_flag` indicates the splitting direction of the current block. When `mtt_split_cu_vertical_flag` is true (or 1), the current block can be split vertically, and when `mtt_split_cu_vertical_flag` is false (or 0), the current block can be split horizontally. `mtt_split_cu_binary_flag` indicates whether the current block is split based on a binary split. When `mtt_split_cu_binary_flag` is true (or 1), the current block can be split into two sub-blocks, and when `mtt_split_cu_binary_flag` is false (or 0), the current block can be split into three sub-blocks.
[0102] Therefore, when both `mtt_split_cu_vertical_flag` and `mtt_split_cu_binary_flag` are true, the current block can be vertically split into two sub-blocks. When `mtt_split_cu_vertical_flag` is true and `mtt_split_cu_binary_flag` is false, the current block can be vertically split into three sub-blocks. When `mtt_split_cu_vertical_flag` is false and `mtt_split_cu_binary_flag` is true, the current block can be horizontally split into two sub-blocks. When both `mtt_split_cu_vertical_flag` and `mtt_split_cu_binary_flag` are false, the current block can be horizontally split into three sub-blocks.
[0103] The method described above for signaling segmentation information can be applied to a coded block, and the method can be recursively applied to multiple sub-blocks segmented from a coded block.
[0104] Reference Figure 4 S410 can decode multiple sub-blocks based on a predetermined decoding order.
[0105] The decoding order of sub-blocks within the current block can be determined based on one or more predefined scan order candidates. These one or more scan order candidates may include at least one of a horizontal Z-scan order or a vertical Z-scan order.
[0106] Suppose the current block is divided into four sub-blocks based on quad-slicing. In this case, the four sub-blocks can be decoded sequentially in a horizontal Z-scan order. Here, the horizontal Z-scan order can refer to the order in which the four sub-blocks obtained from quad-slicing are decoded: top-left sub-block, top-right sub-block, bottom-left sub-block, and bottom-right sub-block. Alternatively, the four sub-blocks can be decoded sequentially in a vertical Z-scan order. Here, the vertical Z-scan order can refer to the order in which the four sub-blocks obtained from quad-slicing are decoded: top-left sub-block, bottom-left sub-block, top-right sub-block, and bottom-right sub-block.
[0107] You can selectively use one of one or more scan order candidates available for the current block.
[0108] The number and / or type of one or more scan order candidates available for the current block can be deduced based on the type of block segmentation of the current block.
[0109] As an example, when the current block is segmented based on quadruple partitioning, the number of one or more scan order candidates available for the current block can be greater than or equal to 2. On the other hand, when the current block is not segmented based on quadruple partitioning, the number of one or more scan order candidates available for the current block can be 1. Alternatively, when the current block is segmented based on quadruple partitioning, the one or more scan order candidates available for the current block may include a horizontal Z-scan order and a vertical Z-scan order. On the other hand, when the current block is not segmented based on quadruple partitioning, the one or more scan order candidates available for the current block may include a horizontal Z-scan order, but may not include a vertical Z-scan order.
[0110] The index information of one or more scan sequence candidates can be explicitly indicated by a signal.
[0111] As an example, the current block can be divided into four sub-blocks based on quad-segmentation, and the index information of either the horizontal Z-scan order or the vertical Z-scan order can be signaled for the four sub-blocks.
[0112] Specifically, when the current block is split into four sub-blocks (i.e., when both split_cu_flag and split_qt_flag are true), qt_scan_flag can be signaled. In other words, qt_scan_flag can be index information used to select either the horizontal Z-scan order or the vertical Z-scan order when the current block is split into four sub-blocks based on a quadruple partition.
[0113] When qt_scan_flag is true (or 1), the four sub-blocks can be decoded sequentially according to the horizontal Z-scan order. On the other hand, when qt_scan_flag is false (or 0), the four sub-blocks can be decoded sequentially according to the vertical Z-scan order. Conversely, when qt_scan_flag is true, the four sub-blocks can be decoded sequentially according to the vertical Z-scan order, while when qt_scan_flag is false, the four sub-blocks can be decoded sequentially according to the horizontal Z-scan order.
[0114] Tables 1 and 2 below show the syntax and semantics corresponding to the above implementation. As shown in Table 1, when split_qt_flag indicates that the current block is split into four sub-blocks (i.e., when split_qt_flag is true), qt_scan_flag can be signaled for the current block.
[0115] [Table 1]
[0116] [Table 2]
[0117] The sub-blocks within the current block can be decoded in the decoding order according to the selected scan order candidate.
[0118] As an example, the decoding process for each sub-block may also include block segmentation processing for the corresponding sub-block. The segmentation information for the block segmentation of the corresponding sub-block can be additionally signaled, in the same way as the method described above for signaling segmentation information.
[0119] Whether the above adaptive scan order is used / available can be notified by using signals using advanced syntax.
[0120] As an example, whether an adaptive scan order is used / available can be signaled via the Sequence Parameter Set (SPS). Table 3 is the syntax table for signaling whether an adaptive scan order is used / available via the Sequence Parameter Set. In Table 3, `sps_adaptive_scan_order_flag` can indicate whether an adaptive scan order is used / available within the sequence currently using `sps_seq_parameter_set_id` among multiple sequences belonging to the bitstream. For example, when `sps_adaptive_scan_order_flag` is true (or 1), an adaptive scan order can be used or available in the sequence currently using `sps_seq_parameter_set_id` among multiple sequences belonging to the bitstream. On the other hand, when `sps_adaptive_scan_order_flag` is false (or 0), an adaptive scan order may not be used or available in the sequence currently using `sps_seq_parameter_set_id` among multiple sequences belonging to the bitstream.
[0121] [Table 3]
[0122] As an example, whether an adaptive scan order is used / available can be signaled via a frame parameter set (PPS). Table 4 is the syntax table for signaling whether an adaptive scan order is used / available via a frame parameter set. In Table 4, `pps_adaptive_scan_order_flag` can indicate whether an adaptive scan order is used / available within the frame currently using `pps_seq_parameter_set_id` among multiple frames belonging to the bitstream. For example, when `pps_adaptive_scan_order_flag` is true (or 1), an adaptive scan order can be used or available in the frame currently using `pps_seq_parameter_set_id` among multiple frames belonging to the bitstream. On the other hand, when `pps_adaptive_scan_order_flag` is false (or 0), an adaptive scan order may not be used or available in the frame currently using `pps_seq_parameter_set_id` among multiple frames belonging to the bitstream.
[0123] [Table 4]
[0124] As an example, whether the adaptive scan order is used / available can be signaled via the frame header (PH). Table 5 is the syntax table for signaling whether the adaptive scan order is used / available via the frame header. In Table 5, ph_adaptive_scan_order_flag can indicate whether the adaptive scan order is used / available in the current frame. For example, when ph_adaptive_scan_order_flag is true (or 1), the adaptive scan order can be used or is available in the current frame. On the other hand, when ph_adaptive_scan_order_flag is false (or 0), the adaptive scan order may not be used or is not available in the current frame.
[0125] When `sps_adaptive_scan_order_flag` is true, `ph_adaptive_scan_order_flag` can be notified via a signal. When `sps_adaptive_scan_order` is false, `ph_adaptive_scan_order_flag` can be assumed to be false without being notified via a signal.
[0126] [Table 5]
[0127] The advanced syntax for signaling whether an adaptive scan order is used / available is not limited to SPS, PPS, or PH. As an example, the advanced syntax can be signaled at the level of at least one of the following: tile, slice, or CTU row / column.
[0128] The index information (qt_scan_flag) of one or more scan order candidates as described above can be signaled based on at least one of the high-level syntax.
[0129] As an example, when `sps_adaptive_scan_order_flag` is true, `qt_scan_flag` can be notified via a bitstream signal; otherwise, it can be notified via a signal without a bitstream. Similarly, when `pps_adaptive_scan_order_flag` is true, `qt_scan_flag` can be notified via a bitstream signal; otherwise, it can be notified via a signal without a bitstream. Likewise, when `ph_adaptive_scan_order_flag` is true, `qt_scan_flag` can be notified via a bitstream signal; otherwise, it can be notified via a signal without a bitstream.
[0130] The qt_scan_flag can be signaled based on the split_qt_flag indicating whether the current block was split based on a quadruple partition.
[0131] As an example, when split_qt_flag is true, qt_scan_flag can be notified via a bitstream signal; otherwise, qt_scan_flag can be notified via a bitstream signal.
[0132] The qt_scan_flag can be signaled based on at least one of the split_qt_flag for the current block and the advanced syntax described above.
[0133] As an example, when both `split_qt_flag` and `sps_adaptive_scan_order_flag` are true, `qt_scan_flag` can be notified via a bitstream signal. On the other hand, when at least one of `split_qt_flag` and `sps_adaptive_scan_order_flag` is false, `qt_scan_flag` can be notified via a bitstream signal.
[0134] Alternatively, when both `split_qt_flag` and `pps_adaptive_scan_order_flag` are true, `qt_scan_flag` can be notified via a bitstream signal. On the other hand, when at least one of `split_qt_flag` and `pps_adaptive_scan_order_flag` is false, `qt_scan_flag` can be notified via a bitstream signal without a bitstream signal.
[0135] Alternatively, as shown in Table 6, when both split_qt_flag and ph_adaptive_scan_order_flag are true, qt_scan_flag can be notified via a bitstream signal. On the other hand, when at least one of split_qt_flag and ph_adaptive_scan_order_flag is false, qt_scan_flag can be notified without a bitstream signal.
[0136] [Table 6]
[0137] An additional signal can be used to indicate whether an adaptive scan order is used in the encoding unit. This signal can be expressed as `cu_adaptive_scan_order_flag`.
[0138] As an example, when `cu_adaptive_scan_order_flag` is true (or 1), an adaptive scan order can be used for the current frame. In this case, the index information of one or more scan order candidates can be signaled. On the other hand, when `cu_adaptive_scan_order_flag` is false (or 0), an adaptive scan order may not be used for the current frame. In this case, the default scan order (e.g., the horizontal Z scan order) can be used for the current block. When `cu_adaptive_scan_order_flag` is not signaled, it can be deduced to be false.
[0139] The cu_adaptive_scan_order_flag can be signaled based on at least one of the advanced syntaxes mentioned above.
[0140] As an example, when `sps_adaptive_scan_order_flag` is true, it can be notified to `cu_adaptive_scan_order_flag` via a bitstream signal; otherwise, it can be notified without a bitstream signal. Similarly, when `pps_adaptive_scan_order_flag` is true, it can be notified to `cu_adaptive_scan_order_flag` via a bitstream signal; otherwise, it can be notified without a bitstream signal. Likewise, when `ph_adaptive_scan_order_flag` is true, it can be notified to `cu_adaptive_scan_order_flag` via a bitstream signal; otherwise, it can be notified without a bitstream signal.
[0141] The cu_adaptive_scan_order_flag can be signaled based on the split_qt_flag, which indicates whether the current block was split based on a quadruple partition.
[0142] As an example, when split_qt_flag is true, cu_adaptive_scan_order_flag can be notified by signal via bitstream; otherwise, cu_adaptive_scan_order_flag can be notified by signal without bitstream.
[0143] The cu_adaptive_scan_order_flag can be signaled based on at least one of the split_qt_flag for the current block and the advanced syntax described above.
[0144] As an example, when both `split_qt_flag` and `sps_adaptive_scan_order_flag` are true, `cu_adaptive_scan_order_flag` can be notified via a bitstream signal. Conversely, when at least one of `split_qt_flag` and `sps_adaptive_scan_order_flag` is false, `cu_adaptive_scan_order_flag` can be notified without a bitstream signal. Similarly, when both `split_qt_flag` and `pps_adaptive_scan_order_flag` are true, `cu_adaptive_scan_order_flag` can be notified via a bitstream signal. Conversely, when at least one of `split_qt_flag` and `pps_adaptive_scan_order_flag` is false, `cu_adaptive_scan_order_flag` can be notified without a bitstream signal. When both `split_qt_flag` and `ph_adaptive_scan_order_flag` are true, `cu_adaptive_scan_order_flag` can be notified via a bitstream signal. Conversely, when at least one of `split_qt_flag` and `ph_adaptive_scan_order_flag` is false, `cu_adaptive_scan_order_flag` can be notified via a bitstream signal.
[0145] As described above, the decoding process for each sub-block can include block partitioning for the corresponding sub-block. In this case, at least one of `qt_scan_flag` or `cu_adaptive_scan_order_flag` can be signaled along with the partitioning information for the block partitioning at the node of the corresponding sub-block. Therefore, when the current block is partitioned into multiple first sub-blocks and the first sub-blocks are partitioned into multiple second sub-blocks, the decoding order of the first sub-blocks within the current block can be different from the decoding order of the second sub-blocks within the first sub-blocks. Here, when the partitioning depth of the current block is N, the partitioning depth of the first sub-blocks can be (N+1), and the partitioning depth of the second sub-blocks can be (N+2).
[0146] The current frame can be divided into multiple top-level block units (or coding tree units, CTUs), and an adaptive decoding order can be applied to multiple top-level block units, as will be explained below. Figures 5 to 7 Let me describe it in detail.
[0147] Figure 5 This example illustrates the decoding order based on the raster scan order of multiple CTUs within a single frame.
[0148] like Figure 5 As shown, decoding is performed in the order of the first CTU column to the last CTU column within a frame, and can also be performed in the order of the leftmost CTU to the rightmost CTU within each CTU column. It can be defined as the decoding order based on the raster scan order. In other words, it can refer to a scan order in which all CTUs included in the first CTU column are decoded, and then the CTUs included in the second CTU column are decoded sequentially.
[0149] Figure 6 and Figure 7 An example of adaptive scan order being used for multiple CTUs within a single frame is shown. The decoding order, which combines Z-scan order and raster scan order, can be used within a predefined number of CTU columns.
[0150] Figure 6 This example illustrates the case where the predefined number of CTU columns is 2. (See also...) Figure 6 The first and second CTUs in the first CTU column can be decoded sequentially, and then the first and second CTUs in the second CTU column can be decoded sequentially in Z-scan order. Subsequently, the third and fourth CTUs in the first CTU column can be decoded sequentially in raster scan order, and then the third and fourth CTUs in the second CTU column can be decoded sequentially in Z-scan order.
[0151] Figure 7 This example illustrates the case where the predefined number of CTU columns is 4. (See also...) Figure 7 Decoding can be performed in units of four CTU columns within the frame, following the Z-scan order. Then, decoding can be performed again in units of four CTU columns from the CTU columns located in the raster scan order, following the Z-scan order.
[0152] In this way, encoding efficiency can be improved by using a decoding order in which the Z-scan order and the raster scan order are combined within a predefined number of CTU columns.
[0153] The number of predefined CTU columns can be an integer that is a multiple of 2. Furthermore, the number of predefined CTU columns can be a fixed value predefined for the encoding and decoding devices. Alternatively, information specifying the number of CTU columns can be signaled separately. This information can be signaled at the level of at least one of the sequence parameter set, picture parameter set, picture header, or slice header. As an example, information signaled regarding the decoding order using a combination of Z-scan order and raster scan order can be shown in Table 7 below.
[0154] [Table 7]
[0155] According to Table 7, a signal can be used to indicate whether an adaptive scan order in CTU units is used in the current sequence using sps_seq_parameter_set_id among multiple sequences in the bitstream. The corresponding flag can be expressed as sps_adaptive_ctu_scan_order_flag.
[0156] When `sps_adpative_ctu_scan_order_flag` is true, an adaptive scan order in units of CTUs can be used in the sequence currently using `sps_seq_parameter_id` among multiple sequences within the bitstream. Conversely, when `sps_adpative_ctu_scan_order_flag` is false, an adaptive scan order in units of CTUs may not be used in the sequence currently using `sps_seq_parameter_id` among multiple sequences within the bitstream.
[0157] Only when `sps_adaptive_ctu_scan_order_flag` is true can information about the number of CTU columns using the adaptive scan order (`sps_num_of_scan_lines`) in the corresponding sequence, scanned in Z order, be explicitly signaled. The number of CTU columns using the adaptive scan order can be expressed as an integer value that is a multiple of 2. For example, when `sps_num_of_scan_lines` is N, the number of CTU columns using the adaptive scan order can be derived as (2...). N). Alternatively, when the value of sps_num_of_scan_lines is N, the number of CTU columns using adaptive scan order can be derived as 2. N .
[0158] Figure 8 An illustrative configuration of a decoding device 300 performing an image decoding method according to the present disclosure is shown.
[0159] Reference Figure 8 The decoding device 300 may include a memory 360 and at least one processor 800.
[0160] Processor 800 can divide the current block into multiple sub-blocks. The division can be performed based on one of a predefined block division type. The block division type can include at least one of a quadruple partition, a binary partition, or a triangular partition. A binary partition can include at least one of a symmetric binary partition or an asymmetric binary partition. For this purpose, processor 800 can obtain partitioning information indicating one of the block division types from the bitstream, which is referenced... Figure 4 The description is the same. Processor 800 can use tree-based block partitioning.
[0161] Processor 800 can decode multiple sub-blocks based on a predetermined decoding order. The decoding order of sub-blocks within the current block can be determined based on one or more predefined scan order candidates, which is consistent with the reference... Figure 4 The description is the same.
[0162] Processor 800 can determine whether an adaptive scan order is used / available based on the high-level grammar. As an example, the high-level grammar can be signaled at the level of at least one of SPS, PPS, or PH, which is related to the reference... Figure 4 The same applies to what is described. However, it is not limited to this and can signal the higher-level syntax at the level of at least one of the chunks, slices, and / or CTU rows / columns.
[0163] Additionally, as referenced Figure 4As described, processor 800 can obtain index information (qt_scan_flag) indicative of one or more scan order candidates from the bitstream based on at least one of the high-level syntaxes. Alternatively, processor 800 can obtain qt_scan_flag from the bitstream based on split_qt_flag, which indicates whether the current block is split based on a quadruple partition. Alternatively, processor 800 can obtain qt_scan_flag from the bitstream based on split_qt_flag for the current block and at least one of the high-level syntaxes described above.
[0164] Processor 800 can obtain from the bitstream a syntax (cu_adaptive_scan_order_flag) indicating whether an adaptive scan order is used in the encoding unit, which is consistent with the reference... Figure 4 The description is the same.
[0165] Processor 800 can divide the current frame into multiple top-level block units (or coding tree units, CTUs) and decode these top-level block units based on an adaptive decoding order. The adaptive decoding order of the multiple top-level block units is related to a reference... Figures 5 to 7 The description is the same.
[0166] Figure 9 An image encoding method performed by an encoding device 200 as an embodiment of the present disclosure is illustrated.
[0167] Reference Figure 9 This can divide the current block into multiple sub-blocks S900.
[0168] Segmentation can be performed based on one of the predefined block segmentation types. The block segmentation type can include at least one of quadruple segmentation, binary segmentation, or triangular segmentation. Here, binary segmentation can include at least one of symmetric binary segmentation or asymmetric binary segmentation. Tree-based block segmentation can be used in the image coding method according to this disclosure.
[0169] The segmentation information used to indicate one of the block segmentation types can be encoded and signaled, and the signaling method for segmentation information is similar to that used in the reference. Figure 4 The description is the same.
[0170] Reference Figure 9 S910 can encode multiple sub-blocks based on a predetermined encoding order.
[0171] The encoding order of sub-blocks within the current block can be determined based on one or more predefined scan order candidates. Here, one or more scan order candidates may include at least one of a horizontal Z-scan order or a vertical Z-scan order.
[0172] You may selectively use one or more scan order candidates available for the current block. The number and / or type of the one or more scan order candidates available for the current block may be determined based on the type of block partitioning of the current block.
[0173] Index information indicating the selection of any one of one or more scan sequence candidates can be encoded and signaled. The signaling method for index information is similar to that used in reference [reference / ... Figure 4 The description is the same.
[0174] The sub-blocks within the current block can be encoded according to the selected scan order candidate in the encoding order. As an example, the encoding process for each sub-block can also include block segmentation processing for the corresponding sub-block. The segmentation information for the block segmentation of the corresponding sub-block can be additionally encoded and signaled, which is the same as the method described above for signaling segmentation information.
[0175] The advanced syntax regarding whether the aforementioned adaptive scan order is used / available can be encoded and signaled. (Advanced syntax and reference) Figure 4 The description is the same.
[0176] The index information (qt_scan_flag) indicating one of the above-mentioned scan order candidates can be encoded and signaled based on at least one of the high-level syntaxes. Alternatively, the qt_scan_flag can be encoded and signaled based on split_qt_flag, which indicates whether the current block is split based on a quad partition. Alternatively, the qt_scan_flag can be encoded and signaled based on the split_qt_flag for the current block and at least one of the above-mentioned high-level syntaxes.
[0177] The syntax (cu_adaptive_scan_order_flag) used to determine and indicate whether an adaptive scan order is used in an encoding unit can be additionally encoded and signaled.
[0178] As an example, when it is determined that the adaptive scan order is used for the current block, `cu_adaptive_scan_order_flag` can be encoded as true (or 1). When `cu_adaptive_scan_order_flag` is true, index information indicating one or more scan order candidates can be encoded and signaled. On the other hand, when it is determined that the adaptive scan order is not used for the current block, `cu_adaptive_scan_order_flag` can be encoded as false (or 0). In this case, the encoded order for the current block can be deduced as the default scan order (e.g., the horizontal Z scan order). When `cu_adaptive_scan_order_flag` is not encoded, it can be deduced as false.
[0179] The cu_adaptive_scan_order_flag can be encoded in a bitstream based on at least one of the above advanced syntaxes.
[0180] As an example, when `sps_adaptive_scan_order_flag` is true, `cu_adaptive_scan_order_flag` can be encoded in the bitstream; otherwise, it may not be encoded in the bitstream. Similarly, when `pps_adaptive_scan_order_flag` is true, it can be encoded in the bitstream; otherwise, it may not be encoded in the bitstream. Likewise, when `ph_adaptive_scan_order_flag` is true, it can be encoded in the bitstream; otherwise, it may not be encoded in the bitstream.
[0181] The cu_adaptive_scan_order_flag can be encoded based on split_qt_flag, which indicates whether the current block is split based on a quadruple partition.
[0182] As an example, when split_qt_flag is true, cu_adaptive_scan_order_flag can be encoded in the bitstream; otherwise, cu_adaptive_scan_order_flag may not be encoded in the bitstream.
[0183] The cu_adaptive_scan_order_flag can be encoded based on at least one of the split_qt_flag for the current block and the advanced syntax described above.
[0184] As an example, when both `split_qt_flag` and `sps_adaptive_scan_order_flag` are true, `cu_adaptive_scan_order_flag` can be encoded in the bitstream. Conversely, when at least one of `split_qt_flag` and `sps_adaptive_scan_order_flag` is false, `cu_adaptive_scan_order_flag` may not be encoded in the bitstream. Similarly, when both `split_qt_flag` and `pps_adaptive_scan_order_flag` are true, `cu_adaptive_scan_order_flag` can be encoded in the bitstream. Conversely, when at least one of `split_qt_flag` and `pps_adaptive_scan_order_flag` is false, `cu_adaptive_scan_order_flag` may not be encoded in the bitstream. When both `split_qt_flag` and `ph_adaptive_scan_order_flag` are true, `cu_adaptive_scan_order_flag` may not be encoded in the bitstream. Conversely, when at least one of `split_qt_flag` and `ph_adaptive_scan_order_flag` is false, `cu_adaptive_scan_order_flag` may not be encoded in the bitstream.
[0185] As described above, the encoding process for each sub-block can include block partitioning for that sub-block. In this case, at least one of `qt_scan_flag` or `cu_adaptive_scan_order_flag` can be encoded together with the partitioning information for the block partitioning at the nodes of the corresponding sub-block. Therefore, when the current block is partitioned into multiple first sub-blocks and the first sub-blocks are partitioned into multiple second sub-blocks, the encoding order of the first sub-blocks within the current block can be different from the encoding order of the second sub-blocks within the first sub-blocks. Here, when the partitioning depth of the current block is N, the partitioning depth of the first sub-blocks can be (N+1), and the partitioning depth of the second sub-blocks can be (N+2).
[0186] The current frame can be divided into multiple top-level block units (or coding tree units, CTUs), and an adaptive coding order can be applied to multiple top-level block units. This is because, through reference... Figures 5 to 7 The adaptive decoding order discussed can be equally applied to the adaptive encoding order, so overlapping descriptions will be omitted.
[0187] Figure 10 An illustrative configuration of an encoding device 200 performing an image encoding method according to the present disclosure is shown.
[0188] Reference Figure 10 The encoding device 200 may include a memory 270 and at least one processor 1000.
[0189] Processor 1000 can divide the current block into multiple sub-blocks. The division can be performed based on one of a predefined block division type. The block division type can include at least one of quadruple division, binary division, or triangular division, and the binary division can include at least one of symmetric binary division or asymmetric binary division. Processor 1000 can encode the division information used to indicate one of the block division types. Processor 1000 can use tree-based block division.
[0190] The processor 1000 can encode multiple sub-blocks based on a predetermined encoding order. The encoding order of sub-blocks within the current block can be determined based on one or more predefined scan order candidates, which is consistent with the reference... Figure 9 The description is the same.
[0191] Processor 1000 can determine whether an adaptive scan order is used / available, and can encode the high-level grammar based on this determination. As an example, the high-level grammar can be encoded at the level of at least one of SPS, PPS, or PH, which is consistent with the reference... Figure 4 The same applies to what is described. However, it is not limited to this and can encode high-level syntax at the level of at least one of the following: chunks, slices, and / or CTU rows / columns.
[0192] Processor 1000 may encode index information (qt_scan_flag) indicating one of the aforementioned scan order candidates in the bitstream based on at least one of the high-level syntaxes. Alternatively, processor 1000 may encode qt_scan_flag in the bitstream based on split_qt_flag, which indicates whether the current block is split based on a quadruple partition. Alternatively, processor 1000 may also encode qt_scan_flag in the bitstream based on split_qt_flag for the current block and at least one of the aforementioned high-level syntaxes.
[0193] The processor 1000 can additionally encode the syntax (cu_adaptive_scan_order_flag) indicating whether an adaptive scan order is used in the coding unit within the bitstream, which is consistent with the reference. Figure 9 The description is the same.
[0194] Processor 1000 can segment the current image into multiple top-level block units (or coding tree units, CTUs) and encode these top-level block units based on an adaptive decoding order. This is achieved by referencing... Figures 5 to 7 The adaptive decoding order discussed can be equally applied to the adaptive encoding order, so overlapping descriptions will be omitted.
[0195] In the above embodiments, the method is described based on a flowchart as a series of steps or blocks. However, the corresponding embodiments are not limited to this order of steps. Some steps may occur simultaneously with other steps or in a different order, as described above. In addition, those skilled in the art will understand that the steps shown in the flowchart are not exclusive. Other steps may be included, or one or more steps in the flowchart may be deleted, without affecting the scope of the embodiments of this disclosure.
[0196] The methods described above according to embodiments of the present disclosure can be implemented in software, and the encoding and / or decoding devices according to the present disclosure can be included in an apparatus for performing image processing, such as a TV, computer, smartphone, set-top box, display device, etc.
[0197] In this disclosure, when the implementation is implemented as software, the above-described method can be implemented as a module (process, function, etc.) performing the above-described functions. The module can be stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), another chipset, logic circuitry, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. In other words, the implementations described herein can be executed by implementation on a processor, microprocessor, controller, or chip. For example, the functional units shown in the various figures can be executed by implementation on a computer, processor, microprocessor, controller, or chip. In this case, information for implementation (e.g., information about instructions) or algorithms can be stored in a digital storage medium.
[0198] Furthermore, decoding and encoding devices employing embodiments of this disclosure can be included in multimedia broadcasting transmitting and receiving devices, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video conferencing devices, real-time communication devices similar to video communication, mobile streaming devices, storage media, cameras, devices for providing video-on-demand (VOD) services, OTT (over-the-top) video devices, devices for providing internet streaming services, three-dimensional (3D) video devices, virtual reality (VR) devices, augmented reality (AR) devices, video telephony devices, transportation terminals (e.g., vehicle (including autonomous vehicle) terminals, aircraft terminals, ship terminals, etc.), and medical video devices, and can be used to process video signals or data signals. For example, OTT (over-the-top) video devices can include game consoles, Blu-ray players, internet-connected televisions, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0199] Furthermore, the processing methods applying the embodiments of this disclosure can be generated in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data with data structures according to the embodiments of this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include all types of storage devices and distributed storage devices that store computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical media storage devices. Additionally, computer-readable recording media include media implemented in the form of carrier waves (e.g., transmission via the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired / wireless communication networks.
[0200] Furthermore, the embodiments of this disclosure can be implemented by a computer program product using program code, and the program code can be executed on a computer using the embodiments of this disclosure. The program code can be stored on a computer-readable medium.
[0201] Figure 11 Examples of content streaming systems to which embodiments of the present disclosure can be applied are shown.
[0202] Reference Figure 11 A content streaming system that applies embodiments of the present disclosure may generally include an encoding server, a streaming server, a network server, a media storage device, a user device, and a multimedia input device.
[0203] An encoding server generates a bitstream by compressing content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data and then sends it to a streaming server. As another example, when multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server can be omitted.
[0204] A bitstream can be generated by an encoding method or bitstream generation method that applies the embodiments of this disclosure, and the streaming server can temporarily store the bitstream during the sending or receiving of the bitstream.
[0205] A streaming server sends multimedia data to a user device via a web server based on a user request, and the web server acts as a medium to inform the user what services are available. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this scenario, the content streaming system may include a separate control server, which controls the commands / responses between the various devices in the content streaming system.
[0206] A streaming server can receive content from media storage and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a specific time period.
[0207] Examples of user devices may include mobile phones, smartphones, laptops, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, touchscreen PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc.
[0208] In a content streaming system, each server can operate as a distributed server, and in this case, the data received from each server can be distributed and processed.
[0209] The claims set forth herein can be combined in various ways. For example, the technical features of the method claims of this disclosure can be combined and implemented as an apparatus, and the technical features of the apparatus claims of this disclosure can be combined and implemented as a method. Furthermore, the technical features of the method claims and the apparatus claims of this disclosure can be combined and implemented as an apparatus, and the technical features of the method claims and the apparatus claims of this disclosure can be combined and implemented as a method.
Claims
1. A method, the method comprising: The current block is divided into multiple sub-blocks based on one of the predefined block partitioning types; as well as The multiple sub-blocks are decoded according to a predetermined decoding order. The decoding order is determined based on one or more candidate scan orders, and The one or more candidate scan sequences include at least one of a horizontal Z-scan sequence or a vertical Z-scan sequence.
2. The method according to claim 1, wherein, The decoding order is determined based on index information indicating one of the one or more scan order candidates.
3. The method according to claim 2, wherein, The index information is signaled based on information indicating whether the current block has been divided into four sub-blocks.
4. The method according to claim 3, wherein, The method also includes obtaining a high-level syntax from the bitstream regarding whether to use an adaptive scan order.
5. The method according to claim 4, wherein, The advanced syntax is signaled in at least one of the sequence parameter set, the screen parameter set, or the screen header.
6. The method according to claim 5, wherein, The index information is signaled based on the high-level syntax.
7. The method according to claim 1, wherein, The current block, including the current block, is divided into multiple top-level block units, and The multiple top-level block units are decoded based on a decoding order that combines the raster scan order and the Z scan order.
8. A method, the method comprising: The current block is divided into multiple sub-blocks based on one of the predefined block partitioning types; as well as The multiple sub-blocks are encoded according to a predetermined encoding order. The encoding order is determined based on one or more scan order candidates, and The one or more candidate scan sequences include at least one of a horizontal Z-scan sequence or a vertical Z-scan sequence.
9. A computer-readable storage medium storing a bit stream generated by the method according to claim 8.
10. A method, the method comprising: Obtain a bitstream of image information, wherein the bitstream is generated by dividing a current block into multiple sub-blocks based on one of a predefined block segmentation type and encoding the multiple sub-blocks according to a predetermined encoding order; and Send data including the bit stream. The encoding order is determined based on one or more scan order candidates, and The one or more candidate scan sequences include at least one of a horizontal Z-scan sequence or a vertical Z-scan sequence.