Encoding / decoding device and device for transmitting data
By deriving residuals and predicted samples using BDPCM information and combining them with deblocking filtering to optimize video coding, the problem of high cost in high-resolution image/video coding is solved, and coding efficiency and deblocking filtering effect are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2020-07-09
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies suffer from high transmission and storage costs during the encoding of high-resolution and high-quality images/videos, and lack effective methods for improving the efficiency of deblocking filtering and BDPCM encoding.
The residual samples and predicted samples of the current block are derived using BDPCM information to generate a reconstructed image. Deblocking filtering is performed when appropriate, and the video coding process is optimized by deriving boundary strength and orientation information.
It improves image/video compression efficiency and enhances the deblocking filtering efficiency based on BDPCM, especially the coding efficiency of chroma images.
Smart Images

Figure CN117041573B_ABST
Abstract
Description
[0001] This application is a divisional application of the original invention patent application No. 202080055790.7 (International Application No.: PCT / KR2020 / 008984, Application Date: July 9, 2020, Invention Title: Method and Apparatus for Encoding Images Based on Deblocking Filtering). Technical Field
[0002] This disclosure relates to video coding technology, and more specifically, to a video coding method and apparatus based on deblocking filtering in a video coding system. Background Technology
[0003] Today, the demand for high-resolution and high-quality images / videos, such as 4K, 8K, or even higher Ultra High Definition (UHD) images / videos, is constantly growing across various fields. As image / video data becomes higher resolution and higher quality, the amount of information or bits transmitted increases compared to traditional image data. Therefore, transmission and storage costs increase when using media such as traditional wired / wireless broadband lines to transmit image data or when using existing storage media to store image / video data.
[0004] In addition, there is increasing interest and demand for immersive media such as virtual reality (VR) and artificial reality (AR) content or holograms, and broadcasting of images / videos with image characteristics that differ from real images such as game images is on the rise.
[0005] Therefore, there is a need for efficient image / video compression techniques to effectively compress, transmit, store, and reproduce information with high resolution and high quality images / videos that have the various characteristics described above. Summary of the Invention
[0006] Technical issues
[0007] One aspect of this disclosure is to provide a method and apparatus for increasing image coding efficiency.
[0008] This disclosure also provides methods and apparatus for increasing the efficiency of transform index coding in video coding based on deblocking filtering.
[0009] This disclosure also provides methods and apparatus for deblocking filtering of BDPCM-encoded video.
[0010] This disclosure also provides a video coding method and apparatus for chroma components based on BDPCM coding.
[0011] Technical solution
[0012] In one aspect, a video decoding method performed by a decoding device is provided. The method may include: receiving a bitstream including BDPCM (Block Differential Pulse Code Modulation or Block-Based Delta Pulse Code Modulation) information; deriving residual samples of the current block based on the BDPCM information; deriving predicted samples of the current block based on the BDPCM information; generating a reconstructed image based on the residual samples and the predicted samples; and performing deblocking filtering on the reconstructed image, wherein deblocking filtering may not be performed when BDPCM is applied to the current block.
[0013] BDPCM information may include flags indicating whether BDPCM is applied to the current block, and when the flag is 1, the boundary strength (bS) used for deblocking filtering can be derived to be zero.
[0014] The current block can include a luminance-coded block or a chrominance-coded block.
[0015] The tree type of the current block can be a single tree type, the current block can be a chroma-coded block, and when the flag information is 1, the boundary strength can be derived as 1.
[0016] The steps for deriving residual samples may include: deriving the quantization transform coefficients of the current block based on BDPCM; and deriving the transform coefficients by performing dequantization of the quantization transform coefficients.
[0017] The quantization transformation coefficients can be derived based on the directional information for the direction in which BDPCM is performed.
[0018] Intra-prediction samples for the current block can be derived based on the direction of BDPCM execution.
[0019] In another aspect, a video coding method performed by an encoding device is provided. This method may include: deriving a prediction sample for the current block based on BDPCM (Block Differential Pulse Code Modulation or Block-based Decremental Pulse Code Modulation); deriving a residual sample for the current block based on the prediction sample; generating a reconstructed image based on the residual sample and the prediction sample; performing deblocking filtering on the reconstructed image; deriving quantized residual information for the current block based on BDPCM; and encoding the BDPCM information used for BDPCM and the quantized residual information, wherein deblocking filtering may not be performed when BDPCM is applied to the current block.
[0020] According to another embodiment of the present disclosure, a digital storage medium may be provided that stores image data including encoded image information and bitstream generated according to an image encoding method performed by an encoding device.
[0021] According to another embodiment of the present disclosure, a digital storage medium can be provided that stores image data including encoded image information and bitstreams to enable a decoding device to perform an image decoding method.
[0022] Technical effect
[0023] According to this disclosure, the overall image / video compression efficiency can be increased.
[0024] According to this disclosure, the efficiency of deblocking filtering in BDPCM-based video coding can be increased.
[0025] According to this disclosure, the efficiency of deblocking filtering for BDPCM-based chroma images can be increased.
[0026] The effects achievable through the specific examples of this disclosure are not limited to those listed above. For example, various technical effects may exist that can be understood or derived from this disclosure by one of ordinary skill in the art. Therefore, the specific effects of this disclosure are not limited to those expressly described herein, but may include various effects that can be understood or derived from the technical features of this disclosure. Attached Figure Description
[0027] Figure 1 Examples of video / image coding systems to which this disclosure can be applied are illustrated schematically.
[0028] Figure 2 This is a diagram that schematically illustrates the configuration of a video / image encoding device to which this disclosure can be applied.
[0029] Figure 3 This is a diagram that schematically illustrates the configuration of a video / image decoding device to which this disclosure can be applied.
[0030] Figure 4 It is a control flowchart used to describe the deblocking filtering process according to the implementation method.
[0031] Figure 5 This is a diagram illustrating a sample located at the boundary of a block.
[0032] Figure 6 This is a diagram illustrating a method for determining bS according to an embodiment of the present disclosure.
[0033] Figure 7 This is a control flowchart used to describe a video decoding method according to embodiments of the present disclosure.
[0034] Figure 8 This is a control flowchart used to describe a video encoding method according to embodiments of the present disclosure.
[0035] Figure 9 The structure of a content streaming system applying this disclosure is illustrated. Detailed Implementation
[0036] While this disclosure may be readily modified and includes various embodiments, specific embodiments thereof have been illustrated by way of example in the accompanying drawings and will now be described in detail. However, this is not intended to limit this disclosure to the specific embodiments disclosed herein. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the technical concept of this disclosure. The singular form may include the plural form unless the context clearly indicates otherwise. Terms such as “comprising” and “having” are intended to indicate the presence of the features, numbers, steps, operations, elements, components, or combinations thereof used in the following description, and should therefore not be construed as pre-excluding the possibility of the presence or addition of one or more different features, numbers, steps, operations, elements, components, or combinations thereof.
[0037] Furthermore, for ease of description of their different features and functions, the components in the accompanying drawings described herein are illustrated independently; however, this does not imply that each component is implemented by a separate piece of hardware or software. For example, any two or more of these components may be combined to form a single component, and any single component may be divided into multiple components. Embodiments in which components are combined and / or divided will fall within the scope of this disclosure, provided they do not depart from the spirit of this disclosure.
[0038] In the following description, preferred embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Furthermore, in the drawings, the same reference numerals are used for the same components, and repeated descriptions of the same components will be omitted.
[0039] This document relates to video / image coding. For example, the methods / examples disclosed in this document may relate to the VVC (Video Coding Universal) standard (ITU-T Rec.H.266), next-generation video / image coding standards after VVC, or other video coding-related standards (e.g., HEVC (High Efficiency Video Coding) standard (ITU-T Rec.H.265), EVC (Essential Video Coding) standard, AVS2 standard, etc.).
[0040] This document provides various implementations related to video / image encoding, and these implementations can be combined and performed together unless otherwise specified.
[0041] In this document, video can refer to a collection of images over a period of time. Generally, an image is a unit representing a specific time period, while a slice / patch is a unit that constitutes a part of an image. A slice / patch can include one or more coding tree units (CTUs). An image can consist of one or more slices / patches. An image can consist of one or more patch groups. A patch group can include one or more patches.
[0042] A pixel or primitive (pel) can refer to the smallest unit that makes up a picture (or image). Alternatively, "sample" can be used as the term corresponding to a pixel. A sample can typically represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chrominance component. Alternatively, a sample can refer to a pixel value in the spatial domain, or, when the pixel value is transformed to the frequency domain, it can refer to the transform coefficients in the frequency domain.
[0043] A unit can represent the basic unit of image processing. A unit may include a specific region and at least one of the information associated with that region. A unit may include a luminance block and two chrominance (e.g., cb, cr) blocks. Depending on the context, units and terms such as blocks and regions may be used interchangeably. Typically, an M×N block may include a set (or array) of samples or transform coefficients consisting of M columns and N rows.
[0044] In this document, the terms " / " and "," should be interpreted as indicating "and / or". For example, the expression "A / B" can mean "A and / or B". Additionally, "A, B" can mean "A and / or B". Furthermore, "A / B / C" can mean "at least one of A, B, and / or C". Additionally, "A / B / C" can mean "at least one of A, B, and / or C".
[0045] Additionally, in this document, the term "or" should be interpreted as indicating "and / or". For example, the expression "A or B" could include 1) only A, 2) only B, and / or 3) both A and B. In other words, the term "or" in this document should be interpreted as indicating "additionally or alternatively".
[0046] In this disclosure, "at least one of A and B" can mean "only A", "only B" or "both A and B". Furthermore, in this disclosure, the expression "at least one of A or B" or "at least one of A and / or B" can be interpreted as "at least one of A and B".
[0047] Furthermore, in this disclosure, "at least one of A, B, and C" may mean "A only", "B only", "C only" or "any combination of A, B, and C". Additionally, "at least one of A, B, or C" or "at least one of A, B, and / or C" may mean "at least one of A, B, and C".
[0048] Additionally, the parentheses used in this disclosure can indicate "for example". Specifically, when indicated as "prediction (intra-frame prediction)", it can mean that "intra-frame prediction" is proposed as an example of "prediction". In other words, "prediction" in this disclosure is not limited to "intra-frame prediction", and "intra-frame prediction" is proposed as an example of "prediction". Furthermore, when indicated as "prediction (i.e., intra-frame prediction)", this can also mean that "intra-frame prediction" is proposed as an example of "prediction".
[0049] The technical features described individually in one of the accompanying drawings of this disclosure may be implemented individually or simultaneously.
[0050] Figure 1 Examples of video / image coding systems to which this disclosure can be applied are illustrated schematically.
[0051] Reference Figure 1 A video / image encoding system may include a first device (source device) and a second device (receiving device). The source device may transmit encoded video / image information or data to the receiving device in the form of a file or stream via a digital storage medium or network.
[0052] The source device may include a video source, an encoding device, and a transmitter. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display, and the display may be configured as a separate device or an external component.
[0053] Video sources can be obtained through processes that capture, synthesize, or generate video / images. Video sources may include video / image capture devices and / or video / image generation devices. Video / image capture devices may include, for example, one or more cameras, video / image archives including previously captured video / images, etc. Video / image generation devices may include, for example, computers, tablets, and smartphones, and can generate video / images (electronically). For example, virtual video / images can be generated by computers, etc. In this case, the video / image capture process can be replaced by a process that generates related data.
[0054] Encoding devices can encode input video / images. They can perform a series of processes such as prediction, transformation, and quantization for compression and coding efficiency. The encoded data (encoded video / image information) can be output as a bitstream.
[0055] A transmitter can send encoded video / image information or data, output in bitstream form, to a receiver in a receiving device via a digital storage medium or network, either as a file or a stream. Digital storage media can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmitter can include elements for generating media files according to a predetermined file format and may include elements for transmission via a broadcast / communication network. The receiver can receive / extract the bitstream and send the received / extracted bitstream to a decoding device.
[0056] Decoding devices can decode video / images by performing a series of processes such as dequantization, inverse transform, and prediction, which correspond to the operations of encoding devices.
[0057] The renderer can render decoded video / images. The rendered video / images can then be displayed on a monitor.
[0058] Figure 2 This diagram schematically illustrates the configuration of a video / image encoding apparatus to which this disclosure may be applied. In the following, the term "video encoding apparatus" may include an image encoding apparatus.
[0059] Reference Figure 2 The encoding device 200 may include an image segmenter 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 250, a filter 260, and a memory 270. The predictor 220 may include an inter-frame predictor 221 and an intra-frame predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 250 may be referred to as a reconstructor or a reconstruction block generator. According to embodiments, the image segmenter 210, predictor 220, residual processor 230, entropy encoder 240, adder 250, and filter 260 described above may be constituted by one or more hardware components (e.g., an encoder chipset or processor). Furthermore, the memory 270 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may further include the memory 270 as an internal / external component.
[0060] Image partitioner 210 can divide an input image (or picture or frame) input to encoding device 200 into one or more processing units. As an example, a processing unit may be referred to as a coding unit (CU). In this case, starting from a coding tree unit (CTU) or a maximum coding unit (LCU), the coding units can be recursively partitioned according to a quadtree-binary-tritree (QTBTTT) structure. For example, based on a quadtree structure, a binary tree structure, and / or a ternary tree structure, a coding unit can be partitioned into multiple coding units of varying depths. In this case, for example, a quadtree structure can be applied first, and a binary tree structure and / or a ternary tree structure can be applied later. Alternatively, a binary tree structure can be applied first. The encoding process according to this disclosure can be performed based on the final coding units without further partitioning. In this case, the maximum coding unit can be directly used as the final coding unit based on the encoding efficiency according to the image characteristics. Alternatively, the coding units can be recursively partitioned into deeper coding units as needed, thereby allowing the optimally sized coding unit to be used as the final coding unit. Here, the encoding process may include processes such as prediction, transformation, and reconstruction, which will be described later. As another example, the processing unit may further include a prediction unit (PU) or a transformation unit (TU). In this case, the prediction unit and the transformation unit may be separate from or distinct from the final encoding unit described above. The prediction unit may be a unit for predicting samples, and the transformation unit may be a unit for deriving the transform coefficients and / or a unit for deriving the residual signal from the transform coefficients.
[0061] Depending on the context, units and terms such as blocks and regions can be used to represent each other. Typically, an M×N block can represent a set of samples or transform coefficients consisting of M columns and N rows. Samples can typically represent pixels or pixel values, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. Samples can be used as a term corresponding to pixels or primitives (pellets) in a picture (or image).
[0062] Subtractor 231 subtracts the predicted signal (predicted block, predicted sample array) output from predictor 220 from the input image signal (original block, original sample array) to generate a residual signal (residual block, residual sample array), and the generated residual signal is sent to converter 232. Predictor 220 can perform prediction on the processing target block (hereinafter referred to as "current block") and can generate a prediction block that includes the prediction samples of the current block. Predictor 220 can determine whether to apply intra-frame prediction or inter-frame prediction based on the current block or CU. As discussed later in the description of each prediction mode, the predictor can generate various prediction-related information such as prediction mode information and send the generated information to entropy encoder 240. The prediction information can be encoded in entropy encoder 240 and output as a bitstream.
[0063] Intra-predictor 222 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the reference samples can be located near or separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. Non-directional modes can include, for example, DC mode and planar mode. Depending on the level of detail in the prediction direction, the directional modes can include, for example, 33 or 65 directional prediction modes. However, this is just an example, and more or fewer directional prediction modes can be used depending on the settings. Intra-predictor 222 can determine the prediction mode to be applied to the current block by using the prediction modes applied to neighboring blocks.
[0064] Inter-frame predictor 221 can derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on blocks, sub-blocks, or samples, according to the correlation between motion information between neighboring blocks and the current block. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. The reference image including the reference block and the reference image including the temporally neighboring block may be the same as or different from each other. The temporally neighboring block may be referred to as a juxtaposed reference block, a juxtaposed CU (colCU), etc., and the reference image including the temporally neighboring block may be referred to as a juxtaposed image (colPic). For example, inter-frame predictor 221 can configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive the motion vector and / or reference image index of the current block. Inter-frame prediction can be performed based on various prediction modes. For example, in jump mode and merge mode, the inter-frame predictor 221 can use motion information of neighboring blocks as motion information of the current block. In jump mode, unlike merge mode, residual signals cannot be sent. In motion information prediction (motion vector prediction, MVP) mode, motion vectors of neighboring blocks can be used as motion vector predictors, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0065] Predictor 220 can generate prediction signals based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform prediction on a block based on an intra-block copy (IBC) prediction mode or a palette mode. The IBC prediction mode or palette mode can be used for content image / video encoding such as games, etc. Although IBC essentially performs prediction within the current block, its execution is similar to inter-frame prediction in that it derives a reference block within the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0066] The predicted signals generated by the inter-frame predictor 221 and / or the intra-frame predictor 222 can be used to generate the reconstructed signal or the residual signal. The transformer 232 can generate transform coefficients by applying transform techniques to the residual signal. For example, the transform techniques can include at least one of Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), Karhunen-Loève Transform (KLT), Graph-Based Transform (GBT), or Conditional Nonlinear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transform obtained based on the predicted signal generated using all previously reconstructed pixels. Furthermore, the transform processing can be applied to square pixel blocks of the same size, or to blocks of variable size that are not square.
[0067] Quantizer 233 quantizes the transform coefficients and sends them to entropy encoder 240, which encodes the quantized signal (information about the quantized transform coefficients) and outputs the encoded signal in a bitstream. The information about the quantized transform coefficients can be referred to as residual information. Quantizer 233 can rearrange the block-type quantized transform coefficients into a one-dimensional vector based on the coefficient scan order and generate information about the quantized transform coefficients based on this one-dimensional vector form. Entropy encoder 240 can perform various encoding methods such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). Entropy encoder 240 can encode information required for video / image reconstruction, other than the quantized transform coefficients (e.g., values of syntax elements), either together or separately. The encoded information (e.g., encoded video / image information) can be transmitted or stored in bitstream form on a unit-by-unit basis in the Network Abstraction Layer (NAL). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), and Video Parameter Set (VPS). Additionally, the video / image information may include general constraint information. In this disclosure, information and / or syntax elements sent from the encoding device to / signaled to the decoding device may be included in the video / image information. The video / image information can be encoded using the encoding process described above and included in the bitstream. The bitstream can be transmitted over a network or stored in a digital storage medium. Here, the network may include broadcast networks, communication networks, and / or the like, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmitter (not shown) that sends the signal output from the entropy encoder 240 or a memory (not shown) that stores it may be configured as an internal / external element of the encoding device 200, or the transmitter may be included in the entropy encoder 240.
[0068] The quantized transform coefficients output from quantizer 233 can be used to generate a prediction signal. For example, by applying dequantization and inverse transform using vectorized transform coefficients via dequantizer 234 and inverse transformer 235, the residual signal (residual block or residual sample) can be reconstructed. Adder 155 adds the reconstructed residual signal to the prediction signal output from inter-frame predictor 221 or intra-frame predictor 222, thereby generating a reconstructed signal (reconstructed image, reconstructed block, reconstructed sample array). When there is no residual for the processing target block, as in the case of applying a jump mode, the prediction block can be used as the reconstructed block. Adder 250 can be referred to as a reconstructor or reconstructed block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the target image, and, as described later, for inter-frame prediction of the next image by filtering.
[0069] In addition, luminance mapping with chroma scaling (LMCS) can be applied in image encoding and / or reconstruction processing.
[0070] Filter 260 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, filter 260 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be stored in memory 270, specifically in the DPB of memory 270. Various filtering methods can include, for example, deblocking filtering, sample adaptive offset, adaptive ring filter, bilateral filter, etc. As discussed later in the description of each filtering method, filter 260 can generate various filtering-related information and send the generated information to entropy encoder 240. The filtering information can be encoded in entropy encoder 240 and output as a bitstream.
[0071] The modified reconstructed image sent to memory 270 can be used as a reference image in inter-frame predictor 221. Accordingly, the encoding device can avoid prediction mismatch between the encoding device 100 and the decoding device when applying inter-frame prediction, and can also improve encoding efficiency.
[0072] The memory 270DPB can store modified reconstructed images for use as reference images in the inter-frame predictor 221. The memory 270 can store motion information of blocks in the current image from which motion information has been derived (or encoded) and / or motion information of blocks in reconstructed images. The stored motion information can be sent to the inter-frame predictor 221 to be used as motion information for neighboring blocks or temporally neighboring blocks. The memory 270 can store reconstructed samples of reconstructed blocks in the current image and send them to the intra-frame predictor 222.
[0073] Figure 3This is a diagram that schematically illustrates the configuration of a video / image decoding device to which this disclosure can be applied.
[0074] Reference Figure 3 The video decoding device 300 may include an entropy decoder 310, a residual processor 320, a predictor 330, an adder 340, a filter 350, and a memory 360. The predictor 330 may include an inter-frame predictor 331 and an intra-frame predictor 332. The residual processor 320 may include a dequantizer 321 and an inverse transformer 321. According to embodiments, the entropy decoder 310, residual processor 320, predictor 330, adder 340, and filter 350 described above may be constituted by one or more hardware components (e.g., a decoder chipset or processor). Additionally, the memory 360 may include a decoded picture buffer (DPB) and may be constituted by a digital storage medium. The hardware components may also include the memory 360 as an internal / external component.
[0075] When the input includes a bitstream containing video / image information, the decoding device 300 can interact with data already prepared therein. Figure 2 The processing of video / image information in the encoding device correspondingly reconstructs the image. For example, the decoding device 300 can deduce units / blocks based on information related to block segmentation obtained from the bitstream. The decoding device 300 can perform decoding by using processing units applied in the encoding device. Therefore, the decoding processing unit can be, for example, an encoding unit, which can be segmented along a quadtree structure, binary tree structure, and / or ternary tree structure using encoding tree units or maximum encoding units. One or more transform units can be derived using encoding units. And, the reconstructed image signal decoded and output by the decoding device 300 can be reproduced by a reproducer.
[0076] Decoding device 300 can receive data from... in the form of a bitstream. Figure 2The signal output by the encoding device can be decoded by the entropy decoder 310. For example, the entropy decoder 310 can parse the bitstream to derive information (e.g., video / image information) required for image reconstruction (or picture reconstruction). The video / image information may also include information about various parameter sets such as Adaptive Parameter Set (APS), Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), etc. In addition, the video / image information may also include general constraint information. The decoding device can further decode the picture based on the information about the parameter sets and / or general constraint information. In this disclosure, the signaling / receiving information and / or syntax elements, which will be described subsequently, can be decoded and obtained from the bitstream through the decoding process. For example, the entropy decoder 310 can decode the information in the bitstream based on encoding methods such as Exponential Golomb coding, CAVLC, CABAC, etc., and can output the values of the syntax elements required for image reconstruction and the quantized values of the transform coefficients of the residuals. More specifically, the CABAC entropy decoding method can receive bins corresponding to each syntax element in the bitstream, determine a context model using information about the target syntax element and the decoding information of neighboring and target blocks, or information about symbols / bins decoded in previous steps, predict the bin generation probability based on the determined context model, and perform arithmetic decoding on the bins to generate symbols corresponding to each syntax element value. Here, the CABAC entropy decoding method can update the context model after determining it using information about symbols / bins decoded for the context model of the next symbol / bin. Prediction information from the information decoded in the entropy decoder 310 can be provided to the predictors (inter-frame predictor 332 and intra-frame predictor 331), and the residual values (i.e., quantization transform coefficients) and associated parameter information that have undergone entropy decoding in the entropy decoder 310 can be input to the residual processor 320. The residual processor 320 can derive residual signals (residual blocks, residual samples, residual sample arrays). Additionally, filtering information from the information decoded in the entropy decoder 310 can be provided to the filter 350. Furthermore, a receiver (not shown) that receives the signal output from the encoding device can also configure the decoding device 300 as an internal / external component, and the receiver can be a component of the entropy decoder 310. Additionally, the decoding device according to this disclosure can be referred to as a video / image / picture encoding device, and the decoding device can be divided into an information decoder (video / image / picture information decoder) and a sample decoder (video / image / picture sample decoder). The information decoder may include the entropy decoder 310, and the sample decoder may include at least one of a dequantizer 321, an inverse transformer 322, an adder 340, a filter 350, a memory 360, an inter-frame predictor 332, and an intra-frame predictor 331.
[0077] The dequantizer 321 can output transform coefficients by dequantizing the quantized transform coefficients. The dequantizer 321 can rearrange the quantized transform coefficients into two-dimensional blocks. In this case, the rearrangement can be performed based on the order of coefficient scans already performed in the encoding device. The dequantizer 321 can perform dequantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain the transform coefficients.
[0078] The dequantizer 322 obtains the residual signal (residual block, residual sample array) by performing an inverse transform on the transform coefficients.
[0079] The predictor can perform predictions on the current block and generate a prediction block that includes prediction samples for the current block. The predictor can determine whether to apply intra-frame prediction or inter-frame prediction to the current block based on information about the prediction output from the entropy decoder 310, and specifically, can determine the intra-frame / inter-frame prediction mode.
[0080] The predictor can generate a predicted signal based on various prediction methods. For example, the predictor can apply intra-frame prediction or inter-frame prediction to the prediction of a block, and can also apply intra-frame prediction and inter-frame prediction simultaneously. This can be referred to as combined intra-frame and inter-frame prediction (CIIP). Additionally, the predictor can perform intra-block copying (IBC) for the prediction of a block. Intra-block copying can be used for content image / video encoding such as in games with screen content coding (SCC). Although IBC essentially performs prediction within the current block, its execution is similar to inter-frame prediction in that it derives a reference block within the current block. That is, IBC can use at least one of the inter-frame prediction techniques described in this disclosure.
[0081] The intra-predictor 331 can predict the current block by referencing samples in the current image. Depending on the prediction mode, the reference samples can be located near or separate from the current block. In intra-prediction, the prediction mode can include multiple non-directional modes and multiple directional modes. The intra-predictor 331 can determine the prediction mode applied to the current block by using the prediction modes applied to neighboring blocks.
[0082] Inter-frame predictor 332 can deduce the predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference image. In this case, to reduce the amount of motion information transmitted in inter-frame prediction mode, motion information can be predicted based on the correlation between motion information of neighboring blocks and the current block, on a block, sub-block, or sample basis. Motion information may include motion vectors and reference image indices. Motion information may also include inter-frame prediction direction (L0 prediction, L1 prediction, Bi prediction, etc.) information. In the case of inter-frame prediction, neighboring blocks may include spatially neighboring blocks existing in the current image and temporally neighboring blocks existing in the reference image. For example, inter-frame predictor 332 can configure a motion information candidate list based on neighboring blocks and deduce the motion vector and / or reference image index of the current block based on received candidate selection information. Inter-frame prediction can be performed based on various prediction modes, and the information about the prediction may include information indicating the mode of inter-frame prediction for the current block.
[0083] Adder 340 can generate a reconstruction signal (reconstructed image, reconstruction block, reconstruction sample array) by adding the obtained residual signal to the prediction signal (prediction block, prediction sample array) output from predictor 330. When there is no residual for processing the target block, as in the case of applying the jump mode, the prediction block can be used as the reconstruction block.
[0084] Adder 340 can be referred to as a reconstructor or reconstruction block generator. The generated reconstructed signal can be used for intra-frame prediction of the next processing target block in the current block, and as described later, it can be output by filtering or used for inter-frame prediction of the next image.
[0085] In addition, luminance mapping with chroma scaling (LMCS) can be applied in image decoding processing.
[0086] Filter 350 can improve subjective / objective video quality by applying filtering to the reconstructed signal. For example, filter 350 can generate a modified reconstructed image by applying various filtering methods to the reconstructed image, and the modified reconstructed image can be sent to memory 360, specifically to the DPB of memory 360. Various filtering methods can include, for example, deblocking filtering, adaptive sample shifting, adaptive ring filtering, bilateral filtering, etc.
[0087] The (modified) reconstructed image stored in the DPB of memory 360 can be used as a reference image in inter-frame predictor 332. Memory 360 can store motion information of blocks in the current image from which motion information has been derived (or decoded) and / or motion information of blocks in a reconstructed image. The stored motion information can be sent to inter-frame predictor 260 to be used as motion information of neighboring blocks or temporally neighboring blocks. Memory 360 can store reconstructed samples of reconstructed blocks in the current image and send them to intra-frame predictor 331.
[0088] The examples described in this specification in the predictor 330, dequantizer 321, inverse transformer 322 and filter 350 of the decoding device 300 can be similarly or correspondingly applied to the predictor 220, dequantizer 234, inverse transformer 235 and filter 260 of the encoding device 200, respectively.
[0089] As described above, prediction is performed to improve compression efficiency during video encoding. Accordingly, a prediction block can be generated that includes prediction samples for the current block, which is the target block for encoding. Here, the prediction block includes prediction samples in the spatial domain (or pixel domain). The prediction block can be derived identically in both the encoding and decoding devices, and the encoding device can improve image encoding efficiency by signaling to the decoding device information about the residual between the original block and the prediction block (residual information), not the original sample values of the original block itself. The decoding device can derive a residual block including residual samples based on the residual information, generate a reconstructed block including reconstructed samples by adding the residual block to the prediction block, and generate a reconstructed image including the reconstructed block.
[0090] Residual information can be generated through transformation and quantization processes. For example, an encoding device can derive a residual block between the original block and the prediction block, derive transform coefficients by performing a transform process on the residual samples (residual sample array) included in the residual block, and derive quantized transform coefficients by performing a quantization process on the transform coefficients. This allows it to signal the associated residual information to the decoding device (via a bitstream). Here, the residual information can include the value information, position information, transform technique, transform kernel, quantization parameters, etc., of the quantized transform coefficients. The decoding device can perform quantization / dequantization processes based on the residual information and derive residual samples (or residual sample blocks). The decoding device can generate a reconstructed block based on the prediction block and the residual block. The encoding device can derive the residual block by performing dequantization / inverse transform on the quantized transform coefficients to serve as a reference for inter-frame prediction of the next image, and can generate a reconstructed image based on this.
[0091] Furthermore, depending on the implementation, block differential pulse code modulation or block-based delta pulse code modulation (BDPCM) techniques can be used. BDPCM can also be referred to as quantization-based residual block delta pulse code modulation (RDPCM).
[0092] In the case of predicting blocks by applying BDPCM, rows or columns of the block are predicted line-by-line using reconstructed samples. In this case, the reference pixels used can be unfiltered samples. The BDPCM direction can indicate whether prediction is made in the vertical or horizontal direction. Prediction errors can be quantized in the spatial domain, and pixels can be reconstructed by adding the dequantized prediction error to the prediction. As an alternative to BDPCM, quantized residual domain BDPCM can be proposed, and the prediction direction or its signaling can be the same as that applied to the spatial domain BDPCM. In other words, as in delta pulse code modulation (DPCM), quantization coefficients are stacked through quantized residual domain BDPCM, and the residuals can then be reconstructed by dequantization. Therefore, quantized residual domain BDPCM can be used in the sense that DPCM is applied to the end of residual coding. In the following text, the quantized residual domain used refers to the residual derived from the prediction without transformation, and the quantized residual domain means the domain of the residual samples used for quantization.
[0093] For a block of size M (rows) × N (columns), make r (i,j) (0≤i≤M-1, 0≤j≤N-1) is the prediction residual after performing intra-frame prediction horizontally (copying the left adjacent pixel value line by line across the prediction block) or vertically (copying the upper adjacent line to each line in the prediction block) using unfiltered samples from the upper block boundary samples or the left block boundary samples. Furthermore, Q(r) (i,j (0≤i≤M-1, 0≤j≤N-1) represents the residual r (i,j) The quantized version. Here, residual means the difference between the original block value and the predicted block value.
[0094] Subsequently, when BDPCM is applied to quantized residual samples, the derivation of elements... Modified M×N array
[0095] When signaling to the vertical BDPCM As shown in the following formula.
[0096] [Formula 1]
[0097]
[0098] For horizontal prediction, when similar rules are applied, the residual quantization sample is as shown in the following equation.
[0099] [Equation 2]
[0100]
[0101] The residual quantized sample is sent to the decoding device.
[0102] In the decoding device, the above calculation is reversed to produce Q(r). (i,j) (0≤i≤M-1, 0≤j≤N-1).
[0103] For vertical prediction, the following formula can be applied.
[0104] [Formula 3]
[0105]
[0106] In addition, for horizontal prediction, the following formula can be applied.
[0107] [Formula 4]
[0108]
[0109] The inversely quantized residual Q -1 (Q(r i,j Added to the intra-block prediction value to produce reconstructed sample values.
[0110] The main benefit of this approach is that inverse BDPCM can be performed simply by adding predictors during coefficient resolution or immediately after resolution.
[0111] As described above, BDPCM can be applied to the quantized residual domain, and the quantized residual domain can include the quantized residuals (or quantized residual coefficients). In this case, the transform can be skipped from the residuals. That is, for residual samples, the transform can be skipped, but quantization can still be performed. Alternatively, the quantized residual domain can also include the quantized transform coefficients. A flag indicating whether BDPCM is applicable can be signaled at the sequence hierarchy (SPS), and this flag can be signaled only when the transform skip mode is available in the SPS.
[0112] When applying BDPCM, intra-block prediction in the residual domain for quantization can be performed by copying samples in a prediction direction similar to the intra-prediction direction (e.g., vertical or horizontal prediction). The residual is quantized, and the incremental value between the quantized residual and the predictor for the vertical or horizontal direction (i.e., the quantized residual in the horizontal or vertical direction), i.e., the difference, is... Encoded.
[0113] When encoding CUs in intra-frame prediction, flag information indicating whether BDPCM is applicable can be sent at the CU level. That is, the flag information indicates whether regular intra-frame coding or BDPCM is applied. When BDPCM is applied, a BDPCM prediction direction flag indicating whether the prediction direction is horizontal or vertical can be sent. The block is then predicted using unfiltered reference samples through a regular horizontal or vertical intra-frame prediction process. Residuals are quantized, and the difference between each quantized residual and the predictor (e.g., quantized residuals already quantized at adjacent positions in the horizontal or vertical direction along the BDPCM prediction direction) is encoded.
[0114] When BDPCM is applicable, flag information can be sent at the CU level when the CU size is equal to or the same as the MaxTsSize (maximum transform skip size) of the luma sample, and when the CU is encoded in intra-frame prediction. Here, MaxTsSize means the maximum block size allowed for transform skip mode.
[0115] The grammatical elements described above and their semantics are represented in the following table.
[0116] [Table 1]
[0117]
[0118] [Table 2]
[0119]
[0120] The syntax element “intra_bdpcm_flag” in Table 1 indicates whether BDPCM is applied to the current luma coding block. When the value of “intra_bdpcm_flag” is equal to 1, the transformation used for the coding block can be skipped, and the prediction mode used for the coding block can be configured to be horizontal or vertical via “intra_bdpcm_dir_flag”, which indicates the prediction direction. When “intra_bdpcm_flag” does not exist, its value is considered to be zero.
[0121] When "intra_bdpcm_dir_flag" indicating the prediction direction is zero, the BDPCM prediction direction is horizontal; when "intra_bdpcm_dir_flag" is 1, the BDPCM prediction direction is vertical.
[0122] Furthermore, as mentioned above, an in-loop filtering process can be performed on the reconstructed image. A modified reconstructed image can be generated through this in-loop filtering process, and this modified reconstructed image can be output as a decoded image in the decoding device. Additionally, the modified reconstructed image can be stored in the decoded image buffer or memory of the encoding / decoding device, and then used as a reference image in the inter-frame prediction process during image encoding / decoding.
[0123] The in-loop filtering process may include a deblocking filtering process, a Sample Adaptive Shift (SAO) process, and / or an Adaptive Loop Filter (ALF) process. In this case, one or a portion of the deblocking filtering process, the SAO process, the ALF process, and the bilateral filter process may be applied sequentially, or all of them may be applied sequentially. For example, the deblocking filtering process may be applied to the reconstructed image, and then the SAO process may be performed. Alternatively, for example, the ALF process may be performed after the deblocking filtering process has been applied to the reconstructed image. This can be performed in the same manner in the encoding device.
[0124] Deblocking filtering is a filtering scheme that removes distortion occurring at the boundaries between blocks in a reconstructed image. Based on the deblocking filtering process, the target boundary is derived in the reconstructed image, the boundary strength (bS) of the target boundary is determined, and deblocking filtering can be performed on the target boundary based on bS. bS can be determined based on factors such as the prediction patterns of two block patterns adjacent to the target boundary, the difference in motion vectors, whether the reference image is identical, and the presence of non-zero effective coefficients.
[0125] SAO (Side Array Filtering) is a method used to compensate for the offset difference between the reconstructed image and the original image on a sample-by-sample basis, and can be applied based on types such as band offset and edge offset. According to SAO, samples can be classified into different categories based on each SAO type, and an offset value can be added to each sample based on the category. SAO filtering information can include information about whether SAO is applied, SAO type information, SAO offset value information, etc. After applying deblocking filtering, SAO can also be applied to the reconstructed image.
[0126] An Adaptive Loop Filter (ALF) is a filtering scheme based on filter coefficients on a sample-by-sample basis, using the filter shape for image reconstruction. The encoding device determines whether an ALF is applied, the ALF shape, and / or the ALF filter coefficients by comparing the reconstructed image with the original image and signals the result to the decoding device. In other words, the ALF filtering information can include information about whether an ALF is applied, ALF shape information, and ALF filter coefficient information. An ALF can also be applied to the reconstructed image after deblocking filtering.
[0127] Figure 4This is a control flowchart describing the deblocking filtering process according to an implementation method.
[0128] Deblocking filtering is applied to the reconstructed image. Deblocking filtering is performed on each CU of the reconstructed image in the same order as the decoding process. First, vertical edges are filtered (horizontal filtering), then horizontal edges are filtered (vertical filtering). Deblocking filtering can be applied to the edges of transform blocks of the image as well as the edges of all coded blocks or sub-blocks. The output of the deblocking filter can be referred to as the modified reconstructed image or the modified reconstructed sample / sample array.
[0129] like Figure 4 As shown, the encoding and decoding devices can deduce the target boundaries to be filtered in the reconstructed image (step S1400).
[0130] Subsequently, the encoding and decoding devices can derive the boundary strength bS (step S1410).
[0131] bS can be determined based on two blocks facing the target boundary. For example, bS can be determined based on the following table.
[0132] [Table 3]
[0133]
[0134]
[0135] Here, p and q represent samples of two blocks facing the target boundary. For example, p0 can represent a sample of the left or top block facing the target boundary, and q0 can represent a sample of the right or bottom block facing the target boundary. When the edge direction of the target boundary is vertical, p0 can represent a sample of the left block facing the target boundary, and q0 can represent a sample of the right block facing the target boundary. When the edge direction of the target boundary is horizontal, p0 can represent a sample of the top block facing the target boundary, and q0 can represent a sample of the bottom block facing the target boundary.
[0136] The encoding and decoding devices can apply filtering based on bS (step S1420).
[0137] When bS is zero, filtering is not applied to the target boundary. Filtering can be performed based on filter strength (strong or weak) and / or filter length.
[0138] Furthermore, the filter strength based on the reconstructed average brightness level can be derived as follows. Figure 5 This is a diagram illustrating a sample located at the boundary of a block.
[0139] In HEVC, the filter strength for deblocking filtering can be determined from the average quantization parameter qP. LThe derived variables β and t C Control. In VVC, deblocking filtering can be achieved by adding an offset to qP based on the reconstructed average luminance level. L This controls the intensity of the deblocking filter. The reconstructed brightness level LL can be derived as shown in the following equation.
[0140] [Formula 5]
[0141] LL=((p 0,0 +p 0,3 +q 0,0 +q 0,3 )>>2) / (1<<bitDepth)
[0142] Here, you can Figure 5 The sample value p in the middle is identified i,k and q i,k The position of , where i is from 0 to 3, and k is from 0 to 3.
[0143] variable qP L It can be derived as shown in the following formula.
[0144] [Formula 6]
[0145] qP L =((Qp) Q +Qp P +1)>>1)+qpOffset
[0146] Here, Q pQ and Q pP They respectively represent samples q 0,0 and p 0,0 The quantization parameters, and the offset value qpOffset that depends on the transformation process, can be signaled in the Sequence Parameter Set (SPS).
[0147] Furthermore, stronger filtering can be applied to luminance samples. For samples located at a side boundary that belong to a large block, a replicated linear filter (a stronger deblocking filter) can be applied. A sample belonging to a large block can be defined as a sample belonging to each boundary where the vertical boundary width is equal to or greater than 32 or the horizontal boundary height is equal to or greater than 32.
[0148] In addition, strong filtering of chroma samples can be performed as shown in the following formula.
[0149] [Formula 7]
[0150] p2′=(3*p3+2*p2+p1+p0+q0+4)>>3
[0151] p1′=(2*p3+p2+2*p1+p0+q0+q1+4)>>3
[0152] p0′=(p3+p2+p1+2*p0+q0+q1+q2+4)>>3
[0153] The chroma filtering expressed in Equation 7 is performed on an 8×8 chroma sample raster. Strong filtering of the chroma samples can be performed on both sides of the block boundary. Here, chroma filtering is selected when the two edges of the chroma sample are equal to or greater than 8 in chroma sample units, and chroma filtering is performed if the following three conditions are met: First, the boundary strength (bS) and the block size are determined; the second and third determinations are to determine whether the filtering is on / off and to determine the strong filter, respectively, which is essentially the same as the determination for HEVC luma blocks. Examples of determining the boundary strength (bS) of the chroma block are shown in the following table.
[0154] [Table 4]
[0155]
[0156] As shown in Table 4, deblocking of chroma samples can be performed when bS equals 2, or when bS equals 1 if a large block boundary is detected. The second and third conditions are essentially the same as those applied to the strong filtering determination of HEVC luminance samples.
[0157] In VVC, deblocking filters can be applied to sub-block boundaries, and as shown in the example, deblocking filters can be performed on an 8×8 grid. Deblocking filtering can be applied to both CU boundaries and sub-block boundaries aligned with the 8×8 grid.
[0158] Subblock boundaries may include prediction unit (PU) boundaries introduced by subblock-based temporal motion vector prediction (STMVP) and affine modes, and transform unit (TU) boundaries introduced by subblock transform (SBT) and ISP (intra-framing sub-segmentation) modes.
[0159] For sub-blocks on an 8×8 grid using SBT or ISP, the same procedure as that applied to the TU deblocking filter in HEVC is used. Deblocking filtering is applied to the TU boundaries on the 8×8 grid when any of the sub-blocks spanning the edge has non-zero coefficients.
[0160] For sub-blocks on a raster grid with STMVP and affine modes, the same procedure as that applied to the TU deblocking filter in HEVC is used. The deblocking filter is applied to the PU boundaries on the 8×8 raster, taking into account the difference between the reference image and motion vector of adjacent sub-blocks.
[0161] In the following section, a deblocking filtering method is proposed for in-loop filtering of blocks for which BDPCM is applied. In the example, during the encoding and decoding of an image or video, when the luma block is BDPCM encoded in a single-tree type, the boundary strength (bS) of the corresponding chroma block can be set to the same as that of the luma block. That is, in the case of a single-tree type, to determine whether deblocking filtering is applied to blocks encoded in BDPCM, the bS of the chroma block can be determined as follows when calculating bS.
[0162] Figure 6 This is a diagram illustrating a method for determining bS according to an embodiment of the present disclosure.
[0163] According to the example, when both blocks are edge-based BDPCM encoded, bS can be derived to be zero. That is, for chroma blocks, regardless of whether BDPCM is applied, if the corresponding luma block is BDPCM encoded, the bS of the chroma block can be derived to be zero in a single-tree type.
[0164] exist Figure 6 In the diagram, the large block on the left can represent the luminance block, and the small block on the right can represent the chrominance block based on the components (Cb and Cr).
[0165] like Figure 6 As shown, bS in the block boundary can be determined according to the encoding scheme of two adjacent luma blocks, and the bS of the chroma block can be set to be the same as the bS of the luma block.
[0166] The two adjacent luma blocks shown above can be encoded using both BDPCM and inter-frame prediction, and in this case, bS can be set to 1 or 0. Similarly, the two adjacent luma blocks shown below can be encoded using both BDPCM and intra-frame prediction, and in this case, bS can also be set to 1 or 0. In both cases, the bS of the chroma block can be set to the same value as the bS of the luma block.
[0167] Additionally, when two adjacent luma blocks are both encoded in BDPCM (similar to the luma block shown in the middle), the bS used for the boundary between the two blocks can be set to zero. In this case, the bS used for deblocking filtering can also be set to zero for the chroma block corresponding to the luma block.
[0168] Whether to encode luminance blocks according to the BDPCM scheme can be signaled using information such as the intra_bdpcm_flag flag shown in Table 1.
[0169] Therefore, when bS is set to zero, deblocking filtering is not performed. That is, when two adjacent luma blocks are encoded according to the BDPCM scheme of the implementation, the bS of both the luma and chroma blocks is set to zero, and deblocking filtering is not performed.
[0170] The method for determining bS is shown in the table below.
[0171] [Table 5]
[0172]
[0173]
[0174]
[0175] In Table 5, as described above, when the target boundary is a vertical boundary, the left block can be designated as P and the right block as Q, based on the target boundary. Furthermore, when the target boundary is a horizontal boundary, the upper block can be designated as P and the lower block as Q, based on the target boundary.
[0176] The variable bS[xDi][yDj] used for bS can be derived as one of 0, 1, and 2.
[0177] As shown in Table 5, when both samples p0 and q0 are in a coded block with an intra_bdpcm_flag equal to 1, bS[xDi][yDj] is set to 0. In this case, deblocking filtering is not performed.
[0178] Otherwise, if sample p0 or q0 is in a coding block of a coding unit encoded using intra-prediction mode, set bS[xDi][yDj] to equal 2.
[0179] Additionally, if the block edge is also a transform block edge and the sample p0 or q0 is in a transform block containing one or more non-zero transform coefficient levels, bS[xDi][yDj] is set to equal to 1.
[0180] Alternatively, when encoding two blocks with different prediction modes, or using different reference images or different numbers of motion vectors, bS[xDi][yDj] can be set to equal 1.
[0181] Additionally, according to another example, bdpcm flag information can be signaled as shown in the table below.
[0182] [Table 6]
[0183]
[0184] Table 6 shows the "sps_bdpcm_enabled_flag" that signals the sequence parameter set (SPS). When the syntax element "sps_bdpcm_enabled_flag" is 1, this indicates whether BDPCM is applied to the coding unit performing intra-frame prediction; that is, "intra_bdpcm_luma_flag" and "intra_bdpcm_chroma_flag" exist in the coding unit.
[0185] In the absence of the syntax element "sps_bdpcm_enabled_flag", this value is considered to be zero.
[0186] [Table 7]
[0187]
[0188]
[0189] The syntax elements “intra_bdpcm_luma_flag” and “intra_bdpcm_chroma_flag” shown in Table 6 indicate whether BDPCM is applied to the current luma coding block or the current chroma coding block. When either “intra_bdpcm_luma_flag” or “intra_bdpcm_chroma_flag” is 1, the transformation for the coding block can be skipped, and the prediction mode for the coding block can be set in the horizontal or vertical direction using “intra_bdpcm_luma_dir_flag” or “intra_bdpcm_chroma_dir_flag”, which indicates the prediction direction. If either “intra_bdpcm_luma_flag” or “intra_bdpcm_chroma_flag” is absent, its value is considered zero.
[0190] The "intra_bdpcm_chroma_flag" signal can be sent to chroma-encoded blocks that are either single-tree or dual-tree chroma-encoded blocks.
[0191] When either "intra_bdpcm_luma_dir_flag" or "intra_bdpcm_chroma_dir_flag" representing the prediction direction is zero, the BDPCM prediction direction is horizontal; and when either "intra_bdpcm_luma_dir_flag" or "intra_bdpcm_chroma_dir_flag" is 1, the BDPCM prediction direction is vertical.
[0192] The method for determining bS based on the flag information shown in Tables 6 and 7 is as follows.
[0193] [Table 8]
[0194]
[0195] As shown in Table 8, according to the example, whether two adjacent luma blocks are encoded in BDPCM can be indicated by intra_bdpcm_luma_flag. When the value of intra_bdpcm_luma_flag is 1, bS[xDi][yDj] for the luma block can be set to zero.
[0196] Additionally, whether two adjacent chroma blocks are encoded in BDPCM can be indicated by intra_bdpcm_chroma_flag. When intra_bdpcm_chroma_flag is 1, bS[xDi][yDj] for the chroma block can be set to zero.
[0197] In other words, a signal can be sent to each of the luma and chroma blocks to notify the flag information used for BDPCM, and when the flag value is 1, bS used for deblocking filtering can be set to zero. That is, for blocks encoded according to the BDPCM scheme, deblocking filtering can be omitted.
[0198] The following figures are provided to illustrate specific examples of this disclosure. Since the specific names of the devices or signals / messages / fields shown in the figures are for illustrative purposes only, the technical features of this disclosure are not limited to the specific names used in the following figures.
[0199] Figure 7 This is a control flowchart used to describe a video decoding method according to embodiments of the present disclosure.
[0200] Decoding device 300 receives encoded information such as BDPCM information from the bitstream (step S710).
[0201] BDPCM information may include BDPCM flag information, which indicates whether BDPCM is applied to the current block and direction information for the direction in which BDPCM is executed.
[0202] The BDPCM flag value can be 1 when BDPCM is applied to the current block, and the BDPCM flag value can be zero when BDPCM is not applied to the current block.
[0203] Furthermore, the tree type of the current block can be classified as either a single-tree (SINGLE_TREE) or a dual-tree (DUAL_TREE) based on whether the chroma block corresponding to the luma block has a separate partitioning structure. A single-tree block is defined when the chroma block has the same partitioning structure as the luma block, and a dual-tree block is defined when the chroma block has a different partitioning structure than the luma block. According to the example, BDPCM can be applied independently to either the luma block or the chroma block of the current block. When BDPCM is applied to the luma block, the transform index for the luma block may not be received, and when BDPCM is applied to the chroma block, the transform index for the chroma block may not be received.
[0204] When the current block has a double-tree structure, BDPCM can be applied to only one component block, and even when the current block has a single-tree structure, BDPCM can be applied to only one component block.
[0205] Furthermore, the orientation information of BDPCM can indicate the horizontal or vertical direction. Based on the orientation information, quantization information and prediction samples can be derived.
[0206] The decoding device 300 can derive the quantization transform coefficients of the current block, i.e., the untransformed quantized residual sample based on BDPCM, and derive the residual sample by performing dequantization on the quantized residual sample (step S720).
[0207] When BDPCM is applied to the current block, the residual information received in the decoding device 300 can be the difference between the quantized residuals. Depending on the BDPCM direction, the difference between the quantized residual of a previous vertical or horizontal line and the quantized residual of a specific line can be received. The decoding device 300 can add the quantized residual of the previous vertical or horizontal line to the difference in the received quantized residuals and derive the quantized residual of the specific line. The quantized residual can be derived based on Equation 3 or Equation 4.
[0208] As mentioned above, when BDPCM is applied to the current block, the dequantized transform coefficients can be derived as residual samples that have not undergone the transform process.
[0209] Intra-predictor 331 can perform intra-prediction for the current block based on BDPCM information (i.e., direction information for performing BDPCM) and derive prediction samples (step S730).
[0210] When BDPCM is applied to the current block, intra-frame prediction can be performed using BDPCM, which means that BDPCM can be applied only to intra-frame slices or intra-frame coded blocks predicted in intra-frame mode.
[0211] Intra-frame prediction can be performed based on the orientation information of BDPCM, and the intra-frame prediction mode of the current block can be either horizontal orientation mode or vertical orientation mode.
[0212] The decoding device 300 can generate a reconstructed image based on the derived residual samples and predicted samples (step S740).
[0213] The decoding device 300 can perform deblocking filtering on the reconstructed image. Deblocking filtering is one of the in-loop filtering methods based on BDPCM information (step S750). In this case, when the BDPCM flag information is 1, the boundary strength (bS) is derived to be zero, and deblocking filtering may not be performed.
[0214] According to the example, when the luma block is BDPCM encoded in a single-tree type, the boundary strength (bS) of the corresponding chroma block can be set to be the same as that of the luma block.
[0215] When two blocks are encoded using BDPCM based on the edges between the blocks, bS can be derived to be zero. That is, for chroma blocks, regardless of whether BDPCM is applied, if the corresponding luma block is BDPCM encoded, the bS of the chroma block can be derived to be zero in a single-tree type.
[0216] Therefore, when bS is set to zero, deblocking filtering is not performed. That is, when two adjacent luma blocks are encoded according to the BDPCM scheme of the implementation, the bS of both the luma and chroma blocks is set to zero, and deblocking filtering is not performed.
[0217] Alternatively, according to another example, whether two adjacent luma blocks are encoded in BDPCM can be indicated by `intra_bdpcm_luma_flag`. When `intra_bdpcm_luma_flag` is 1, `bS[xDi][yDj]` for the luma block can be set to zero. Additionally, whether two adjacent chroma blocks are encoded in BDPCM can be indicated by `intra_bdpcm_chroma_flag`. When `intra_bdpcm_chroma_flag` is 1, `bS[xDi][yDj]` for the chroma block can be set to zero.
[0218] In other words, a signal can be sent to each of the luma and chroma blocks to notify the flag information used for BDPCM, and when the flag value is 1, bS used for deblocking filtering can be set to zero. That is, for blocks encoded according to the BDPCM scheme, deblocking filtering can be omitted.
[0219] In summary, the BDPCM information received in the decoding device may include flag information indicating whether BDPCM is applied to the current block, and when the flag information is 1, the boundary strength (bS) used for deblocking filtering can be derived to be zero.
[0220] In this case, the current block can be either a luminance-coded block or a chrominance-coded block, and if the tree type of the current block is a single tree type and a chrominance-coded block, the boundary strength can be deduced to be 1 when the flag information is 1.
[0221] The following figures are provided to illustrate specific examples of this disclosure. Since the specific names of the devices or signals / messages / fields shown in the figures are for illustrative purposes only, the technical features of this disclosure are not limited to the specific names used in the following figures.
[0222] Figure 8 This is a control flowchart used to describe a video encoding method according to embodiments of the present disclosure.
[0223] The encoding device 200 can derive the prediction sample of the current block based on BDPCM (step S810).
[0224] The encoding device 200 can derive intra-prediction samples for the current block based on a specific direction in which BDPCM is performed. The specific direction can be vertical or horizontal, and prediction samples for the current block can be generated according to the intra-prediction mode.
[0225] Furthermore, the tree type of the current block can be classified as either a single-tree (SINGLE_TREE) or a dual-tree (DUAL_TREE) based on whether the chroma block corresponding to the luma block has a separate partitioning structure. A single-tree block is defined as the chroma block having the same partitioning structure as the luma block, while a dual-tree block is defined as the chroma block having a different partitioning structure than the luma block. According to the example, BDPCM can be applied independently to either the luma block or the chroma block of the current block.
[0226] When the current block has a double-tree structure, BDPCM can be applied to only one component block, and even when the current block has a single-tree structure, BDPCM can be applied to only one component block.
[0227] Alternatively, according to the example, BDPCM can be applied only if the width of the current block is a first threshold or less and the height of the current block is a second threshold or less. The first and second thresholds can be 32 and are set to the maximum height or maximum width of the transform block to which the transform is performed.
[0228] The encoding device 200 can derive the residual sample of the current block based on the prediction block (step S820) and generate a reconstructed image based on the residual sample and the prediction sample (step S830).
[0229] The encoding device 200 can perform deblocking filtering on the reconstructed image, which is one of the in-loop filtering methods based on BDPCM information (step S840). In this case, when the BDPCM flag information is 1, the boundary strength (bS) is derived to be zero, and deblocking filtering may not be performed.
[0230] According to the example, when the luma block is BDPCM encoded in a single-tree type, the boundary strength (bS) of the corresponding chroma block can be set to be the same as that of the luma block.
[0231] When two blocks are encoded using BDPCM based on the edges between the blocks, bS can be derived to be zero. That is, for chroma blocks, regardless of whether BDPCM is applied, if the corresponding luma block is BDPCM encoded, the bS of the chroma block can be derived to be zero in a single-tree type.
[0232] Therefore, when bS is set to zero, deblocking filtering is not performed. That is, when two adjacent luma blocks are encoded according to the BDPCM scheme of the implementation, the bS of both the luma and chroma blocks is set to zero, and deblocking filtering is not performed.
[0233] Alternatively, according to another example, depending on whether two adjacent luma blocks are encoded in BDPCM, bS[xDi][yDj] for the luma block can be set to zero, and depending on whether two adjacent chroma blocks are encoded in BDPCM, bS[xDi][yDj] for the chroma block can be set to zero.
[0234] Additionally, a signal can be sent to each of the luma and chroma blocks to notify the flag information used for BDPCM, and when the flag value is 1, the bS used for deblocking filtering can be set to zero. That is, for blocks encoded according to the BDPCM scheme, deblocking filtering can be omitted.
[0235] Subsequently, the encoding device 200 can derive the quantized residual information based on BDPCM.
[0236] The encoding device 200 can derive quantized residual information from the quantized residual sample of a specific line and the difference between the quantized residual sample of a previous vertical or horizontal line and the quantized residual sample of the specific line. That is, the difference between the quantized residual and the conventional residual is generated as residual information, which can be derived based on Equation 1 or Equation 2.
[0237] The encoding device 200 can encode the quantized residual information and encoding information (e.g., BDPCM information of BDPCM) of the current block (step S850).
[0238] BDPCM information may include BDPCM flag information, which indicates whether BDPCM is applied to the current block and direction information for the direction in which BDPCM is executed.
[0239] When BDPCM is applied to the current block, the BDPCM flag value can be encoded as 1, and when BDPCM is not applied to the current block, the BDPCM flag value can be encoded as zero.
[0240] In addition, as mentioned above, when the current block's tree structure is a double tree, BDPCM can be applied to only one component block, and even when the current block has a single tree structure, BDPCM can be applied to only one component block.
[0241] The orientation information of BDPCM can indicate the horizontal or vertical direction.
[0242] In this disclosure, at least one of quantization / dequantization and / or transformation / inverse transformation may be omitted. When quantization / dequantization is omitted, the quantization transformation coefficients may be referred to as transformation coefficients. When transformation / inverse transformation is omitted, the transformation coefficients may be referred to as coefficients or residual coefficients, or, for the sake of consistency, may still be referred to as transformation coefficients.
[0243] Furthermore, in this disclosure, quantization transform coefficients and transform coefficients can be referred to as transform coefficients and scaling transform coefficients, respectively. In this case, residual information can include information about the transform coefficients, and this information can be signaled via residual coding syntax. Transform coefficients can be derived based on residual information (or information about transform coefficients), and scaling transform coefficients can be derived through the inverse transform (scaling) of the transform coefficients. Residual samples can be derived based on the inverse transform (scaling) of the scaling transform coefficients. These details can also be applied / expressed in other parts of this disclosure.
[0244] In the above embodiments, the method is explained based on a flowchart using a series of steps or blocks. However, this disclosure is not limited to the order of the steps, and a step may be performed in a different order or sequence than described above, or a step may be performed concurrently with other steps. Furthermore, those skilled in the art will understand that the steps shown in the flowchart are not exclusive, and another step may be incorporated or one or more steps in the flowchart may be deleted without affecting the scope of this disclosure.
[0245] The methods described above according to this disclosure can be implemented in software form, and the encoding and / or decoding devices according to this disclosure can be included in devices for image processing such as televisions, computers, smartphones, set-top boxes, and display devices.
[0246] When the embodiments of this disclosure are implemented by software, the above methods can be implemented as modules (steps, functions, etc.) for performing the above functions. These modules can be stored in memory and can be executed by a processor. The memory can be internal or external to the processor and can be connected to the processor in various well-known ways. The processor may include application-specific integrated circuits (ASICs), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this disclosure can be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units shown in each figure can be implemented and executed on a computer, processor, microprocessor, controller, or chip.
[0247] Furthermore, the decoding and encoding devices using this disclosure can include multimedia broadcast transceivers, mobile communication terminals, home theater video devices, digital cinema video devices, surveillance cameras, video chat devices, real-time communication devices (such as video communication), mobile streaming devices, storage media, cameras, video-on-demand (VoD) service providers, over-the-top (OTT) video devices, internet streaming service providers, three-dimensional (3D) video devices, video telephony devices, and medical video devices, and can be used to process video signals or data signals. For example, over-the-top (OTT) video devices can include game consoles, Blu-ray players, internet access TVs, home theater systems, smartphones, tablet PCs, digital video recorders (DVRs), etc.
[0248] Furthermore, the processing methods of this disclosure can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having the data structure according to this disclosure can also be stored in a computer-readable recording medium. Computer-readable recording media include various storage devices and distributed storage devices for storing computer-readable data. Computer-readable recording media can include, for example, Blu-ray discs (BD), Universal Serial Bus (USB), ROM, PROM, EPROM, EEPROM, RAM, CD-ROM, magnetic tape, floppy disks, and optical data storage devices. In addition, computer-readable recording media include media implemented in the form of a carrier wave (e.g., transmission over the Internet). Furthermore, bitstreams generated by encoding methods can be stored in computer-readable recording media or transmitted via wired or wireless communication networks. Additionally, embodiments of this disclosure can be implemented as computer program products by program code, and the program code can be executed on a computer according to embodiments of this disclosure. The program code can be stored on a computer-readable carrier.
[0249] Figure 9 The structure of a content streaming system applying this disclosure is illustrated.
[0250] Furthermore, the content streaming system using this disclosure can generally include an encoding server, a streaming server, a web server, a media storage device, a user device, and a multimedia input device.
[0251] An encoding server is used to compress content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data to generate a bitstream, and then sends it to a streaming server. As another example, in cases where the multimedia input device, such as a smartphone, camera, or camcorder, directly generates the bitstream, the encoding server can be omitted. The bitstream can be generated by applying the encoding method or bitstream generation method disclosed herein. Furthermore, the streaming server can temporarily store the bitstream during the sending or receiving process.
[0252] The streaming server sends multimedia data to the user's device via a web server based on the user's request. The web server acts as a tool to notify the user of available services. When a user requests a desired service, the web server transmits the request to the streaming server, and the streaming server sends the multimedia data to the user. In this context, the content streaming system may include a separate control server, which in this case controls the commands / responses between the corresponding devices within the content streaming system.
[0253] A streaming server can receive content from media storage devices and / or encoding servers. For example, when receiving content from an encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a predetermined period of time.
[0254] For example, user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigators, board-type PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, head-mounted displays (HMDs)), digital TVs, desktop computers, digital signage, etc. The servers in the content streaming system can operate as distributed servers, and in this case, data received by each server can be processed in a distributed manner.
[0255] The claims disclosed herein can be combined in various ways. For example, the technical features of the method claims can be combined to be implemented or performed in a device, and the technical features of the device claims can be combined to be implemented or performed in a method. Furthermore, the technical features of the method claims and the device claims can be combined to be implemented or performed in a device, and the technical features of the method claims and the device claims can be combined to be implemented or performed in a method.
Claims
1. A decoding device for video decoding, the decoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: Obtain a bitstream including block-based incremental pulse code modulation (BDPCM) flag information; The residual sample of the current block is derived based on the BDPCM flag information; Based on the BDPCM flag information, the predicted sample of the current block is derived; A reconstructed image is generated based on the residual samples and the predicted samples; Derive the boundary strength bS of the target boundary of the current block in the reconstructed image; Based on the bS, determine whether the deblocking filter is applied to the target boundary; as well as Based on the determined results, the deblocking filter is performed on the reconstructed image. The BDPCM flag information is related to whether BDPCM is applied to the current block. Wherein, the tree type of the current block is a single tree type, and the current block is a chroma-coded block. Wherein, based on the BDPCM flag information being 1, the bS of the target boundary of the chroma coding block is derived to be 0, and Where bS is 0 based on the target boundary, the deblocking filter is not applied to the target boundary.
2. An encoding device for video encoding, the encoding device comprising: Memory; as well as At least one processor, connected to the memory, is configured to: Derive the predicted sample for the current block based on block-based incremental pulse code modulation (BDPCM); The residual sample of the current block is derived based on the predicted sample; A reconstructed image is generated based on the residual samples and the predicted samples; Derive the boundary strength bS of the target boundary of the current block in the reconstructed image; Based on the bS, determine whether the deblocking filter is applied to the target boundary; Based on the determined result, the deblocking filter is performed on the reconstructed image; The quantized residual information is derived based on the aforementioned BDPCM; as well as The BDPCM flag information used for the BDPCM and the quantized residual information are encoded. The BDPCM flag information is related to whether the BDPCM is applied to the current block. Wherein, the tree type of the current block is a single tree type, and the current block is a chroma-coded block. Wherein, based on the application of the BDPCM to the current block, the bS of the target boundary of the chroma-coded block is derived to be 0, and Where bS is 0 based on the target boundary, the deblocking filter is not applied to the target boundary.
3. An apparatus for transmitting video data, the apparatus comprising: At least one processor is configured to obtain a bitstream for the video, wherein the bitstream is generated based on the following operations: deriving a prediction sample of the current block based on block-based incremental pulse code modulation (BDPCM); deriving a residual sample of the current block based on the prediction sample; generating a reconstructed image based on the residual sample and the prediction sample; deriving a boundary strength bS of the target boundary of the current block in the reconstructed image; determining whether deblocking filtering is applied to the target boundary based on bS; performing the deblocking filtering on the reconstructed image based on the determination result; deriving quantized residual information based on the BDPCM; and encoding BDPCM flag information for the BDPCM and the quantized residual information; and A transmitter configured to transmit the data comprising the bit stream. The BDPCM flag information is related to whether the BDPCM is applied to the current block. Wherein, the tree type of the current block is a single tree type, and the current block is a chroma-coded block. Wherein, based on the application of the BDPCM to the current block, the bS of the target boundary of the chroma-coded block is derived to be 0, and Where bS is 0 based on the target boundary, the deblocking filter is not applied to the target boundary.