Image encoding / decoding method and device, and recording medium on which bitstream is stored
By employing DIPMs derived from a divided reference region and applying filters, the method improves encoding efficiency and accuracy in intra prediction for high-resolution images, addressing limitations in existing technologies.
Patent Information
- Application Number
- PCT/KR2025/009416
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-04
- Filing Date
- 2025-07-02
- Publication Date
- 2026-01-08
AI Technical Summary
Existing image compression technologies face challenges in efficiently encoding and decoding high-resolution, high-quality images due to limitations in intra prediction methods, particularly in deriving optimal prediction modes for image blocks.
The method and device utilize Derived Intra Prediction Modes (DIPMs) based on a predetermined reference region, which is divided into sub-regions, and apply filters to derive and adaptively utilize these modes for improved encoding efficiency, incorporating diverse MPM candidates and reference sample lines for more accurate prediction.
This approach enhances encoding efficiency by deriving more optimized DIPMs, reducing complexity, and improving the accuracy of intra prediction while adaptively utilizing reference sample lines, thereby enhancing the encoding process.
Smart Images

Figure KR2025009416_08012026_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device, and recording medium storing bitstream
[0001] The present invention relates to a video encoding / decoding method and device, and a recording medium storing a bitstream.
[0002] Recently, the demand for high-resolution, high-quality images, such as HD (High Definition) images and UHD (Ultra High Definition) images, is increasing in various application fields, and accordingly, high-efficiency image compression technologies are being discussed.
[0003] There are various technologies for image compression, such as inter prediction technology that predicts pixel values included in the current picture from pictures before or after the current picture, intra prediction technology that predicts pixel values included in the current picture using pixel information within the current picture, and entropy encoding technology that assigns short codes to values with high frequency of appearance and long codes to values with low frequency of appearance, and these technologies can be used to effectively compress and transmit or store image data.
[0004] The present disclosure provides a method and device for inducing DIPM.
[0005] The present disclosure provides a method and device for constructing a reference area for DIPM.
[0006] The present disclosure provides a signaling method and device for adaptive utilization of DIPM mode.
[0007] The present disclosure provides a DIPM-based intra prediction method and device.
[0008] The video decoding method and device according to the present disclosure can derive one or more DIPMs (derived intra prediction modes) for a current block based on a predetermined reference region, generate a prediction block of the current block based on the one or more DIPMs, and reconstruct the current block based on the prediction block. Here, the reference region can be divided into a plurality of sub-regions.
[0009] In the image decoding method and device according to the present disclosure, the one or more DIPMs can be derived by applying a filter to samples of some areas within a sub-area.
[0010] In the image decoding method and device according to the present disclosure, some areas within the plurality of sub-areas to which the filter is applied may include the same number of samples.
[0011] In the image decoding method and device according to the present disclosure, some areas within the plurality of sub-areas to which the filter is applied may include different numbers of samples.
[0012] In the image decoding method and device according to the present disclosure, the prediction block can be generated based on one or more DIPMs and samples belonging to a predetermined reference sample line.
[0013] In the image decoding method and device according to the present disclosure, the predetermined reference sample line can be determined based on a sub-region used to derive one or more DIPMs.
[0014] In the image decoding method and device according to the present disclosure, the predetermined reference sample line can be set as a reference sample line at a pre-defined position.
[0015] In the image decoding method and device according to the present disclosure, the predetermined reference sample line can be determined based on whether the DIPM derived for the current block is a planar mode.
[0016] In the image decoding method and device according to the present disclosure, at least one of the one or more DIPMs can be added as an MPM candidate to the MPM list of the current block.
[0017] In the video decoding method and device according to the present disclosure, the MPM list may include intra prediction modes of neighboring blocks of the current block as MPM candidates. At least one of one or more DIPMs derived for the neighboring blocks may be set as the intra prediction mode of the neighboring block.
[0018] In the image decoding method and device according to the present disclosure, a primary mode among the one or more DIPMs derived for the current block may be stored in the current block. The primary mode may be a mode having the largest amplitude value among the one or more DIPMs.
[0019] The video encoding method and device according to the present disclosure can derive one or more DIPMs (derived intra prediction modes) for a current block based on a predetermined reference region, generate a prediction block of the current block based on the one or more DIPMs, generate a residual block of the current block based on the prediction block of the current block, derive transform coefficients of the current block based on the residual block, and encode residual information regarding the transform coefficients. Here, the reference region can be divided into a plurality of sub-regions.
[0020] A computer-readable digital storage medium is provided, which stores encoded video / image information that causes a decoding device according to the present disclosure to perform a video decoding method.
[0021] A computer-readable digital storage medium storing video / image information generated by a video encoding method according to the present disclosure is provided.
[0022] A method and device for transmitting video / image information generated by a video encoding method according to the present disclosure are provided.
[0023] According to the present disclosure, by deriving and utilizing DIPM, which is a mode more optimized for the current block, in an encoding device and a decoding device, the encoding efficiency of intra prediction can be improved.
[0024] According to the present disclosure, a more accurate DIPM can be derived based on various or variable reference regions, thereby improving the encoding efficiency of intra prediction.
[0025] According to the present disclosure, DIPM-related information can be efficiently signaled.
[0026] According to the present disclosure, the encoding efficiency of intra prediction can be improved while reducing the complexity of the DIPM derivation process by utilizing some areas within the reference area.
[0027] According to the present disclosure, the encoding efficiency of intra prediction can be improved by adaptively utilizing a reference sample line for intra prediction.
[0028] According to the present disclosure, by defining more diverse MPM candidates, it is possible to derive more accurate intra prediction modes.
[0029] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0030] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0031] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0032] FIG. 4 illustrates a decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.
[0033] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.
[0034] FIG. 6 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0035] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.
[0036] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0037] The present disclosure may be modified in various ways and encompasses numerous embodiments. Specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, but rather to encompass all modifications, equivalents, and alternatives falling within the spirit and technical scope of the present disclosure. Similar reference numerals have been used to designate similar components throughout the description of each drawing.
[0038] While terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present disclosure, a first component could be referred to as a "second component," and similarly, a second component could also be referred to as a "first component." The term "and / or" includes a combination of multiple related items described herein or any of multiple related items described herein.
[0039] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0040] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0041] The present disclosure relates to video / image coding. For example, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the versatile video coding (VVC) standard. In addition, the methods / embodiments disclosed in this specification can be applied to methods disclosed in the essential video coding (EVC) standard, the AOMedia Video 1 (AV1) standard, the second generation of audio video coding standard (AVN2), or the next generation of video / image coding standards (e.g., H.267 or H.268).
[0042] This specification presents various embodiments of video / image coding, and unless otherwise stated, the embodiments may be performed in combination with each other.
[0043] In this specification, a video may refer to a set of images over time. A picture generally refers to a unit representing one image at a specific time point, and a slice / tile is a unit that constitutes part of a picture in coding. A slice / tile may include one or more coding tree units (CTUs). A picture may be composed of one or more slices / tiles. A tile is a rectangular area consisting of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs that has a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs that has a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged consecutively according to the CTU raster scan, while tiles within a picture may be arranged consecutively according to the tile raster scan. A slice may contain an integer number of complete tiles or an integer number of contiguous complete CTU rows within a picture, which may be exclusively contained within a single NAL unit. Meanwhile, a picture may be divided into two or more subpictures. A subpicture may be a rectangular region of one or more slices within a picture.
[0044] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term "sample" can be used as a counterpart to a pixel. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component.
[0045] A unit may represent a basic unit of image processing. A unit may include at least one of a specific region of a picture and information related to the region. One unit may include one luma block and two chroma (e.g., cb, cr) blocks. In some cases, the term "unit" may be used interchangeably with terms such as "block" or "area." In general, an MxN block may include a set (or array) of samples (or sample array) or transform coefficients consisting of M columns and N rows.
[0046] As used herein, "A or B" can mean "only A," "only B," or "both A and B." In other words, as used herein, "A or B" can be interpreted as "A and / or B." For example, as used herein, "A, B or C" can mean "only A," "only B," "only C," or "any combination of A, B and C."
[0047] As used herein, a slash ( / ) or a comma can mean "and / or." For example, "A / B" can mean "A and / or B." Accordingly, "A / B" can mean "only A," "only B," or "both A and B." For example, "A, B, C" can mean "A, B, or C."
[0048] In this specification, "at least one of A and B" may mean "only A", "only B" or "both A and B". Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted identically to "at least one of A and B".
[0049] Additionally, in this specification, “at least one of A, B and C” can mean “only A,” “only B,” “only C,” or “any combination of A, B and C.” Additionally, “at least one of A, B or C” or “at least one of A, B and / or C” can mean “at least one of A, B and C.”
[0050] Additionally, parentheses used herein may mean "for example." Specifically, when "prediction (intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be suggested as an example of "prediction." Furthermore, even when "prediction (i.e., intra-prediction)" is indicated, "intra-prediction" may be suggested as an example of "prediction."
[0051] Technical features individually described in a single drawing in this specification may be implemented individually or simultaneously.
[0052] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0053] Referring to FIG. 1, a video / image coding system may include a first device (source device) and a second device (receiving device).
[0054] A source device can transmit encoded video / image information or data to a receiving device via a digital storage medium or a network in the form of a file or streaming. The source device may include a video source, an encoding device, and a transmitting device. The receiving device may include a receiving device, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. The transmitter may be included in the encoding device. The receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be configured as a separate device or an external component.
[0055] A video source may obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source may include a video / image capture device and / or a video / image generation device. The video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. The video / image generation device may include a computer, a tablet, a smartphone, etc., and may (electronically) generate video / images. For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process of generating related data.
[0056] An encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, to improve compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0057] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device via a digital storage medium or network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit can include an element for generating a media file via a predetermined file format and an element for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0058] The decoding device can decode the video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding device.
[0059] The renderer can render decoded video / images. The rendered video / images can be displayed through the display unit.
[0060] FIG. 2 is a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and in which encoding of a video / image signal is performed.
[0061] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a prediction unit (predictor) 220, a residual processor (residual processor) 230, an entropy encoder (entropy encoder) 240, an adder (adder) 250, a filter (filter) 260, and a memory (memory) 270. The prediction unit (220) may include an inter prediction unit (221) and an intra prediction unit (222). The residual processor (230) may include a transformer (transformer) 232, a quantizer (quantizer) 233, a dequantizer (dequantizer) 234, and an inverse transformer (inverse transformer) 235. The residual processing unit (230) may further include a subtractor (231). The addition unit (250) may be called a reconstructor or a recontructed block generator. The image segmentation unit (210), the prediction unit (220), the residual processing unit (230), the entropy encoding unit (240), the addition unit (250), and the filtering unit (260) described above may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to an embodiment. In addition, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0062] The image segmentation unit (210) can segment an input image (or picture, frame) input to the encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively segmented from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0063] For example, a single coding unit may be split into multiple coding units with deeper depths based on a quad-tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad-tree structure may be applied first, and the binary tree structure and / or the ternary structure may be applied later. Alternatively, the binary tree structure may be applied before the quad-tree structure. The coding procedure according to the present specification may be performed based on the final coding unit that is no longer split. In this case, based on coding efficiency according to image characteristics, etc., the largest coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively split into coding units of lower depths, and the coding unit with the optimal size may be used as the final coding unit. Here, the coding procedure may include procedures such as prediction, transformation, and restoration, which will be described later.
[0064] As another example, the processing unit may further include a prediction unit (PU) or a transform unit (TU). In this case, the prediction unit and the transform unit may each be split or partitioned from the final coding unit described above. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from a transform coefficient.
[0065] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the case. In general, an MxN block can represent a set of samples or transform coefficients consisting of M columns and N rows. A sample can generally represent a pixel or a pixel value, and can represent only the pixel / pixel value of the luminance component, or only the pixel / pixel value of the chrominance component. A sample can be used as a term corresponding to a pixel or pel in a picture (or image).
[0066] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, a unit that subtracts a prediction signal (prediction block, prediction sample array) from an input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).
[0067] The prediction unit (220) can perform a prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The prediction unit (220) can determine whether intra prediction or inter prediction is applied on a current block or CU basis. The prediction unit (220) can generate various information related to prediction, such as prediction mode information, as described later in the description of each prediction mode, and transmit the information to the entropy encoding unit (240). The information related to prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0068] The intra prediction unit (222) can predict the current block by referring to samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located a certain distance away from the current block depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one of a DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of detail in the prediction direction. However, this is only an example, and a greater or lesser number of directional modes may be used depending on the settings. The intra prediction unit (222) may also determine the prediction mode applied to the current block by using the prediction mode applied to the template region.
[0069] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of the motion information between the template region and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the template region can include a spatial template region (spatial neighboring block) existing in the current picture and a temporal template region (temporal neighboring block) existing in the reference picture. The reference picture including the reference block and the reference picture including the temporal template region may be the same or different. The temporal template region may be called a collocated reference block, a collocated CU (colCU), etc., and a reference picture including the temporal template region may be called a collocated picture (colPic). For example, the inter prediction unit (221) may configure a motion information candidate list based on template regions, and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, and for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the template region as the motion information of the current block. In the case of skip mode, unlike the merge mode, a residual signal may not be transmitted.In the motion vector prediction (MVP) mode, the motion vector of the template area is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0070] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within a picture can be signaled based on information about the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or a residual signal.
[0071] The transform unit (232) can apply a transform technique to the residual signal to generate transform coefficients. For example, the transform technique can include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transform obtained from a graph when the relationship information between pixels is expressed as a graph. CNT refers to a transform obtained based on generating a prediction signal using all previously restored pixels. In addition, the transform process can be applied to a pixel block having a square size and the same size, or can be applied to a block of a non-square variable size.
[0072] The quantization unit (233) quantizes the transform coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information about the quantized transform coefficients) and output it as a bitstream. The information about the quantized transform coefficients can be called residual information. The quantization unit (233) can rearrange the quantized transform coefficients in a block form into a one-dimensional vector form based on the coefficient scan order, and can also generate information about the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form.
[0073] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), etc. The entropy encoding unit (240) can also encode information necessary for video / image restoration (e.g., values of syntax elements, etc.) together or separately from quantized transform coefficients.
[0074] Encoded information (e.g., encoded video / image information) can be transmitted or stored in the form of a bitstream in units of NAL (network abstraction layer) units. The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. In the present specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted via a network or stored in a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media, such as a USB, SD, CD, DVD, Blu-ray, HDD, or SSD. The signal output from the entropy encoding unit (240) may be configured as an internal / external element of the encoding device (200) by a transmitting unit (not shown) and / or a storing unit (not shown), or the transmitting unit may be included in the entropy encoding unit (240).
[0075] The quantized transform coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, by applying inverse quantization and inverse transformation to the quantized transform coefficients through the inverse quantization unit (234) and the inverse transform unit (235), a residual signal (residual block or residual samples) can be reconstructed. The addition unit (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the reconstructed residual signal to the prediction signal output from the inter prediction unit (221) or the intra prediction unit (222). When there is no residual for the block to be processed, such as when skip mode is applied, the predicted block can be used as a reconstructed block. The addition unit (250) may be called a reconstructor or a reconstructed block generation unit. The generated restoration signal can be used for intra prediction of the next processing target block within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0076] The filtering unit (260) can improve subjective / objective picture quality by applying filtering to the restoration signal. For example, the filtering unit (260) can apply various filtering methods to the restoration picture to generate a modified restoration picture, and store the modified restoration picture in the memory (270), specifically, in the DPB of the memory (270). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0077] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter prediction unit (221). Through this, when inter prediction is applied, the encoding device can avoid prediction mismatch between the encoding device (200) and the decoding device, and can also improve encoding efficiency.
[0078] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter prediction unit (221). The memory (270) can store motion information of a block from which motion information in the current picture is derived (or encoded) and / or motion information of blocks in an already restored picture. The stored motion information can be transferred to the inter prediction unit (221) to be used as motion information in a spatial template area or motion information in a temporal template area. The memory (270) can store restored samples of restored blocks in the current picture and transfer them to the intra prediction unit (222).
[0079] FIG. 3 is a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and in which decoding of a video / image signal is performed.
[0080] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-prediction unit (332) and an intra-prediction unit (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).
[0081] The entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) described above may be configured by a single hardware component (e.g., a decoding device chipset or processor) depending on the embodiment. In addition, the memory (360) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0082] When a bitstream including video / image information is input, the decoding device (300) can restore the image corresponding to the process in which the video / image information is processed in the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division-related information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied in the encoding device. Accordingly, the processing unit of decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a ternary tree structure. One or more transform units may be derived from the coding unit. Then, the restored image signal decoded and output through the decoding device (300) can be reproduced through a reproduction device.
[0083] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through the entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information on various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). In addition, the video / image information may further include general constraint information. The decoding device can decode the picture further based on the information on the parameter set and / or the general constraint information. The signaling / received information and / or syntax elements described later in this specification can be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information in a bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transform coefficients for residuals. More specifically, the CABAC entropy decoding method receives a bin corresponding to each syntax element in the bitstream, determines a context model using information of the syntax element to be decoded and decoding information of the surrounding and decoding target blocks or information of symbols / bins decoded in the previous step, and predicts the occurrence probability of the bin according to the determined context model to perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the information of the decoded symbol / bin for the context model of the next symbol / bin after determining the context model. Information regarding prediction among the information decoded by the entropy decoding unit (310) is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and residual values on which entropy decoding is performed by the entropy decoding unit (310), i.e., quantized transform coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive a residual signal (residual block, residual samples, residual sample array). In addition, information regarding filtering among the information decoded by the entropy decoding unit (310) can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of a decoding device (300), or the receiving unit may be a component of an entropy decoding unit (310).
[0084] Meanwhile, a decoding device according to the present specification may be called a video / video / picture decoding device, and the decoding device may be divided into an information decoding device (video / video / picture information decoding device) and a sample decoding device (video / video / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), the inverse transformation unit (322), the addition unit (340), the filtering unit (350), the memory (360), the inter prediction unit (332), and the intra prediction unit (331).
[0085] The inverse quantization unit (321) can inverse quantize the quantized transform coefficients and output the transform coefficients. The inverse quantization unit (321) can rearrange the quantized transform coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scanning order performed in the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transform coefficients using quantization parameters (e.g., quantization step size information) and obtain transform coefficients.
[0086] In the inverse transform unit (322), the transform coefficients are inversely transformed to obtain a residual signal (residual block, residual sample array).
[0087] The prediction unit (320) can perform a prediction on the current block and generate a predicted block including prediction samples for the current block. The prediction unit (320) can determine whether intra-prediction or inter-prediction is applied to the current block based on the information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter-prediction mode.
[0088] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) can apply intra prediction or inter prediction for prediction of a single block, and can also apply intra prediction and inter prediction simultaneously. This can be called combined inter and intra prediction (CIIP) mode. In addition, the prediction unit can be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode can be used for content image / video coding such as games, such as screen content coding (SCC). IBC basically performs prediction within the current picture, but can be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC can utilize at least one of the inter prediction techniques described herein. Palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information about the palette table and palette index may be included and signaled in the video / image information.
[0089] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located in the neighborhood of the current block, or may be located a certain distance away from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The intra prediction unit (331) may also determine the prediction mode applied to the current block by using the prediction mode applied to the template region.
[0090] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information can be predicted in units of blocks, subblocks, or samples based on the correlation of the motion information between the template region and the current block. The motion information can include a motion vector and a reference picture index. The motion information can further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the template region can include a spatial template region (spatial neighboring block) existing in the current picture and a temporal template region (temporal neighboring block) existing in the reference picture. For example, the inter prediction unit (332) can construct a motion information candidate list based on the template regions, and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter prediction can be performed based on various prediction modes, and information about the prediction can include information indicating an inter prediction mode for the current block.
[0091] The addition unit (340) can generate a restoration signal (restored picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter-prediction unit (332) and / or intra-prediction unit (331)). When there is no residual for the block to be processed, such as when skip mode is applied, the prediction block can be used as the restoration block.
[0092] The addition unit (340) may be referred to as a restoration unit or restoration block generation unit. The generated restoration signal may be used for intra prediction of the next processing target block within the current picture, may be output after filtering as described below, or may be used for inter prediction of the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0093] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can apply various filtering methods to the restored picture to generate a modified restored picture, and transmit the modified restored picture to the memory (360), specifically, to the DPB of the memory (360). The various filtering methods can include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0094] The (modified) reconstructed picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of a block from which motion information is derived (or decoded) in the current picture and / or motion information of blocks in an already reconstructed picture. The stored motion information can be transferred to the inter prediction unit (332) to be used as motion information of a spatial template area or motion information of a temporal template area. The memory (360) can store reconstructed samples of reconstructed blocks in the current picture and transfer them to the intra prediction unit (331).
[0095] In this specification, the embodiments described in the filtering unit (260), the inter prediction unit (221), and the intra prediction unit (222) of the encoding device (200) can be applied to the filtering unit (350), the inter prediction unit (332), and the intra prediction unit (331) of the decoding device (300) in the same or corresponding manner, respectively.
[0096] FIG. 4 illustrates an image decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.
[0097] Referring to FIG. 4, one or more DIPMs (derived intra prediction modes) can be derived for a current block based on a predetermined reference area (S400).
[0098] The reference region, which is referenced to derive the DIPM, may be a previously restored region surrounding the current block. The reference region may include multiple reference sample lines. For example, the reference region may consist of up to 18 reference sample lines. However, this is merely an example, and the number of reference sample lines constituting the reference region may be less than or more than 18.
[0099] A reference region can be divided into multiple sub-regions. An index for identifying each sub-region can be assigned to each sub-region. When the reference region is divided into N sub-regions, indices 0 to (N-1) can be assigned to the first to Nth sub-regions, respectively. A smaller index can be assigned to a sub-region closer to the current block. Here, N can be an integer greater than or equal to 2. Each sub-region can be composed of K reference sample lines. K can be an integer greater than or equal to 1, 2, 3, or more.
[0100] Multiple sub-regions can be defined so as not to overlap each other. That is, a reference sample line belonging to one sub-region may not belong to another sub-region. For example, it is assumed that a reference region is divided into N sub-regions (first to Nth sub-regions), and each sub-region is composed of three reference sample lines. In this case, the first sub-region may be composed of reference sample lines numbered 0 to 2. The second sub-region may be composed of reference sample lines numbered 3 to 5. Similarly, the Nth sub-region may be composed of reference sample lines numbered (3*N-3) to (3*N-1). Here, the Pth reference sample line may mean a line that is P samples away from the left and / or top border of the current block. This can be equally applied to the embodiments described below.
[0101] Alternatively, multiple sub-regions may be defined to overlap each other. That is, at least one reference sample line belonging to one sub-region may belong to another sub-region. For example, a reference region may be composed of five reference sample lines. In this case, a first sub-region may be composed of reference sample lines numbered 0 to 2. A second sub-region may be composed of reference sample lines numbered 1 to 3. A third sub-region may be composed of reference sample lines numbered 2 to 4.
[0102] The number of available sub-areas may vary depending on the size of the current block. Here, the size may be defined as at least one of the width or height, the product of the width and height, the maximum / minimum of the width and height, or the ratio of the width and height.
[0103] For example, if the product of the width and height of the current block is less than X, K sub-regions can be used. If the product of the width and height of the current block is greater than or equal to X, L sub-regions can be used. Here, X, K, and L can each be an integer greater than or equal to 0. K can be an integer less than L. For example, if the product of the width and height of the current block is less than 64, at least 3 sub-regions can be used. If the product of the width and height of the current block is greater than or equal to 64, at least 4 sub-regions can be used. Alternatively, if the product of the width and height of the current block is less than 64, at least 1 sub-region can be used. If the product of the width and height of the current block is greater than or equal to 64, at least 2 sub-regions can be used.
[0104] For example, if at least one of the width or the height of the current block is less than X, K sub-regions can be used. If at least one of the width or the height of the current block is greater than or equal to X, M sub-regions can be used. Here, X, K, and M can each be an integer greater than or equal to 0. K can be an integer less than M. For example, if the width and the height of the current block are less than 8, at least 3 sub-regions can be used. If the width and the height of the current block are greater than or equal to 8, at least 4 sub-regions can be used. Alternatively, if the width and the height of the current block are less than 8, at least 1 sub-region can be used. If the width and the height of the current block are greater than or equal to 8, at least 2 sub-regions can be used.
[0105] For example, if the product of the width and height of the current block is less than X, K sub-regions can be used. If the product of the width and height of the current block is greater than or equal to X and less than Y, L sub-regions can be used. If the product of the width and height of the current block is greater than or equal to Y, J sub-regions can be used. Here, X, Y, K, L, and J can each be an integer greater than or equal to 0. K can be an integer less than L, and L can be an integer less than J. For example, if the product of the width and height of the current block is less than 64, at least 3 sub-regions can be used. If the product of the width and height of the current block is greater than or equal to 64 and less than 256, at least 4 sub-regions can be used. If the product of the width and height of the current block is greater than or equal to 256, at least 5 sub-regions can be used. Alternatively, if the product of the width and height of the current block is less than 64, at least one sub-region may be used. If the product of the width and height of the current block is greater than or equal to 64 and less than 256, at least two sub-regions may be used. If the product of the width and height of the current block is greater than or equal to 256, at least three sub-regions may be used.
[0106] For example, if at least one of the width or the height of the current block is less than X, N sub-regions can be used. If at least one of the width or the height of the current block is greater than or equal to X and less than Y, M sub-regions can be used. If at least one of the width or the height of the current block is greater than or equal to Y, L sub-regions can be used. Here, X, Y, N, M, and L can each be an integer greater than or equal to 0. N can be an integer less than M. M can be an integer less than L. For example, if the width and the height of the current block are less than 8, at least 3 sub-regions can be used. If the width and the height of the current block are greater than or equal to 8 and less than 16, at least 4 sub-regions can be used. If the width and the height of the current block are greater than or equal to 16, at least 5 sub-regions can be used. Alternatively, if the width and the height of the current block are less than 8, at least 1 sub-region can be used. If the width and height of the current block are greater than or equal to 8 and less than 16, at least two sub-areas may be used. If the width and height of the current block are greater than or equal to 16, at least three sub-areas may be used.
[0107] The above example divides block sizes into two or three categories, but this is merely an example. For example, the number of available sub-areas could be divided into N categories based on more granular block size conditions.
[0108] The positions (or numbers) of reference sample lines corresponding to available sub-areas may be pre-defined in the encoding device and the decoding device. Alternatively, available sub-areas may be defined by grouping a predetermined number of reference sample lines in ascending order of the numbers assigned to the reference sample lines.
[0109] Regardless of the size of the current block, the same number of sub-areas may be set to be used.
[0110] By applying a specific filter to the samples belonging to each sub-region, a histogram of gradient (HoG) can be derived for each sub-region.
[0111] Specifically, by applying a predetermined filter to a sample belonging to a sub-region, horizontal variation and vertical variation can be derived from the sample. A gradient can be derived based on the horizontal and vertical variation, and an intra prediction mode mapped to the derived gradient can be determined. The intra prediction mode mapped to the gradient can be an intra prediction mode having a direction most similar to the gradient. A predetermined amplitude value can be assigned / accumulated for the mapped intra prediction mode. Here, the amplitude value can be derived based on at least one of the magnitude of the horizontal variation and the magnitude of the vertical variation. For example, the amplitude value can be defined as the sum of the magnitudes of the horizontal variation and the vertical variation.
[0112] The samples to which a filter is applied within a sub-region may be all samples within the sub-region, or may be a subset of samples within a pre-defined location. Below, we will examine a method for deriving a HoG using samples from a subset of a sub-region.
[0113] Example A
[0114] Some regions within the first to Nth sub-regions to which filters are applied can be configured to contain the same number of samples. Regardless of the location or index of the sub-regions, the HoG can be derived using the same number of samples within the sub-regions.
[0115] For example, each partial region may include at least one of a left sub-region or an upper sub-region. Here, the left sub-region may refer to a region located to the left of the current block among the sub-regions, and the upper sub-region may refer to a region located to the top of the current block among the sub-regions.
[0116] The left sub-region may be defined as an area having a size of 3xH. The height (H) of the left sub-region may be equal to the height of the current block. The upper sub-region may be defined as an area having a size of Wx3. The width (W) of the upper sub-region may be equal to the width of the current block.
[0117] The upper and lower boundaries of the left sub-region may be continuous with the upper and lower boundaries of the current block, respectively. The left and right boundaries of the upper sub-region may be continuous with the left and right boundaries of the current block, respectively.
[0118] Example B
[0119] Some regions within the first to Nth sub-regions to which filters are applied may be configured to include different numbers of samples. Depending on the location or index of the sub-region, the number of samples used to derive the HoG may vary.
[0120] For example, sub-regions with larger indices may be configured to include a larger number of samples in the portion of the filter applied. In other words, sub-regions with larger indices can utilize a larger number of samples to derive the HoG.
[0121] Each partial region may include at least one of a left sub-region or an upper sub-region. Here, the left sub-region may refer to a region located to the left of the current block among the sub-regions, and the upper sub-region may refer to a region located to the upper of the current block among the sub-regions.
[0122] The left sub-region belonging to the first sub-region (i.e., the sub-region with an index of 0) can be defined as an area having a size of 3xH0. The height (H0) of the left sub-region can be equal to the height of the current block. The upper sub-region belonging to the first sub-region can be defined as an area having a size of W0x3. The width (W0) of the upper sub-region can be equal to the width of the current block.
[0123] The upper and lower boundaries of the left sub-region belonging to the first sub-region may be continuous with the upper and lower boundaries of the current block, respectively. The left and right boundaries of the upper sub-region belonging to the first sub-region may be continuous with the left and right boundaries of the current block, respectively.
[0124] The left sub-region belonging to the Nth sub-region (i.e. the sub-region with index (N-1)) is 3xH N-1 can be defined as an area with the size of . Here, the height (H) of the left sub-area N-1) may be greater than the height of the current block. For example, the left sub-region within the Nth sub-region according to the present embodiment may be an area extended in the upper and lower directions by (3*N)-sample lengths from the left sub-region within the Nth sub-region according to embodiment A. Alternatively, the left sub-region within the Nth sub-region according to the present embodiment may be an area extended in the upper direction by (3*N)-sample lengths from the left sub-region within the Nth sub-region according to embodiment A. Alternatively, the left sub-region within the Nth sub-region according to the present embodiment may be an area extended in the lower direction by (3*N)-sample lengths from the left sub-region within the Nth sub-region according to embodiment A.
[0125] The upper sub-area belonging to the Nth sub-area is W N-1 can be defined as an area with a size of x3. The width of the upper sub-area (W N-1 ) may be larger than the width of the current block. For example, the upper sub-region within the Nth sub-region according to the present embodiment may be an area extended in the left and right directions by (3*N)-sample lengths from the upper sub-region within the Nth sub-region according to embodiment A. Alternatively, the upper sub-region within the Nth sub-region according to the present embodiment may be an area extended in the left direction by (3*N)-sample lengths from the upper sub-region within the Nth sub-region according to embodiment A. Alternatively, the upper sub-region within the Nth sub-region according to the present embodiment may be an area extended in the right direction by (3*N)-sample lengths from the upper sub-region within the Nth sub-region according to embodiment A.
[0126] The aforementioned (3*N)-sample length expansion is merely an example. For example, the expansion could be N-sample length, or (2*N)-sample length. Alternatively, the expanded sample length could be defined based on the number of reference sample lines constituting the sub-region.
[0127] The upper and lower boundaries of the left sub-region belonging to the Nth sub-region may be discontinuous with the upper and lower boundaries of the current block, respectively. Alternatively, the upper boundary of the left sub-region belonging to the Nth sub-region may be continuous with the upper boundary of the current block, but the lower boundary of the left sub-region may be discontinuous with the lower boundary of the current block. Alternatively, the upper boundary of the left sub-region belonging to the Nth sub-region may be discontinuous with the upper boundary of the current block, but the lower boundary of the left sub-region may be continuous with the lower boundary of the current block. The left and right boundaries of the upper sub-region belonging to the Nth sub-region may be discontinuous with the left and right boundaries of the current block, respectively. Alternatively, the left boundary of the upper sub-region belonging to the Nth sub-region may be continuous with the left boundary of the current block, but the right boundary of the upper sub-region may be discontinuous with the right boundary of the current block. Alternatively, the left boundary of the upper sub-region belonging to the Nth sub-region may be discontinuous with the left boundary of the current block, but the right boundary of the upper sub-region may be continuous with the right boundary of the current block.
[0128] In the above-described embodiment, for the convenience of explanation, it is assumed that the width of the left sub-region and the height of the upper sub-region are 3, but this is not limited to this. That is, the above-described embodiment can be equally applied even when the width of the left sub-region and the height of the upper sub-region are values other than 3.
[0129] The aforementioned process can be repeated for samples within a sub-region by moving horizontally and / or vertically, thereby deriving one or more intra prediction modes with accumulated amplitude values. Accordingly, a HoG according to the present disclosure can be defined as a group of one or more intra prediction modes with accumulated amplitude values.
[0130] Example 1
[0131] An integrated HoG can be generated based on multiple HoGs derived for multiple sub-regions. The integrated HoG can include at least one of the intra prediction modes included in the multiple HoGs. Each of the intra prediction modes included in the integrated HoG can have an integrated amplitude value. Here, the integrated amplitude value can be derived based on the sum of the amplitude values of the corresponding intra prediction modes within the multiple HoGs.
[0132] Specifically, it is assumed that the reference region is divided into three sub-regions. At this time, HoG[0] for the first sub-region may include mode A, mode B, and mode C. In HoG[0], mode A, mode B, and mode C may have amplitude values of amplitude[0][A], amplitude[0][B], and amplitude[0][C], respectively. HoG[1] for the second sub-region may include mode A and mode B, but may not include mode C. In HoG[1], mode A and mode B may have amplitude values of amplitude[1][A] and amplitude[1][B], respectively. HoG[2] for the third sub-region may include mode A, but may not include mode B and mode C. In HoG[2], mode A may have amplitude value of amplitude[2][A]. The above mode X may mean the Xth intra prediction mode, which may be equally applied in the present specification.
[0133] A single integrated HoG may include mode A, mode B, and mode C contained in HoG[0], HoG[0], and HoG[2]. In the integrated HoG, the integrated amplitude value for mode A may be derived based on the sum of amplitude[0][A], amplitude[1][A], and amplitude[2][A], the integrated amplitude value for mode B may be derived based on the sum of amplitude[0][B] and amplitude[1][B], and the integrated amplitude value for mode C may be derived based on the amplitude value for mode C in HoG[0] (i.e., amplitude[0][C]).
[0134] In this way, if at least two of the plurality of HoGs each include the same intra prediction mode, a single integrated amplitude value for the intra prediction mode can be derived based on the sum of the amplitude values of the corresponding intra prediction modes. For an intra prediction mode included in only one of the plurality of HoGs, the amplitude value of the corresponding intra prediction mode can be set as the integrated amplitude value for the corresponding intra prediction mode.
[0135] The method described above can generate an integrated HoG, which is a group of one or more intra prediction modes with integrated amplitude values. The top M intra prediction modes can be selected from the integrated HoG in descending order of integrated amplitude values. DIPM(s) can be derived based on the top M intra prediction modes. Here, M can be an integer greater than or equal to 1.
[0136] For example, assume that a reference region is divided into five sub-regions, and each sub-region consists of three reference sample lines. In this case, a HoG (i.e., HoG[0]) can be derived based on a first sub-region consisting of reference sample lines 0 to 2. A HoG (i.e., HoG[1]) can be derived based on a second sub-region consisting of reference sample lines 3 to 5. A HoG (i.e., HoG[2]) can be derived based on a third sub-region consisting of reference sample lines 6 to 8. A HoG (i.e., HoG[3]) can be derived based on a fourth sub-region consisting of reference sample lines 9 to 11. A HoG (i.e., HoG[4]) can be derived based on a fifth sub-region consisting of reference sample lines 12 to 14. If a vertical mode is included in at least two of HoG[0], HoG[1], HoG[2], HoG[3], and HoG[4], an integrated amplitude value for the vertical mode can be derived based on the sum of the amplitude values of the corresponding vertical modes.
[0137] Example 2
[0138] An integrated HoG can be generated based on multiple HoGs derived for multiple sub-regions. The integrated HoG can include at least one of the intra prediction modes included in the multiple HoGs. Each of the intra prediction modes included in the integrated HoG can have an integrated amplitude value. Here, the integrated amplitude value can be derived based on a weighted sum of the amplitude values of the corresponding intra prediction modes within the multiple HoGs.
[0139] Specifically, it is assumed that the reference region is divided into three sub-regions. At this time, HoG[0] for the first sub-region may include mode A, mode B, and mode C. In HoG[0], mode A, mode B, and mode C may have amplitude values of amplitude[0][A], amplitude[0][B], and amplitude[0][C], respectively. HoG[1] for the second sub-region may include mode A and mode B, but may not include mode C. In HoG[1], mode A and mode B may have amplitude values of amplitude[1][A] and amplitude[1][B], respectively. HoG[2] for the third sub-region may include mode A, but may not include mode B and mode C. In HoG[2], mode A may have amplitude value of amplitude[2][A].
[0140] A single integrated HoG may include mode A, mode B, and mode C contained in HoG[0], HoG[0], and HoG[2]. In the integrated HoG, the integrated amplitude value for mode A may be derived based on a weighted sum of amplitude[0][A], amplitude[1][A], and amplitude[2][A], the integrated amplitude value for mode B may be derived based on a weighted sum of amplitude[0][B] and amplitude[1][B], and the integrated amplitude value for mode C may be derived based on the amplitude value for mode C in HoG[0] (i.e., amplitude[0][C]).
[0141] In this way, if at least two of the plurality of HoGs each include the same intra prediction mode, a single integrated amplitude value can be derived for the intra prediction mode based on a weighted sum of the amplitude values of the corresponding intra prediction modes. For an intra prediction mode included in only one of the plurality of HoGs, the amplitude value of the corresponding intra prediction mode can be set as the integrated amplitude value for the corresponding intra prediction mode.
[0142] The weights for the weighted sum of amplitude values can be determined based on the index of the HoG to which the intra prediction mode belongs or the location (or index) of the sub-region used to derive the HoG. For example, a smaller weight may be applied to a HoG derived from a sub-region closer to the current block. Conversely, a larger weight may be applied to a HoG derived from a sub-region closer to the current block. Alternatively, a fixed, pre-defined weight may be applied to each sub-region.
[0143] Through the above-described method, an integrated HoG, which is a group of one or more intra prediction modes with integrated amplitude values, can be generated. From the integrated HoG, the top M intra prediction modes can be selected in descending order of amplitude values. Based on the top M intra prediction modes, DIPM(s) can be derived. Here, M can be an integer greater than or equal to 1.
[0144] For example, the reference region can be divided into first to third sub-regions. A HoG (i.e., HoG[0]) can be derived based on the first sub-region consisting of reference sample lines 0 to 2. A HoG (i.e., HoG[1]) can be derived based on the second sub-region consisting of reference sample lines 3 to 5. A HoG (i.e., HoG[2]) can be derived based on the third sub-region consisting of reference sample lines 6 to 8. When a vertical mode is included in at least two of HoG[0], HoG[1], or HoG[2], a predetermined weight can be applied to the amplitude values of the corresponding vertical modes to derive an integrated amplitude value for the vertical mode.
[0145] Example 3
[0146] For each sub-region within a reference region, HoGs can be derived. At least one intra-prediction mode can be selected from the HoG for each sub-region, and a mode list for the corresponding HoG can be generated based on the selected intra-prediction mode. A single integrated mode list can be constructed based on the mode lists for the generated HoGs. DIPM(s) can be derived based on one or more intra-prediction modes belonging to the integrated mode list.
[0147] Depending on the distance between the current block and the sub-region, the number of intra-prediction modes selected from the HoG for that sub-region may vary. For example, a greater number of intra-prediction modes may be selected from the HoG for a sub-region closer to the current block.
[0148] For example, the reference region can be divided into first to third sub-regions. A HoG (i.e., HoG[0]) can be derived based on the first sub-region consisting of reference sample lines 0 to 2. A HoG (i.e., HoG[1]) can be derived based on the second sub-region consisting of reference sample lines 3 to 5. A HoG (i.e., HoG[2]) can be derived based on the third sub-region consisting of reference sample lines 6 to 8. A first mode list can be constructed by selecting three intra prediction modes from HoG[0]. The selected three intra prediction modes can be the top three intra prediction modes in descending order of amplitude values in HoG[0]. A second mode list can be constructed by selecting two intra prediction modes from HoG[1]. The two selected intra prediction modes may be the top two intra prediction modes in descending order of amplitude value in HoG[1]. A third mode list may be formed by selecting one intra prediction mode from HoG[2]. The one selected intra prediction mode may be the intra prediction mode with the largest amplitude value in HoG[2]. A single integrated mode list may be formed based on the intra prediction modes belonging to the first to third mode lists.
[0149] Example 4
[0150] For each sub-region within the reference region, HoGs can be derived. Specific costs for each HoG can be calculated. DIPM(s) can be derived based on the top L HoGs in ascending order of the derived costs. L can be an integer greater than or equal to 1. For example, when L is 1, any one of multiple HoGs can be selected based on the cost, and thus, signaling of an index for specifying any one of the multiple HoGs can be omitted.
[0151] Alternatively, multiple HoGs can be reordered in ascending order of the generated costs. An index specifying at least one of the reordered HoGs can be signaled. Based on the signaled index, DIPM(s) can be derived based on the specified HoG.
[0152] The cost for the above HoG can be calculated based on the difference between the predicted value and the restored value of the template region of the current block. The predicted value of the template region can be generated based on the top T intra prediction modes in descending order of amplitude values within the HoG. T can be an integer greater than or equal to 1. For example, when T is 1, the predicted value of the template region can be generated by performing intra prediction on the template region based on one intra prediction mode. When T is 2, the intra prediction can be performed based on two intra prediction modes to generate two predicted values (i.e., first and second predicted values), respectively, and the predicted value of the template region can be generated by a weighted sum of the first and second predicted values.
[0153] Through the aforementioned Example 4, DIPM can be derived by selectively utilizing some of the HoGs for the sub-regions.
[0154] Example 5
[0155] At least one reference sample line among multiple reference sample lines may be specified for the current block. An Head of Grading (HoG) may be derived based on the sub-region to which the specified reference sample line belongs. DIPM(s) may be derived based on the derived HoG.
[0156] At least one of the plurality of reference sample lines may be specified based on an index explicitly signaled for the current block. Here, the index may specify the number / position of a reference sample line (or sub-region) used to derive the HoG. One or more indices may be signaled for the current block. Alternatively, at least one of the plurality of reference sample lines may be pre-defined in the encoding device and the decoding device.
[0157] For example, if reference sample line 2 is specified for the current block, the HoG can be derived based on the sub-region consisting of reference sample lines 2 to 4. Alternatively, if reference sample line 2 and reference sample line 7 are specified for the current block, the HoG can be derived based on the sub-region consisting of reference sample lines 2 to 4, and the HoG can be derived based on the sub-region consisting of reference sample lines 7 to 9.
[0158] The DIPM derivation method discussed above can be applied to at least one of the luminance component or chrominance component of the current block.
[0159] In the case of the chrominance component of the current block (i.e., the chrominance block), the change in the sample value compared to the corresponding luminance block of the chrominance block may be relatively limited. Accordingly, for each of the two component blocks (Cb component block and Cr component block) constituting the chrominance block, an HoG independent of the luminance block can be derived, and DIPM(s) can be derived based on the derived HoG. That is, the HoG can be derived based on the reference region of the Cb component block, and DIPM(s) for the Cb component block can be derived based on the derived HoG. Here, the reference region of the Cb component block can include at least one of a previously restored region around the Cb component block or the previously restored luminance block. Similarly, the HoG can be derived based on the reference region of the Cr component block, and DIPM(s) for the Cr component block can be derived based on the derived HoG. Here, the reference area of the Cr component block may include at least one of the pre-restored area around the Cr component block or the pre-restored luminance block.
[0160] Alternatively, one HoG may be derived for a chrominance block based on a reference region of the Cb and Cr component blocks. Here, the reference region of the Cb and Cr component blocks may include at least one of a pre-reconstructed region around the Cb component block, a pre-reconstructed region around the Cr component block, or the pre-reconstructed luminance block. In this case, DIPM(s) may be derived based on one HoG derived for the Cb and Cr component blocks, and the derived DIPM(s) may be commonly applied to the Cb and Cr component blocks.
[0161] If samples belonging to a sub-region (or a portion thereof) are located on a CTU boundary, the reference region-based DIPM derivation method may not be applicable. For example, if samples belonging to a sub-region (or a portion thereof) belong to a different CTU or CTU row than the current block, the aforementioned DIPM derivation method based on that sub-region (or portion thereof) may not be applicable.
[0162] Samples located outside the CTU boundary among the samples belonging to the sub-region (or part of the region) may be replaced by padding or mirroring based on samples located inside the CTU boundary. Alternatively, samples located outside the CTU boundary among the samples belonging to the sub-region (or part of the region) may be restricted from being referenced when deriving the DIPM. Here, a sample located outside the CTU boundary may mean a sample located in a different CTU or CTU row from the current block. Conversely, a sample located inside the CTU boundary may mean a sample located in the same CTU or CTU row as the current block.
[0163] In addition to the above CTU boundaries, the above-described embodiments may also be applied equally to cases where samples belonging to a sub-region (or a portion of a sub-region) are located outside boundaries such as slice boundaries, tile boundaries, or picture boundaries. In this case, information regarding whether a reference region-based DIPM derivation method is applied and / or index information for identifying the sub-region may be signaled, or the signaling thereof may be omitted.
[0164] Referring to FIG. 4, a prediction block of a current block can be generated based on one or more DIPMs (S410).
[0165] For example, if one DIPM is derived, intra prediction can be performed based on the DIPM to generate a prediction block of the current block.
[0166] For example, when multiple DIPMs are induced, intra prediction can be performed based on the DIPMs to generate prediction blocks, and a prediction block of the current block can be generated based on a weighted sum of the generated prediction blocks.
[0167] For example, when multiple DIPMs are derived, at least one DIPM among the multiple DIPMs can be selected, and an intra prediction mode of the current block can be derived based on the selected DIPM. Intra prediction can be performed based on the intra prediction mode of the current block to generate a prediction block of the current block.
[0168] A prediction block of the current block may also be generated based on a combination of at least two of a plurality of DIPMs. Specifically, at least two DIPMs may be selected from the plurality of DIPMs, and intra prediction modes of the current block may be set based on the selected DIPMs. A plurality of prediction blocks may be generated based on the set plurality of intra prediction modes, and a prediction block of the current block may be generated based on a weighted sum of the plurality of prediction blocks.
[0169] For the selection of the above DIPM, an index indicating a DIPM among multiple DIPMs for the current block may be explicitly signaled. Alternatively, at least one of the multiple DIPMs may be selected based on the amplitude values of the multiple DIPMs. In this way, a prediction block may be generated by selectively utilizing some, but not all, of the multiple DIPMs.
[0170] The prediction block of the current block can be generated by performing intra prediction based on one or more DIPMs and samples belonging to a predetermined reference sample line.
[0171] The above-described reference sample line may be determined based on a sub-region (or an index of a sub-region) used to derive the DIPM of the current block. For example, if the index of the sub-region used to derive the DIPM of the current block is 1, one or more DIPMs may be derived based on samples of reference sample lines 3 to 5 belonging to the sub-region. A prediction block may be generated based on the derived one or more DIPMs. At this time, the prediction block may be generated by performing intra prediction based on samples of reference sample line 4 among reference sample lines 3 to 5. Alternatively, the prediction block may be generated by performing intra prediction based on samples of reference sample line 3 or 5. That is, intra prediction may be performed based on some, but not all, of the plurality of reference sample lines used to derive the DIPM.
[0172] Alternatively, intra prediction can be performed based on multiple reference sample lines used to derive the DIPM. For example, intra prediction can be performed based on reference sample lines 3 to 5 to generate prediction blocks, and a prediction block of the current block can be generated based on a weighted sum of the generated prediction blocks.
[0173] The above-described reference sample line may be set as a reference sample line at a pre-defined position. For example, the reference sample line at the pre-defined position may be a reference sample line closest to the current block. Even if the DIPM of the current block is derived based on a sub-region with an index of 1, a prediction block of the current block may be generated by performing intra prediction based on the reference sample line closest to the current block (i.e., reference sample line 0) and the pre-derived DIPM, instead of reference sample lines 3 to 5 belonging to the sub-region.
[0174] The above-described reference sample line may also be selected based on a pre-derived DIPM. For example, if the DIPM of the current block is derived in planar mode (or the pre-derived DIPM includes a planar mode), intra prediction based on multiple reference sample lines may not be performed. In this case, intra prediction based on reference sample line 0 may be restricted to be performed. If the DIPM of the current block is derived in DC mode (or the pre-derived DIPM includes a DC mode), intra prediction based on multiple reference sample lines may not be performed. In this case, intra prediction based on reference sample line 0 may be restricted to be performed.
[0175] If the samples belonging to the multiple reference sample lines are located on the CTU boundary, intra prediction based on the multiple reference sample lines may not be performed. For example, if the samples belonging to the multiple reference sample lines belong to a different CTU or CTU row than the current block, intra prediction based on the multiple reference sample lines may not be performed.
[0176] Among the samples belonging to the plurality of reference sample lines, samples located outside the CTU boundary may be replaced by performing padding or mirroring based on samples located inside the CTU boundary. Alternatively, samples located outside the CTU boundary among the samples belonging to the plurality of reference sample lines may be restricted from being referenced for intra prediction of the current block. Here, a sample located outside the CTU boundary may refer to a sample located in a different CTU or CTU row from the current block. Conversely, a sample located inside the CTU boundary may refer to a sample located in the same CTU or CTU row as the current block.
[0177] In addition to the above CTU boundary, the above-described embodiment can be equally applied to cases where samples belonging to multiple reference sample lines are located outside boundaries such as slice boundaries, tile boundaries, and picture boundaries.
[0178] A most probable mode (MPM) list can be constructed for the current block. At least one intra prediction mode can be derived for the current block based on at least one of the multiple MPM candidates in the MPM list. To this end, an MPM index can be signaled to specify at least one of the multiple MPM candidates. A prediction block of the current block can be generated based on the derived intra prediction mode.
[0179] At least one of the DIPMs derived based on the aforementioned reference area can be added as an MPM candidate to the MPM list of the current block.
[0180] For example, a primary mode for each sub-region within a reference region may be added to the MPM list. Here, the primary mode may be a mode with the largest amplitude value among multiple DIPMs derived based on the sub-region.
[0181] In the MPM list, a DIPM derived based on the existing DIPM derivation method (hereinafter referred to as a conventional DIPM) may be added as an MPM candidate. In this case, it may be determined whether a conventional DIPM identical to the first mode exists in the MPM list. If it is determined that a conventional DIPM identical to the first mode does not exist, the first mode may be added to the MPM list.
[0182] The MPM list may include intra prediction modes of neighboring blocks of the current block as MPM candidates. Here, the neighboring blocks may include at least one of neighboring blocks or non-adjacent blocks of the current block.
[0183] When a DIPM derivation method based on a reference region is applied to a neighboring block, at least one of the one or more DIPMs derived based on the reference region of the neighboring block may be set to the intra prediction mode of the neighboring block and added to the MPM list. At this time, the MPM candidate may be formed by combining it with the conventional DIPM of other neighboring blocks to which the conventional DIPM derivation method is applied, or the MPM candidate may be formed only with the neighboring block(s) to which the reference region-based DIPM derivation method is applied.
[0184] When a reference region-based DIPM derivation method is applied to a neighboring block, a primary mode among one or more DIPMs derived based on the reference region of the neighboring block may be set as the intra prediction mode of the neighboring block and added to the MPM list. Here, the primary mode may be a mode having the largest amplitude value among one or more DIPMs derived based on the reference region of the neighboring block. Alternatively, when a reference region-based DIPM derivation method is applied to a neighboring block, the intra prediction mode of the neighboring block may be set as a specific intra prediction mode and added to the MPM list. For example, the specific intra prediction mode may be a planar mode.
[0185] When the existing DIPM derivation method is applied to a surrounding block, the primary mode among one or more DIPMs derived based on the existing DIPM derivation method may be set as the intra prediction mode of the surrounding block and added to the MPM list. Here, the primary mode may be the mode with the largest amplitude value among one or more DIPMs derived based on the existing DIPM derivation method.
[0186] Referring to FIG. 4, the current block can be restored based on the predicted block of the current block (S420).
[0187] Transform coefficients can be derived based on residual information signaled through the bitstream. A residual block can be generated by applying at least one of inverse quantization or inverse transformation to the derived transform coefficients. A reconstruction block of the current block can be generated based on the prediction block of the current block and the residual block.
[0188] If the current block is encoded based on intra prediction, a given intra prediction mode may be stored in the current block.
[0189] For example, when the existing DIPM derivation method is applied to the current block, the primary mode according to the existing DIPM derivation method can be stored in the current block. Here, the existing DIPM derivation method may mean a method of deriving DIPM based on the first sub-region described above (i.e., a sub-region with an index of 0). The primary mode may be a mode having the largest amplitude value among the DIPMs derived through the existing DIPM derivation method. Alternatively, the primary mode may be a mode corresponding to index information specifying any one of the DIPMs derived through the existing DIPM derivation method.
[0190] When the aforementioned reference area-based DIPM derivation method is applied to the current block, the primary mode according to the reference area-based DIPM derivation method may be stored in the current block. Here, the primary mode may be the mode with the largest amplitude value among the DIPMs derived through the reference area-based DIPM derivation method.
[0191] When the aforementioned reference region-based DIPM derivation method is applied to the current block, the primary mode according to the existing DIPM derivation method may be stored in the current block. Alternatively, when the aforementioned reference region-based DIPM derivation method is applied to the current block, a specific intra prediction mode may be stored in the current block. For example, the specific intra prediction mode may be a planar mode.
[0192] In the proposed method according to the present disclosure, weights for indices or HoG integration may be selected based on at least one of a prediction mode (e.g., intra mode, inter mode, etc.), a width of a current block, a height of the current block, a number of samples belonging to the current block, a position of a sub-block within the current block, an explicitly signaled syntax element, a statistical characteristic of surrounding samples of the current block, or whether a secondary transform is used. Here, the index may include at least one of an index for specifying at least one of a plurality of HoGs, an index for specifying at least one of a plurality of sub-regions, an index indicating a reference sample line, or an index for indicating at least one of a plurality of DIPMs.
[0193] Information about the selected index / weight can be binarized using a predetermined binarization method and transmitted to the decoding device. During the binarization process, context modeling can be used to save binarization bits. Considering the total number of predefined sub-regions in the encoding and decoding devices, the information can also be binarized using a binarization method such as truncated binary, truncated unary, or fixed-length.
[0194] The inverse transformation according to the present disclosure can be performed based on at least one of a first transformation or a second transformation. The transformation type for the inverse transformation can be selected based on at least one of a neural network matrix mode, the width of the current block, the height of the current block, the number of samples belonging to the current block, the position of a sub-block within the current block, an explicitly signaled syntactic element, or a statistical characteristic of surrounding samples of the current block. For example, for a block to which the proposed method according to the present disclosure is applied, the first transformation can be performed based on a transformation method based on MTS (multiple transform selection). For a block to which the proposed method according to the present disclosure is applied, the transformation kernel for the second transformation can be selected based on a DIPM derived for the current block.
[0195] The applicability (or availability) of the proposed method according to the present disclosure can be signaled in the high-level syntax (HLS). The high-level syntax can include at least one of a VPS, an SPS, a PPS, a picture header, a slice header, or decoding capability information (DCI). For example, the applicability of the proposed method according to the present disclosure can be determined on a per-PPS basis.
[0196] The applicability (or availability) of the proposed method according to the present disclosure may be adaptively determined by the encoding device and the decoding device without the aforementioned signaling.
[0197] Information regarding the applicability of the proposed method according to the present disclosure may be additionally signaled, and the applicability of the proposed method may be determined based on the signaled information. For example, a flag indicating the applicability of the proposed method may be signaled at the coding tree unit (CTU) or coding unit (CU) level. If the flag indicates that the proposed method is applied (e.g., if the flag is 1 or True), the aforementioned index may be signaled.
[0198] The proposed method according to the present disclosure can be applied when the decoder-side intra-mode derivation (DIMD) mode is defined in HLS. Furthermore, the applicability of the proposed method according to the present disclosure can be determined through signaling of additional information within the DIMD mode.
[0199] If the proposed method according to the present disclosure is determined to be applicable to the current block based on the size / shape of the current block or whether certain conditions are satisfied, a flag indicating whether the proposed method is applicable may be signaled. For example, if the height of the current block is more than four times the width of the current block, the proposed method according to the present disclosure may be determined to be unavailable to the current block, in which case the flag indicating whether the proposed method is applicable may not be signaled.
[0200] The applicability of the proposed method according to the present disclosure may be implicitly determined depending on whether certain conditions are satisfied.
[0201] Information regarding the applicability (or availability) of the proposed method according to the present disclosure may be defined in HLS. Based on this information, information regarding the applicability (or availability) of the proposed method may be adaptively signaled at the coding unit level. For example, if the information regarding the applicability (or availability) of the proposed method in the SPS is false, the proposed method may be determined not to be applied at the coding unit level, and signaling of the information regarding the applicability of the proposed method at the coding unit level may be omitted.
[0202] FIG. 5 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.
[0203] Referring to FIG. 5, the decoding device (300) may include a DIPM derivation unit (500), a prediction block generation unit (510), and a restoration unit (520). The DIPM derivation unit (500) and the prediction block generation unit (510) may be provided in the intra prediction unit (331) of FIG. 3.
[0204] The DIPM induction unit (500) can perform the DIPM induction process according to S400. The prediction block generation unit (510) can perform the prediction block generation process according to S410. The restoration unit (520) can perform the current block restoration process according to S420.
[0205] FIG. 6 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0206] One or more DIPMs can be derived for the current block based on a given reference area (S600). The DIPM derivation method is as described with reference to FIG. 4.
[0207] A prediction block of the current block can be generated based on the DIPM(s) derived from S600 (S610). The method for generating the prediction block is as described with reference to FIG. 4.
[0208] Transform coefficients of the current block can be derived based on the residual block of the current block (S620). The residual block of the current block can be generated based on the prediction block generated in S610. The transform coefficients can be derived by performing at least one of transformation or quantization on the residual block.
[0209] A bitstream can be generated by encoding residual information about the transform coefficients of the current block (S630).
[0210] FIG. 7 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.
[0211] Referring to FIG. 7, the encoding device (200) may include a DIPM derivation unit (700), a prediction block generation unit (710), a transform coefficient derivation unit (720), and a residual information encoding unit (730).
[0212] The DIPM derivation unit (700) and the prediction block generation unit (710) may be provided in the intra prediction unit (222) of FIG. 2. The transform coefficient derivation unit (720) may be provided in the residual processing unit (230) of FIG. 2. The residual information encoding unit (730) may be provided in the entropy encoding unit (240).
[0213] The DIPM derivation unit (700) can perform the DIPM derivation process according to S600. The prediction block generation unit (710) can perform the prediction block generation process according to S610. The transform coefficient derivation unit (720) can perform the transform coefficient derivation process according to S620. The residual information encoding unit (730) can perform the residual information encoding process according to S630.
[0214] In the embodiments described above, the methods are described based on a flowchart as a series of steps or blocks. However, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the embodiments of this document.
[0215] The method according to the embodiments of the present document described above can be implemented in the form of software, and the encoding device and / or decoding device according to the present document can be included in a device that performs image processing, such as a TV, a computer, a smartphone, a set-top box, a display device, etc.
[0216] When the embodiments in this document are implemented as software, the above-described method can be implemented as a module (process, function, etc.) that performs the above-described function. The module can be stored in memory and executed by a processor. The memory can be internal or external to the processor and can be connected to the processor by various well-known means. The processor can include an application-specific integrated circuit (ASIC), another chipset, logic circuit, and / or data processing device. The memory can include a read-only memory (ROM), a random access memory (RAM), flash memory, a memory card, a storage medium, and / or other storage devices. That is, the embodiments described in this document can be implemented and performed on a processor, a microprocessor, a controller, or a chip. For example, the functional units illustrated in each drawing can be implemented and performed on a computer, a processor, a microprocessor, a controller, or a chip. In this case, information for implementation (e.g., information on instructions) or an algorithm can be stored on a digital storage medium.
[0217] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in a multimedia broadcasting transmitting and receiving device, a mobile communication terminal, a home cinema video device, a digital cinema video device, a surveillance camera, a video conversation device, a real-time communication device such as a video communication, a mobile streaming device, a storage medium, a camcorder, a video-on-demand (VoD) service providing device, an OTT (Over the top video) device, an Internet streaming service providing device, a three-dimensional (3D) video device, a VR (virtual reality) device, an AR (argumente reality) device, a video phone video device, a transportation terminal (ex. a vehicle (including an autonomous vehicle) terminal, an airplane terminal, a ship terminal, etc.), and a medical video device, and may be used to process a video signal or a data signal. For example, the OTT (Over the top video) device may include a game console, a Blu-ray player, an Internet-connected TV, a home theater system, a smartphone, a tablet PC, a DVR (Digital Video Recorder), etc.
[0218] In addition, the processing method to which the embodiment(s) of the present specification are applied can be produced in the form of a computer-executable program and can be stored in a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of the present specification can also be stored in a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium can include, for example, a Blu-ray disc (BD), a universal serial bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. In addition, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission via the Internet). In addition, a bitstream generated by an encoding method can be stored in a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0219] Additionally, the embodiments of the present disclosure may be implemented as a computer program product by program code, and the program code may be executed on a computer by the embodiments of the present disclosure. The program code may be stored on a computer-readable carrier.
[0220] FIG. 8 illustrates an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0221] Referring to FIG. 8, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0222] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and camcorders into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and camcorders directly generate bitstreams, the encoding server may be omitted.
[0223] The above bitstream can be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0224] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, in which case the control server controls commands / responses between each device within the content streaming system.
[0225] The streaming server can receive content from a media repository and / or an encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0226] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.
[0227] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0228] The claims set forth in this specification may be combined in various ways. For example, the technical features of the method claims of this specification may be combined and implemented as a device, and the technical features of the device claims of this specification may be combined and implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a device, and the technical features of the method claims and the technical features of the device claims of this specification may be combined and implemented as a method.
Claims
1. A step of deriving one or more DIPMs (derived intra prediction modes) for a current block based on a predetermined reference area; A step of generating a prediction block of the current block based on one or more of the DIPMs; and Including a step of restoring the current block based on the above predicted block, A method wherein the above reference area is divided into multiple sub-areas.
2. In paragraph 1, A method wherein one or more of the above DIPMs are derived by applying a filter to samples of a portion of a sub-region.
3. In paragraph 2, A method wherein some areas within the plurality of sub-areas to which the filter is applied contain the same number of samples.
4. In paragraph 2, A method wherein some areas within the plurality of sub-areas to which the filter is applied contain different numbers of samples.
5. In paragraph 1, A method wherein the above prediction block is generated based on one or more DIPMs and samples belonging to a predetermined reference sample line.
6. In paragraph 5, A method wherein the above-described reference sample line is determined based on a sub-region used to derive one or more DIPMs.
7. In paragraph 5, A method in which the above-described reference sample line is set as a reference sample line at a pre-defined position.
8. In paragraph 5, A method wherein the above-described reference sample line is determined based on whether the DIPM derived for the current block is in planar mode.
9. In paragraph 1, A method wherein at least one of the above one or more DIPMs is added as an MPM candidate to the MPM list of the current block.
10. In paragraph 9, The above MPM list includes the intra prediction mode of the surrounding blocks of the current block as the MPM candidates, A method wherein at least one of the one or more DIPMs derived for the surrounding block is set as the intra prediction mode of the surrounding block.
11. In paragraph 1, A primary mode of one or more DIPMs derived for the current block is stored in the current block, A method wherein the first mode is a mode having the largest amplitude value among the one or more DIPMs.
12. A step of deriving one or more DIPMs (derived intra prediction modes) for the current block based on a predetermined reference area; A step of generating a prediction block of the current block based on one or more of the DIPMs; A step of generating a residual block of the current block based on a prediction block of the current block; A step of deriving transform coefficients of the current block based on the residual block; and A step of encoding residual information regarding the above transformation coefficients, A method wherein the above reference area is divided into multiple sub-areas.
13. A computer-readable storage medium storing a bitstream generated by the method according to Article 12.
14. A step of obtaining a bitstream for image information; wherein the bitstream is generated based on a step of deriving one or more DIPMs (derived intra prediction modes) for a current block based on a predetermined reference region, a step of generating a prediction block of the current block based on the one or more DIPMs, a step of generating a residual block of the current block based on the prediction block of the current block, a step of deriving transform coefficients of the current block based on the residual block, and a step of encoding residual information about the transform coefficients, and Including a step of transmitting data including the above bitstream, A method wherein the above reference area is divided into multiple sub-areas.
Citation Information
Patent Citations
Catalyst and method for manufacturing the same
KR1020240140710A
Oil water separator for small vessels
KR1020250031312A
Method and device for exchanging secret keys based on reconfigurable and unclonable cryptographic component
KR1020250052002A
System and method for preventing unauthorized disclosure of secure printed material
KR1020250070275A
Flame retardant composites for batteries and battery cell assemblies
KR1020250131661A