Method, apparatus, and recording medium for image encoding / decoding
The method addresses the challenges of video encoding/decoding by employing filtering techniques for block classification and filtering, resulting in improved data efficiency and video quality.
Patent Information
- Application Number
- PCT/KR2024/018110
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-15
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-22
AI Technical Summary
Current video encoding/decoding technologies face challenges in efficiently compressing and transmitting high-resolution videos, requiring innovative methods to improve data efficiency and quality.
The proposed method involves a device, method, and recording medium that utilize filtering techniques for video encoding/decoding. Specifically, it includes steps such as block classification, assigning a block classification index, and applying a corresponding filter for filtering based on statistical values of transformation coefficients or pixel values of surrounding pixels.
This approach enhances the efficiency of video encoding/decoding by optimizing data compression and transmission, thereby improving video quality and meeting user demands for higher resolution.
Smart Images

Figure KR2024018110_22052025_PF_FP_ABST
Abstract
Description
Method, device and recording medium for video encoding / decoding
[0001] The present invention relates to a method, device, and recording medium for image encoding / decoding. Specifically, the present invention discloses a method, device, and recording medium for image encoding / decoding using filtering.
[0002] This invention claims the benefit of Korean Patent Application No. 10-2023-0158588, filed November 15, 2023, and Korean Patent Application No. 10-2024-0163407, filed November 15, 2024, the entire contents of which are incorporated herein by reference.
[0003] With the continuous development of the information and communication industry, services providing video through broadcasting and the Internet have spread worldwide.
[0004] Users demand higher resolution and higher quality video. To meet these demands, video encoding / decoding technologies tailored to these needs are required. Video encoding technology can create compressed video by compressing the video representing the images into a smaller amount of data. Video decoding technology can use the compressed video to create reconstructed images.
[0005] When it comes to video encoding / decoding, various technologies exist, including segmentation, prediction, transformation, quantization, filtering, and entropy encoding / decoding. By introducing, modifying, improving, and combining these diverse technologies, video and images can be compressed, transmitted, and stored more effectively.
[0006] One embodiment may provide a device, method and recording medium using filtering.
[0007] A decryption method is provided, comprising: a step of performing classification on a block on one side; and a step of performing filtering on the block based on the classification.
[0008] A block classification index can be assigned to the block by the above classification.
[0009] The filtering can be performed using a filter corresponding to the above block classification index.
[0010] The above block classification index can be determined based on statistical values of coefficients derived by transformation for the above block.
[0011] The above statistical value can be the sum of the values of the coefficients or the sum of the absolute values of the coefficients.
[0012] The above block classification index can be determined based on pixel values of surrounding pixels of the block.
[0013] Among multiple available filters, the filter corresponding to the block classification index can be selected.
[0014] A geometric transformation can be performed on the filter coefficients of the above filter.
[0015] On the other hand, an encoding method is provided, comprising: a step of performing classification on a block; and a step of performing filtering on the block based on the classification.
[0016] A block classification index can be assigned to the block by the above classification.
[0017] The filtering can be performed using a filter corresponding to the above block classification index.
[0018] The above block classification index can be determined based on statistical values of coefficients derived by transformation for the above block.
[0019] The above block classification index can be determined based on pixel values of surrounding pixels of the block.
[0020] Among multiple available filters, the filter corresponding to the block classification index can be selected.
[0021] A geometric transformation can be performed on the filter coefficients of the above filter.
[0022] In another aspect, a computer-readable recording medium for storing a bitstream generated by the above encoding method may be provided.
[0023] In another aspect, a computer-readable recording medium for storing a bitstream for image decoding may be provided, wherein the bitstream includes filter information, classification is performed on a block based on the filter information, and filtering is performed on the block based on the classification.
[0024] A block classification index can be assigned to the block by the above classification.
[0025] The filtering can be performed using a filter corresponding to the above block classification index.
[0026] The above block classification index can be determined based on statistical values of coefficients derived by transformation for the above block.
[0027] The above block classification index can be determined based on pixel values of surrounding pixels of the block.
[0028] Among multiple available filters, the filter corresponding to the block classification index can be selected.
[0029] A geometric transformation can be performed on the filter coefficients of the above filter.
[0030] Devices, methods and recording media using filtering are provided.
[0031] Figure 1 illustrates a system for video coding according to one embodiment.
[0032] Figure 2 shows a segmentation structure of an image according to one embodiment.
[0033] Figure 3 illustrates the structure of intra prediction according to one embodiment.
[0034] Figure 4 shows the structure of inter prediction to explain the inter prediction process according to one embodiment.
[0035] Figure 5 shows the order in which spatial candidates are added to the candidate list according to one embodiment.
[0036] Figure 6 illustrates multiple in-loop filters according to an example.
[0037] Figure 7 shows the structure of entropy encoding and entropy decoding according to an example.
[0038] Figure 8 is a flowchart of an encoding method according to one embodiment.
[0039] Figure 9 is a flowchart of a decryption method according to one embodiment.
[0040] Figure 10 can represent DCT8 transform coefficients according to an example.
[0041] Figure 11 can represent Hadamard transform coefficients according to an example.
[0042] Figure 12 illustrates a 2x2 2D transform for 2x2 block classification according to an example.
[0043] Figure 13 shows coefficients derived as a result of a 2x2 2D transformation according to an example.
[0044] Figure 14 shows a 4x4 transform for 2x2 block classification according to an example.
[0045] Figure 15 shows the 4x4 2D transformation result coefficients according to an example.
[0046] Figure 16 shows an 8x8 2D transform for 2x2 block classification according to an example.
[0047] Figure 17 shows the 8x8 2D transformation result coefficients according to an example.
[0048] Figure 18 illustrates a 4x4 2D transform for 4x4 block classification according to an example.
[0049] Figure 19 shows an 8x8 2D transform for 4x4 block classification according to an example.
[0050] Figure 20 illustrates 4x4 2D transform 2x2 windows according to an example.
[0051] Figure 21 illustrates specific parts and sums of specific parts of a 4x4 2D transform according to an example.
[0052] Figure 22 shows the excluded region of a 4x4 2D transformation according to an example.
[0053] Figure 23 illustrates 8x8 2D transform 2x2 windows according to an example.
[0054] Figure 24 shows the sum of specific parts of specific parts of an 8x8 2D transform according to an example.
[0055] Figure 25 shows the excluded region of an 8x8 2D transformation according to an example.
[0056] Figure 26 shows a case where the range in which a sample is calculated according to an example has a square shape.
[0057] Figure 27 shows a case where the range in which samples are calculated according to an example has a rectangular shape.
[0058] Figure 28 shows a case where the range in which a sample is calculated according to an example has a square shape.
[0059] Figure 29 shows a case where the range in which samples are calculated according to an example has a rectangular shape.
[0060] Figure 30 illustrates extracting samples in a vertical direction from surrounding samples according to an example.
[0061] Figure 31 illustrates extracting samples in a horizontal direction from surrounding samples according to an example.
[0062] Figure 32 illustrates extracting samples in the form of a 45 degree diagonal direction from surrounding samples according to an example.
[0063] Figure 33 illustrates extracting samples in the form of a -45 degree diagonal direction from surrounding samples according to an example.
[0064] Figure 34 shows a first vertical sample extraction form in 2x2 block classification according to an example.
[0065] Figure 35 shows a second vertical sample extraction form in 2x2 block classification according to an example.
[0066] Figure 36 shows a third vertical sample extraction form in 2x2 block classification according to an example.
[0067] Figure 37 shows a fourth vertical sample extraction form in 2x2 block classification according to an example.
[0068] Figure 38 shows a fifth vertical sample extraction form in 2x2 block classification according to an example.
[0069] Figure 39 shows a first horizontal sample extraction form in 2x2 block classification according to an example.
[0070] Figure 40 shows a second horizontal sample extraction form in 2x2 block classification according to an example.
[0071] Figure 41 shows a third horizontal sample extraction form in 2x2 block classification according to an example.
[0072] Figure 42 shows a fourth horizontal sample extraction form in 2x2 block classification according to an example.
[0073] Figure 43 shows a fifth horizontal sample extraction form in 2x2 block classification according to an example.
[0074] Figure 44 shows the first -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[0075] Figure 45 shows the second -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[0076] Figure 46 shows the third -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[0077] Figure 47 shows the fourth -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[0078] Figure 48 shows the fifth -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[0079] Figure 49 shows the first 45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[0080] Figure 50 shows the second 45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[0081] Figure 51 shows the third 45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[0082] Figure 52 shows the fourth 45-degree diagonal sample extraction form in 2x2 block classification according to an example.
[0083] Figure 53 shows the fifth 45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[0084] Figure 54 shows a 9x9 diamond-shaped filter according to an example.
[0085] Figure 55 shows a 7x7 diamond-shaped filter according to an example.
[0086] Figure 56 shows a 5x5 diamond-shaped filter according to an example.
[0087] Figure 57 shows an 11x11 cross-shaped filter according to an example.
[0088] Figure 58 shows a 9x9 cross-shaped filter according to an example.
[0089] Figure 59 shows a 7x7 cross-shaped filter according to an example.
[0090] Figure 60 shows an 11x11 diagonally symmetric cross-shaped filter according to an example.
[0091] Figure 61 shows a 9x9 diagonally symmetric cross-shaped filter according to an example.
[0092] Figure 62 shows a 7x7 diagonally symmetric cross-shaped filter according to an example.
[0093] FIG. 63 may represent a 7x7 diagonally symmetric cross-shaped filter according to an example.
[0094] Fig. 64 may represent a 7x7 diagonally symmetric rhombus-shaped filter according to an example.
[0095] FIG. 65 may represent a 5x5 diagonally symmetric square shaped filter according to an example.
[0096] Figure 66 may represent a 7x7 diagonally symmetric plus shape filter according to an example.
[0097] Figure 67 may represent a 7x7 diagonally symmetric thin cross-shaped filter according to an example.
[0098] Figure 68 may represent a 7x7 diagonally symmetric octagonal filter according to an example.
[0099] Figure 69 shows a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[0100] Figure 70 illustrates a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[0101] Figure 71 shows a 5x5 diagonally symmetric rectangular shape filter according to an example.
[0102] Figure 72 shows a 7x7 diagonally symmetric thick rectangular shape filter according to an example.
[0103] Figure 73 illustrates a case where no transformation is performed for a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[0104] Figure 74 illustrates a case where no transformation is performed for a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[0105] Figure 75 illustrates a case where no transformation is performed for a 5x5 diagonally symmetric rectangular shape filter according to an example.
[0106] Figure 76 illustrates a case where a 90 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[0107] Figure 77 illustrates a case where a 90 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[0108] Figure 78 shows a case where a 90 degree rotation transformation is performed on a 5x5 diagonally symmetric rectangular shape filter according to an example.
[0109] Figure 79 illustrates a case where a 180 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[0110] FIG. 80 illustrates a case where a 180-degree rotation transformation is performed on a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[0111] Figure 81 shows a case where a 180-degree rotation transformation is performed on a 5x5 diagonally symmetric rectangular shape filter according to an example.
[0112] Figure 82 illustrates a case where a 270 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[0113] Figure 83 shows a case where a 270 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[0114] Figure 84 shows a case where a 270 degree rotation transformation is performed on a 5x5 diagonally symmetric rectangular shape filter according to an example.
[0115] Figure 85 shows a 1x1 filter shape and a 3x3 filter shape applied to prediction / residual samples of an adaptive loop filter according to an example.
[0116] Fig. 86 shows filter shapes of an online derived adaptive loop filter according to an example.
[0117] Fig. 87 shows different filter forms of an online derived adaptive loop filter according to an example.
[0118] Figure 88 illustrates filter shapes within an adaptive loop filter according to an example.
[0119] Figure 89 shows an added filter form for an adaptive loop filter according to an example.
[0120] Figure 90 shows a diamond 9x9 filter shape of a chroma adaptive loop filter according to an example.
[0121] Figure 91 shows a cross 9x9 filter form of a component-by-component adaptive loop filter according to an example.
[0122] Figure 92 illustrates a diamond 9x9 chroma filter shape and luma residual-based taps introduced into a chroma adaptive loop filter according to an example.
[0123] Figure 93 shows a cross 9x9 component-wise adaptive loop filter shape and a luma residual-based tap introduced into the component-wise adaptive loop filter according to an example.
[0124] Figure 94 shows the architecture of a low complexity operating point model according to an example.
[0125] Figure 95 illustrates processing of a backbone block in a low complexity operating point model according to an example.
[0126] Figure 96 shows parameters for a head block according to an example.
[0127] Figure 97 shows parameters for a backbone according to an example.
[0128] Figure 98 shows the dimensions for a reconstruction input according to an example.
[0129] Figure 99 shows the alternative and input dimensions of DCT-II according to an example.
[0130] Figure 100 illustrates an input-transformed model and dimensionality change through the above model according to an example.
[0131] Figure 101 illustrates an input-transformed model according to another example and dimensionality change through the above model.
[0132] Figure 102 illustrates processing of a backbone block in an input-transformed model according to another example.
[0133] Figure 103 shows a model in which fixed transformations are replaced with trainable components along with transformations, according to an example.
[0134] Figure 104 illustrates processing of a backbone block in a model with replacement applied according to an example.
[0135] Figure 105 shows parameters for the backbone in a model with replacement applied according to an example.
[0136] Figure 106 illustrates a trainable transformation for input adjustment according to an example.
[0137] Figure 107 shows a parallel synthesis of the output of a neural network loop filter and the output of a deblocking filter according to an example.
[0138] The present invention is capable of various modifications. Furthermore, the present invention may have various embodiments. Specific embodiments are described in the accompanying drawings and detailed description.
[0139] It should be understood that the specific examples are not intended to limit the invention to specific embodiments, and that all modifications, equivalents, and substitutes falling within the spirit and scope of the invention are intended to be encompassed within the scope of the invention as embodiments.
[0140] The embodiments are described in sufficient detail to enable those skilled in the art to practice them. It should be understood that the various embodiments, while different from each other, are not necessarily mutually exclusive. For example, it should be understood that the shapes, structures, and characteristics described in connection with one embodiment may be applied to or implemented in other embodiments without departing from the spirit and scope of the present invention. Furthermore, it should be understood that the positions or arrangements of components within one embodiment may be modified without departing from the spirit and scope of the present invention. Accordingly, the following detailed description is not intended to be limiting, and the scope of the exemplary embodiments, if properly described, is defined only by the appended claims and all equivalents to those claimed by such claims.
[0141] A detailed description of the embodiments described below may refer to the drawings for the embodiments. Any description described in the drawings or the descriptions shown in the drawings may be considered part of the detailed description. In the drawings, similar reference numerals may designate the same or similar functions throughout various aspects. The dependencies between components may not be limited to those depicted in the drawings.
[0142] In the embodiments, a singular expression may include, and may be limited to, and / or restricted by, a plural expression, unless the context clearly excludes a plural expression. That is, expressions such as “at least one” and “one or more” in the embodiments may be replaced with “plural.” Terms such as “ / ,” “and / or,” “at least one of,” and “one or more of” described for a plurality of items may mean 1) one item of the plurality of items, 2) some of the plurality of items, 3) a combination of some of the plurality of items, or 4) a combination of the plurality of items. Furthermore, a plural expression may be replaced with a singular expression. The plural may mean an integer greater than or equal to 1, 2, 3, 4, or 5.
[0143] In the embodiments, terms related to numbers, such as "first" and "second," may be used to describe various components. These terms are used only to distinguish one component from another and do not limit the components. For example, without departing from the scope of the present invention, the first component could be referred to as the second component, and similarly, the second component could also be referred to as the first component.
[0144] When a first component transmits (or provides) information to a second component, it can mean that the first component directly transmits information to the second component, or it can mean that the first component transmits information to the second component via another third component. Here, the information that the second component receives (or obtains) can be information transmitted by the first component, or information generated by applying a specific process to information transmitted by the first component.
[0145] The components of the embodiments may be depicted independently to represent different characteristic functions, and this does not imply that each component corresponds to a separate hardware or software configuration unit. That is, the components of the embodiments may be distinguished and listed for convenience of description. Two or more components described in the embodiments may be regarded as a single component. Furthermore, a single component described in the embodiments may be separated into multiple components that perform the functions of the aforementioned component. Embodiments in which such components are integrated and embodiments in which such components are separated are also included in the scope of the present invention, as long as they do not depart from the essence of the present invention.
[0146] The terms used in the embodiments are used only to describe specific embodiments and are not intended to limit the present invention. In the embodiments, terms such as "comprise" or "have" indicate the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the embodiments. These terms do not preclude the presence or addition of other features, numbers, steps, operations, components, parts, or combinations thereof that are not explicitly described in the embodiments. In other words, the description of "comprising" a specific component of an embodiment does not exclude other components other than the specific component, and means that additional components may also be included in the scope of the embodiments of the present invention or the technical idea of the present invention.
[0147] Some of the components of the embodiments may be optional, not essential components for performing the essential functions of the present invention. These optional components may be used to improve performance. The embodiments may be implemented as a structure that includes only the essential components required to implement the essence of the embodiments, excluding the optional components. Such a structure is also within the scope of the embodiments.
[0148] Hereinafter, embodiments are described in detail with reference to the attached drawings to enable those skilled in the art to easily implement the embodiments. In describing the embodiments, if a detailed description of a related known configuration or function is judged to obscure the gist of the present specification, such detailed description will be omitted. Furthermore, identical reference numerals are used for identical components in the drawings, and redundant descriptions of identical components will be omitted.
[0149]
[0150] Interchange between terms in the examples
[0151] Below, terms listed on a single line may be used with the same meaning in the embodiments and may be used interchangeably in the embodiments.
[0152] - 'one or more', 'at least one'
[0153] - 'two or more', 'a plurality of', 'multiple', 'multiple'. (In embodiments, 'one or more' or 'at least one' may be further limited to 'two or more', 'plural', or 'multiple'.)
[0154] - 'Information', 'Signal'
[0155] - 'value', 'predefined value', 'specific value', 'threshold', 'threshold value', 'baseline value', 'reference value'
[0156] - 'statistical value', 'statistics value'
[0157] - 'indicator', 'index', 'index', 'flag', 'information'
[0158] - 'encoder', 'encoding apparatus'
[0159] - 'decoder', 'decoding apparatus'
[0160] - 'Entropy encoding', 'encoding', 'encoding'
[0161] - 'Entropy decryption', 'decoding', 'decoding'
[0162] - 'coding', 'encoding and / or decoding'
[0163] - 'video', 'moving picture', 'image', 'picture', 'frame', 'screen'
[0164] - 'Reference picture', 'Reference video'
[0165] - 'Reference Picture List (RPL),' 'Reference Image List'
[0166] - 'original', 'input', 'source'
[0167] - 'Block', 'Unit', 'Signal'
[0168] - 'square', 'square shape'
[0169] - 'pixel', 'pixels', 'samples', 'pels'
[0170] - 'region', 'area', 'part', 'segment'
[0171] - 'partition', 'split', 'divide'
[0172] - 'quad', 'quarternary'
[0173] - 'luma component', 'luma', 'luminance component', 'luminance', 'Y'
[0174] - 'chroma component', 'chroma', 'chrominance', 'chrominance component', 'Cb and Cr', 'Cb or Cr', 'Cb', 'Cr', 'U and V', 'U or V', 'U', 'V'
[0175] - 'target', 'current' (e.g. target block and current block, or target image and current image)
[0176] - 'neighbor', 'neighboring', 'adjacent', 'neighbor / neighboring' (e.g., neighboring block, adjacent block, and surrounding block)
[0177] - 'collocated', 'collected'
[0178] - 'reconstruction', 'reconstruction', 'decoding'
[0179] - 'reconstructed', 'reconstructed', 'decoded'
[0180] - 'difference', 'difference', 'difference', 'error', 'residual', 'residual'
[0181] - Largest Coding Unit (LCU), Coding Tree Unit (CTU)
[0182] - 'inter', 'inter-screen'
[0183] - 'Inter prediction', 'inter prediction', 'motion compensation'
[0184] - 'Inter mode', 'Inter prediction mode', 'Inter-screen mode', 'Inter-screen prediction mode'
[0185] - 'Motion vector', 'Predicted motion vector', 'Advanced Motion Vector Prediction (AMVP)'
[0186] - 'list', 'candidate list'
[0187] - 'Spatial candidate', 'Spatial merge candidate'
[0188] - 'Temporal candidate', 'Temporal merge candidate'
[0189] - 'Prediction motion vector candidate', 'motion vector predictor'
[0190] - 'Prediction method', 'Prediction mode'
[0191] - 'Intra', 'Intra'
[0192] - 'Intra prediction', 'Intra prediction'
[0193] - 'Intra mode', 'Intra prediction mode'
[0194] - 'Dequantization', 'scaling'
[0195] - 'Quantization matrix', 'Scaling list'
[0196] - 'Quantization matrix coefficients', 'matrix coefficients'
[0197] - 'Transform coefficient level', 'quantized level', 'quantized coefficient', 'quantized transform coefficient', 'quantized transform coefficient level'
[0198] - 'Dequantized coefficient', 'dequantized transform coefficient'
[0199] - 'Scanning type', 'Scanning direction'
[0200] - 'Directional mode', 'Angle mode', 'Angular mode', 'Intra prediction mode'
[0201] - '(mode) number of intra prediction mode', '(mode) index of intra prediction mode', '(mode) value of intra prediction mode', '(mode) angle of intra prediction mode', '(mode) direction of intra prediction mode', '(mode) number of intra prediction direction', '(mode) index of intra prediction direction', '(mode) value of intra prediction direction', '(mode) angle of intra prediction direction'
[0202] - 'Merge Mode', 'Movement Merge Mode'
[0203] - 'Geometric Partitioning Mode (GPM)', 'Triangle Partitioning Mode'
[0204] In addition to the terms exemplified above, terms having the same meaning according to common knowledge in the technical field may be used interchangeably in the embodiments.
[0205]
[0206] The range of information and values of information described in the examples
[0207] In embodiments, information may include constants, flags, indices, variables, coding parameters, elements, syntax elements, motion information, attributes, entities, objects, and data. That is, the term 'information' may be replaced with 'data', 'flag', 'index', 'variable', 'element', 'syntax element', 'motion information', 'attribute', or 'object'.
[0208] Information can have one of multiple values. The 'nth value' can mean the nth value among the multiple values.
[0209] For example, the first value could represent '0' or (logical) false. The second value could represent '1' or (logical) true. Alternatively, the first value could represent '1' or (logical) true. The second value could represent '0' or (logical) false.
[0210] A flag may be information having a value of either '0' or '1'. In embodiments, the flag values '0' and '1' may be replaced with '1' and '0', respectively. For example, information indicating whether a specific process is performed or whether a specific process is applied may be considered a flag.
[0211] When a variable such as i or j is used to represent a row, column, or index, the variable can be an integer greater than or equal to 0 and less than or equal to n - 1. Alternatively, the variable can be an integer greater than or equal to 1 and less than or equal to n. Here, n can be the number of rows, the number of columns, or the number of entities pointed to by the index.
[0212]
[0213] Coding related concepts
[0214] Below, concepts related to coding are described. The descriptions disclosed below can be applied to embodiments.
[0215] Predefined value: A predefined value may refer to a value commonly used in an encoding device and a decoding device. For example, a predefined value may be interpreted as a fixed value. Alternatively, the predefined value may be a value shared by an encoding device and a decoding device through signaling. Alternatively, the predefined value may be a value derived through the same procedure in an encoding device and a decoding device so that the encoding device and the decoding device have a common value. Alternatively, the predefined value may be a common value in an encoding device and a decoding device. The above description of a predefined value may also be applied to predefined information. In the above descriptions, 'value' may be replaced with 'information'.
[0216] Availability: The availability of certain modes for a specific target may mean that a selected mode among the specific modes is used for the specific target. Other modes within the specific mode category may be unavailable. Unavailable modes may not be used for the specific target. The description of a specific mode above may also apply to other specific information. In the descriptions above, "mode" may be replaced with "information."
[0217] Adjacency: The 'direction' of the 'second entity' with respect to the 'first entity' may refer to the 'second entity' that is adjacent to the 'direction' corner / face of the first entity. For example, the 'top left block' with respect to the 'target block' may be a block adjacent to the top left of the target block. Here, the 'first entity' may be a target unit, a target block, or a target sample. The 'direction' may be one of left-above, above, right-above, left, right, left-below, below, and right-below. The 'second entity' may be a unit, a block, or a sample. For the directions of left-top, right-top, left-bottom, and right-bottom, the corners of the first entity and the corners of the second entity may be diagonally adjacent. For the directions of top, left, right and bottom, one side of the first object and one side of the second object can be in contact with each other.
[0218] - For example, the block adjacent to the upper left of the target block may be the block adjacent to the upper left of the block adjacent to the target block. The block adjacent to the upper right of the target block may be the block adjacent to the right of the block adjacent to the upper right of the target block. The block adjacent to the lower left of the target block may be the block adjacent to the lower left of the block adjacent to the target block.
[0219] Coding: Coding can mean encoding and / or decoding of images.
[0220] Signal: A signal can represent information about an image, unit, or block. A specific signal can represent a specific image, a specific unit, or a specific block.
[0221] Video: A video can refer to a single picture that constitutes a video, or it can refer to the video itself. For example, "encoding and / or decoding a video" can mean "encoding and / or decoding a video," or it can mean "encoding and / or decoding one of the pictures that constitute the video."
[0222] - A picture can mean the entire picture, or it can mean a part of a picture, such as a block.
[0223] Target Image: The target image may be an encoding target image, which is the target of encoding, and / or a decoding target image, which is the target of decoding. Furthermore, the target image may be an input image processed by an encoding device, or a restored image processed by a decoding device. The target image may be an image including a target block.
[0224] Subpicture: A picture can be divided into one or more subpictures.
[0225] - A subpicture may be a square or rectangular area within a picture. A subpicture may contain one or more CTUs.
[0226] - A subpicture may include one or more slices and / or one or more tiles. For example, a subpicture may consist of one or more slice rows and one or more slice columns. Alternatively, each subpicture may consist of one or more tile rows and one or more tile columns.
[0227] - A subpicture may include one or more slices that collectively cover a rectangular area within the picture. Accordingly, the boundary of each subpicture may always be the boundary of a slice. Additionally, each vertical subpicture boundary may always be a vertical tile boundary.
[0228] Slice: A slice may contain one or more tiles within a picture. A slice may consist of one or more rows of tiles and one or more columns of tiles.
[0229] Tile: A tile can be a square or rectangular area within a picture. A tile can contain one or more CTUs. A picture can be divided into one or more tile rows and one or more tile columns.
[0230] CTU: An image can be divided into multiple coding tree units (CTUs).
[0231] - A CTU may include one Y coding tree block (CTB) and at least one of a Cb CTB and a Cr CTB related to the Y CTB, and may include information about each CTB. The information may include syntax elements.
[0232] - Each CTU can be partitioned using one or more partitioning methods to form sub-units such as coding units (CUs), prediction units (PUs), and transform units (TUs). The one or more partitioning methods can include quad tree (QT) partitioning, binary tree (BT) partitioning, and ternary tree (TT) partitioning. Additionally, each CTU can be partitioned using multi-type tree (MTT) partitioning that uses a combination of multiple partitioning methods.
[0233] CTB: CTB can refer to one of Y CTB, Cb CTB, and Cr CTB.
[0234] Unit: A unit can be determined for specific processing in coding. A unit can contain information about a specific region within an image. For specific coding processing, an image can be recursively divided into multiple parts. A unit can represent the region to which a specific processing is applied and information about the region.
[0235] - The type of a unit may indicate a specific processing to be applied to the unit. Depending on the type of the unit, a specific processing may be applied to the unit. A 'specific' unit may be a unit for processing designated as 'specific' in coding. For example, the unit may be at least one of an original unit, a CTU, a coding unit, a prediction unit, a residual unit, a reconstructed residual unit, a transformation unit, and a reconstructed unit.
[0236] - A unit may include samples having a two-dimensional shape or arrangement. In this respect, a 'unit' may also mean a 'block'. For example, a block may be at least one of an original block, a CTB, a coding block (CB), a prediction block (PB), a residual block, a reconstructed residual block, a transform block (TB), and a reconstructed block. For example, a division of a unit may mean a division of a block corresponding to the unit.
[0237] - A unit can contain syntax elements. In other words, a block and its syntax elements can be combined to form a unit.
[0238] - A block is an MxN array of samples. Here, M and N can represent positive integer values, and a block can commonly represent a two-dimensional sample array. The current block can represent an encoding target block that is the target of encoding during encoding, and a decoding target block that is the target of decoding during decoding. In addition, the current block can be at least one of a coding block, a prediction block, a residual block, a transform block, and a restoration block. The block can have various sizes and shapes. For example, the shape of the block can be one or more of a tetragon, a rectangular block, a square block, a rectangle whose width is different from its height (that is, an oblong block), a trapezoid, a triangle, a right-angled triangle, and a pentagon. Here, the width and height of the rectangle can be different from each other. In addition, the shape of the block can include other geometric shapes that can be expressed in two dimensions. For example, the shape of a block may be a quadrilateral or a pentagon, which is defined by excluding the area of a right triangle from the area of a rectangle. Here, the right vertex of the right triangle may be one of the vertices of the rectangle. Furthermore, the shape of a block may be a combination of two or more of the aforementioned shapes. Furthermore, the shape of a block may be the remainder of one of the aforementioned shapes after excluding another shape.
[0239] - In embodiments, a rectangle may be limited to a non-square rectangle. When the shape of a particular object is described as a rectangle in an embodiment, such description may additionally imply that the width and height of the particular object are different from each other.
[0240] - In embodiments, a block may be limited to at least one of a vertically oriented block and a horizontally oriented block. A vertically oriented block may mean a block whose vertical length is greater than its horizontal length. A horizontally oriented block may mean a block whose horizontal length is greater than its vertical length.
[0241] - A unit may include a luma component block (i.e., a Y block) and two chroma component blocks (i.e., at least one of a Cb block and a Cr block), and may include information about each block. The information may include syntax elements.
[0242] - Unit information may include unit type, unit size, unit depth, unit encoding order, and unit decoding order.
[0243] Target Unit: A target unit may be a block, an encoding target unit, which is a target of encoding, and / or a decoding target unit, which is a target of decoding. A target unit may be a specific area within a target picture to which one or more specific coding processes are applied. A unit of a specific type may be generated by applying a specific process to a target unit. Alternatively, a target unit may represent a unit having a specific type for a specific coding process.
[0244] Depth: A block can be hierarchically divided into multiple sub-blocks, each with its own depth, according to a tree structure. The multiple sub-blocks created by block division can be called partitions.
[0245] - The depth of a block can indicate the level of the node corresponding to the block when the blocks that make up the image are expressed in a tree structure. Alternatively, the depth of a block can indicate the number of partitions applied until the block is determined. The depth of a block can increase by 1 as the block is further partitioned.
[0246] - In a tree structure, the root node can be considered to have the smallest level, and the leaf node can be considered to have the largest level. The root node can be the topmost node in the tree structure and corresponds to the first undivided block. The level of the root node can be 0 or 1. When the level of the root node is 0, a node with a level of 1 can represent a block determined by splitting the first block once. A node with a level of n can represent a block determined by splitting the first block n times. A leaf node can be the lowest node in the tree structure. A leaf node can be a node that cannot be split further. The depth of a leaf node can be a predefined maximum depth. For example, the maximum depth can be a positive integer such as 3. The root node can mean a CTU. A leaf node can mean at least one of a CU, a PU, and a TU.
[0247] - Depth can have a type depending on the type of partition. QT depth can represent the depth for quadtree partitioning. BT depth can represent the depth for binary partitioning. TT depth can represent the depth for ternary partitioning.
[0248] Sample: A sample can be a base unit that constitutes a block. A sample can be composed of one or more bits. The bit depth can be the number of bits that constitute a sample. A sample can be numbered from 0 to 2 depending on the bit depth. Bd It can be expressed as values up to -1.
[0249] PU: PU may denote a basic unit for prediction-related processing. For example, prediction-related processing may include inter-prediction, intra-prediction, intra-block copy (IBC) prediction, intra-compensation, and motion compensation.
[0250] - A PU can be divided into multiple sub-PUs, each of which has a smaller size than the PU itself. These multiple sub-PUs can also serve as the basis for prediction-related processing. In other words, a prediction unit partition generated by splitting a prediction unit can also be a prediction unit.
[0251] TU: A TU may be a basic unit for processing related to a residual block. The processing related to the residual block may include at least one of a transform, an inverse transform, quantization, inverse quantization, transform coefficient encoding, transform coefficient decoding, entropy encoding, and entropy decoding. - One TU may be split into a plurality of sub-transform units having a size smaller than the size of the TU. The plurality of sub-TUs may also be basic units for processing related to the residual block. In other words, a transform unit partition generated by splitting a transform unit may also be a transform unit.
[0252] - The transformation may include one or more of a primary transformation and a secondary transformation, and the inverse transformation may include one or more of a primary inverse transformation and a secondary inverse transformation.
[0253] Parameter set: A parameter set may correspond to header information among the structures within a bitstream.
[0254] - The parameter set may include at least one of a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), and a decoding parameter set (DPS).
[0255] - Information signaled through a parameter set can be applied to pictures referencing the parameter set. For example, information within a VPS can be applied to pictures referencing the VPS. Information within an SPS can be applied to pictures referencing the SPS. Information within a PPS can be applied to pictures referencing the PPS. A parameter set can refer to a higher-order parameter set. For example, a PPS can refer to an SPS. An SPS can refer to a VPS.
[0256] - Additionally, the parameter set may include tile group information, slice header information, and tile header information. A tile group may mean a group or slice including multiple tiles.
[0257] MPM (Most Probable Mode): MPM can indicate the intra prediction mode that is likely to be used for intra prediction for the target block.
[0258] - One or more different MPMs can be determined based on coding parameters related to the target block and properties of objects related to the target block.
[0259] - One or more MPMs may be determined based on the intra prediction mode of a reference block. There may be multiple reference blocks. Depending on which intra prediction modes are used for one or more reference blocks, one or more different MPMs may be determined. The reference blocks may include spatial neighboring blocks.
[0260] MPM List: An MPM list may contain one or more MPMs. The number of MPMs in an MPM list may be predefined.
[0261] MPM Index: The MPM index can indicate an MPM among one or more MPMs in the MPM list to be used for intra prediction for the target block.
[0262] MPM Usage Directive: The MPM usage directive can indicate whether the MPM list is used for prediction on the target block.
[0263] Prediction mode: The prediction mode may be information indicating a prediction method for a target block, such as a mode used for intra prediction or a mode used for inter prediction. The prediction mode may refer to one of the prediction-related modes described in the embodiments. In addition, the prediction mode may include at least one of an intra mode, an inter mode, and an intra block copy mode.
[0264] Reference image list: The reference image list may be a list containing one or more reference images used for prediction for the target block.
[0265] - There may be multiple reference image lists. Multiple reference image lists may include List 0 (L0), List 1 (L1), etc.
[0266] - One or more reference image lists may be used for inter prediction for a target block. Parts such as 'L0' and 'L1' in the names of information related to inter prediction may refer to reference image lists related to the information.
[0267] Reference picture: A reference picture may be an image referenced for prediction of a target block. Alternatively, the reference picture may be an image containing a reference block. The reference picture may include a previous image of the target image, a target image, and a subsequent image of the target image.
[0268] Reference image index: The reference image index may be an index indicating one reference image among one or more reference images in the reference image list that is used for prediction of the target block.
[0269] Reference Block: A reference block may be a block referenced for encoding / decoding a target block, such as prediction or filtering. For example, a reference block may include reference samples used to derive prediction samples, and may also refer to a block that provides information used for decoding the target block.
[0270] Reference sample: A reference sample may be a sample that is referenced for encoding / decoding of a target block, such as prediction and filtering.
[0271] Inter prediction indicator: The inter prediction indicator can indicate the direction of inter prediction for the target block. The inter prediction can be one of uni-directional prediction and bi-directional prediction. Alternatively, the inter prediction indicator can indicate the number of reference pictures used when generating a prediction block of the target block. Alternatively, the inter prediction indicator can indicate the number of prediction blocks used for inter prediction for the target block. The reference direction can mean the inter prediction indicator. For example, the inter prediction indicator can indicate one of uni-directional and bi-directional. Alternatively, the inter prediction indicator can have a first value of '0' for an inter mode that uses only reference pictures in the L0 reference picture list, a second value of '1' for an inter mode that uses only reference pictures in the L1 reference picture list, and a third value of '2' for an inter mode that uses at least two of the reference pictures in the L0 reference picture list and the reference pictures in the L1 reference picture list.
[0272] Prediction List Utilization Flag: The prediction list utilization flag for a specific reference image list may indicate whether at least one reference image within the specific reference image list is used to generate a prediction block of the target block. For example, a value of the prediction list utilization flag for a specific reference image list of '0' may indicate that a prediction block is not generated using a reference image within the specific reference image list. A value of the prediction list utilization flag for a specific reference image list of '1' may indicate that a prediction block is generated using a reference image within the specific reference image list.
[0273] - An inter prediction indicator can be derived using a prediction list utilization flag. Conversely, a prediction list utilization flag can be derived using an inter prediction indicator. For example, an inter prediction indicator can be derived using prediction list utilization flags for a plurality of reference image lists. If an inter prediction indicator indicates that specific reference lists among a plurality of reference image lists are used, the prediction list utilization flags of the specific reference lists indicated by the inter prediction indicator among the prediction list utilization flags of the plurality of reference image lists can be set to '1', and the prediction list utilization flags of the remaining reference image lists not indicated by the inter prediction indicator can be set to '0'.
[0274] Reference Direction: The reference direction may point to a list of reference images used for prediction of the target block. For example, the reference direction may point to one or more of the reference image list L0 and the reference image list L1.
[0275] - The reference direction may not indicate that the directions of the reference images in the reference image list are limited to the forward direction or the backward direction, but only indicates the reference image list used for prediction of the target block. That is, each of the reference image list L0 and the reference image list L1 may include forward images and backward images. Here, the forward direction may indicate the direction from the target image to the previous image of the target image. Forward inter prediction may be inter prediction that uses the previous image of the target image as a reference image. Backward direction may indicate the direction from the target image to the subsequent image of the target image. Backward inter prediction may be inter prediction that uses the subsequent image of the target image as a reference image.
[0276] - A unidirectional reference direction may mean that one reference image list is used. A bidirectional reference direction may mean that two reference image lists are used. For example, the reference direction may indicate that only the reference image list L0 is used, that only the reference image list L1 is used, or that two reference image lists are used. Additionally, the reference direction may be indicated by an inter prediction indicator.
[0277] Picture Order Count (POC): The POC of a picture can indicate the display order or output order of the picture.
[0278] Motion information: Motion information may be information used to specify a reference block. Motion information may include information used for inter prediction, such as a motion vector (MV), a reference picture index, a reference picture, an inter prediction indicator, and a prediction list utilization flag. Additionally, motion information may include information used in a specific inter prediction mode, such as an MV candidate, an MV candidate index, a merge candidate, and a merge index.
[0279] - For inter prediction of a target block, multiple motion information for multiple reference image lists can be used, respectively. Motion information for a specific reference image list can be used for prediction using the specific reference image list. Multiple (intermediate) prediction blocks can be derived from the multiple motion information. A (final) prediction block for the target block can be generated using statistical values for the multiple (intermediate) prediction blocks.
[0280] MV: MV can be a two-dimensional vector used in inter prediction. It can represent the offset between a target block and a reference block. Alternatively, it can represent the difference between the locations of a target block and a reference block.
[0281] - For example, MV is (mv x , mv y ) can be expressed in the form of mv x can represent the horizontal component, and mv y can represent vertical components.
[0282] -The zero vector can be (0, 0) MV.
[0283] Block Vector (BV): A BV can be a two-dimensional vector used in intra-block copy prediction. A BV can represent the offset between a target block within a target image and a reference block within the target image. In other words, a BV can represent the displacement between the target block and the reference block within the target image.
[0284] - For example, BV is similar to MV (bv x , bv y ) can be expressed in the form of bv x can represent the horizontal component, bv y can represent vertical components.
[0285] -The zero vector can be (0, 0) BV.
[0286] Motion Information Candidate: In a specific prediction, the motion information of the target block can be selected from among motion information candidates determined by a specific method. The motion information candidate may refer to the motion information of a reference block, or it may refer to the reference block itself containing the motion information. Here, the reference block may be a block determined by a specific method for selecting a motion information candidate.
[0287] Candidate List: A candidate list may be a list containing one or more candidates. For example, the candidate list may include a motion information candidate list, a merge candidate list, an MV candidate list, an MPM list, etc. The candidate list may be generated in the same manner by the encoding device and the decoding device. In other words, the candidate list used by the encoding device and the candidate list used by the decoding device may be the same, and the same candidate list may be shared by the encoding device and the decoding device. The encoding device may select a candidate to be used for processing the target block from among the candidates in the candidate list. An indicator indicating the selected candidate may be signaled from the encoding device to the decoding device. The decoding device may use the indicator to specify a candidate to be used for processing the target block from among the candidates in the candidate list. Alternatively, the encoding device and the decoding device may specify a candidate to be used for processing the target block from among the candidates in the candidate list according to the same rule.
[0288] Motion information candidate list: The motion information candidate list may mean a list constructed using one or more motion information candidates.
[0289] Motion information candidate index: The motion information candidate index may be an identifier or indicator that indicates a motion information candidate used for prediction of a target block among the motion information candidates in the motion information candidate list.
[0290] - In a specific inter prediction mode, motion information of other reconstructed blocks may be used to derive motion information of the target block. The other blocks may include neighboring blocks. In this specific inter prediction mode, the motion information for the target block itself is not individually signaled, but other information used to derive motion information of the target block based on the motion information of other reconstructed blocks may be signaled. In this case, the other information may include information indicating which of the other reconstructed blocks' motion information is used to derive motion information of the target block, such as a motion information candidate index.
[0291] - For example, these inter prediction modes may include AMVP mode, merge mode, and skip mode. The motion information candidate index may be a merge index or an MV candidate index.
[0292] - In embodiments, MV may be part of motion information. In embodiments, information about motion information, such as motion information candidates, motion information candidate lists, and motion information candidate indices, may be replaced with information about MVs, such as MV candidates, MV candidate lists, and MV candidate indices, and the description of motion information may also be applied to MVs.
[0293] Merge: Merge can refer to the merging of motion information across multiple blocks, or it can refer to applying motion information from another block to the target block. In other words, merge mode can refer to a mode in which the motion information of the target block is derived from the motion information of neighboring blocks.
[0294] Merge Candidate: A merge candidate may refer to a specific (restored) block used for merging the target block, or may refer to motion information for the specific block. Alternatively, the merge candidate may include motion information for the specific block.
[0295] - Merge candidates for the target block may include spatial merge candidates, temporal merge candidates, history-based candidates, average candidates based on the average of two merge candidates, and zero merge candidates.
[0296] Merge Candidate List: A merge candidate list may be a list constructed using one or more merge candidates.
[0297] Merge Index: A merge index may be an indicator that points to a merge candidate among the merge candidates in the merge candidate list, which is used for prediction of the target block. The motion information of the merge candidate indicated by the merge index among the merge candidates in the merge candidate list may be used as motion information of the target block.
[0298] Neighboring block: A neighboring block can refer to a block adjacent to the target block. Neighboring blocks can include spatial and temporal neighboring blocks. A neighboring block can also refer to a reconstructed neighboring block within a reference image.
[0299] Spatial neighboring blocks: Spatial neighboring blocks can be blocks that are spatially adjacent to the target block.
[0300] - The target block and spatial neighboring blocks can be included within the target image.
[0301] - A spatial neighboring block may include a block whose boundary is at least partially adjacent to a boundary of the target block. Alternatively, a spatial neighboring block may include a block whose distance from the target block is less than or equal to a specific value.
[0302] - A spatial neighboring block may include a block diagonally adjacent to a vertex of the target block.
[0303] - Spatial neighboring blocks may include an upper left block adjacent to the upper left of the target block, an upper block adjacent to the upper right of the target block, an upper right block entered at the upper right of the target block, a left block adjacent to the left of the target block, a right block adjacent to the right of the target block, a lower left block adjacent to the lower left of the target block, a lower block adjacent to the lower bottom of the target block, and a lower right block adjacent to the lower right of the target block.
[0304] Temporal neighboring blocks: Temporal neighboring blocks can be blocks that are temporally adjacent to the target block.
[0305] - A temporal neighboring block may include a collocated block (COL block). A collocated block may be a block within a reconstructed image within a reference image buffer. A collocated picture (col picture) may refer to an image that includes a collocated block. A collocated picture may be an image included in a reference image list.
[0306] - Call blocks can be determined based on the location of the target block within the target image. Two blocks being "temporally adjacent" can mean that the locations of the two blocks satisfy certain conditions.
[0307] - The position of a call block within a call image may be the same as the position of a target block within a target image. Alternatively, the position of a call block within a call image may correspond to the position of a target block within a target image. Here, the correspondence of the positions of blocks may mean that the areas of the blocks are identical, that an area of one block is included in an area of another block, or that one block occupies a specific position of another block.
[0308] - For example, the location of a call block within a call image may be identical to the location of a target block within the target image. Alternatively, a call block may be a block containing a call sample within a call image. A call sample may be a sample having coordinates identical to the coordinates of a specific sample in the target block.
[0309] - A temporal neighboring block may be a block that is temporally adjacent to a spatial neighboring block of the target block.
[0310] Search range: The search range can refer to a two-dimensional region where MVs are searched during inter prediction. For example, when the optimal MV must be derived for processing a target block, the optimal MV can be selected from among the MVs pointing within the search range.
[0311] Transform coefficient: The transform coefficient may be a coefficient generated by performing a transformation on the residual block. Alternatively, the transform coefficient may be a coefficient value generated by performing dequantization on a quantized level.
[0312] Quantized level: A quantized level can be an integer quantity used as input to dequantization.
[0313] Quantization: Quantization can be the process of generating quantized levels for transform coefficients. Quantized levels can be generated by applying quantization to transform coefficients. The transform can also be considered part of quantization.
[0314] Dequantization: Dequantization can be the process of multiplying a quantized level by a factor. By applying dequantization to a quantized level, (restored) transform coefficients can be generated.
[0315] Quantization Parameter (QP): QP can refer to an argument used to generate quantized levels for transform coefficients in quantization. QP can also refer to an argument used to generate (restored) transform coefficients for quantized levels in dequantization. Alternatively, QP can be a value mapped to the quantization step size.
[0316] Delta QP: Delta QP can be the difference between the QP predicted by a specific process and the QP of the target block. In other words, the QP of the target block can be the sum of the predicted QP and the delta QP.
[0317] Quantization matrix: A quantization matrix can be a matrix used in quantization or inverse quantization to improve the subjective or objective quality of an image.
[0318] Quantization matrix coefficients: Quantization matrix coefficients can be each element within a quantization matrix.
[0319] Scan: A scan can refer to the arrangement of values within a block or matrix. The values can be coefficients. For example, a scan can refer to arranging values arranged in a two-dimensional form into a one-dimensional form, or it can refer to rearranging values arranged in a one-dimensional form into a two-dimensional form. An inverse scan can be the opposite arrangement (or rearrangement) of the arrangement performed in a scan.
[0320] Non-zero transform coefficient: A non-zero transform coefficient can mean a transform coefficient with a non-zero value or a quantized level with a non-zero value.
[0321] Bitstream: A bitstream may refer to a sequence of bits containing encoded information generated by encoding an image. A bitstream may contain information according to specific syntax elements. For example, information may contain syntax elements. An encoding device may generate a bitstream containing information according to specific syntax elements. A decoding device may obtain information from the bitstream according to specific syntax elements.
[0322] Signaling: Signaling information may indicate that information is transmitted from an encoding device to a decoding device via a bitstream. For example, the information may include a syntax element. Alternatively, signaling may mean that the encoding device includes information in a bitstream. Information signaled by the encoding device may be used by the decoding device. In signaling, the bitstream may be transmitted via a network and may be included in a recording medium. In embodiments, description of information being signaled may include: 1) for signaling information, the encoding device determines and generates information; 2) the encoding device encodes the information to generate encoded information; 3) (encoded) information is transmitted from the encoding device to the decoding device via a bitstream; 4) the decoding device decodes the encoded information to obtain information; and 5) for signaling information, the decoding device determines and generates information via signaling.
[0323] - An encoding device can perform encoding on information to generate encoded information. The encoded information can be signaled via a bitstream. A decoding device can obtain information by decoding the encoded information.
[0324] - When information is signaled for a specific target, it can mean that the information is used for each specific target, and the processing indicated by the information is applied to each specific target. For example, when information is signaled at a specific unit level, it can mean that the information is used / processed for each specific unit.
[0325] - The information being signaled may include one or more sub-information. Signaling a specific piece of information may mean that each piece of information within one or more sub-information pieces contained within the specific information is signaled.
[0326] Selective Signaling: Signaling of information may be performed selectively. Selective signaling of information may mean that the encoding device selectively includes information in the bitstream (under certain conditions). Selective signaling of information may mean that the decoding device selectively obtains information from the bitstream (under certain conditions).
[0327] Omission of signaling: Signaling for information may be omitted. Omission of signaling for information may mean that the encoding device (under certain conditions) does not include the information in the bitstream. Omission of signaling for information may mean that the decoding device (under certain conditions) does not obtain the information from the bitstream. The decoding device may derive the information for which signaling is omitted using other information of the embodiments.
[0328] Symbol: may mean at least one piece of information of a target unit, such as a syntax element of a target unit or target block, a coding parameter, a quantized level, and a transform coefficient. In addition, a symbol may mean a target of entropy encoding or a result of entropy decoding.
[0329] Entropy encoding: Entropy encoding can allocate fewer bits to symbols with a high probability of occurrence, and more bits to symbols with a low probability of occurrence. This allocation reduces the size of the bitstream representing the symbols as they are represented.
[0330] - Entropy coding can use methods such as Variable Length Coding (VLC) and Context-Adaptive Binary Arithmetic Coding (CABAC). For example, in variable length coding, entropy coding can be performed using a variable length table. For example, in CABAC, a binarization method for symbols and a probability model of symbols / bins can be derived for entropy coding, and arithmetic coding using context can be performed.
[0331] Entropy decoding: Entropy decoding can reverse the processes performed in entropy encoding. Symbols can be generated by entropy decoding a bitstream.
[0332] Parsing: Parsing can mean determining the values of syntactic elements by performing entropy decoding on the encoded information in the bitstream. Alternatively, parsing can mean entropy decoding itself.
[0333] Statistical Value: The values of information related to specific entities described in the embodiments may be used as inputs to specific operations. The statistical value may be a value derived by a specific operation on the values related to these specific entities. For example, the statistical value for specific information may be one or more of an average value, a weighted average value, a weighted sum value, a minimum value, a maximum value, a mode, a median value, an interpolated value, a sum of products, and a product of sums of values of the specific information. Additionally, information of the embodiments having specific values determined by operations, such as constants, variables, and coding parameters, may have specific statistical values according to the embodiments.
[0334]
[0335] Coding parameters
[0336] In embodiments, coding parameters may be information required for coding. The coding parameters may include information signaled from an encoding device to a decoding device, information calculated / derived during the coding process described in the embodiments, and information used for the coding process described in the embodiments.
[0337] In embodiments, the coding parameters include a size of a CTU, a size of a unit, a form of a unit, a shape of a unit, a depth of a unit, a minimum unit size, a maximum unit size, a maximum unit depth, a minimum unit depth, a partition information of a unit, QT partition information, BT partition information, a partition direction of a BT partition, a partition shape of a BT partition, TT partition information, a partition direction of a TT partition, a partition shape of a TT partition, MTT partition information, a combination of MTT partitions, a partition direction of an MTT partition, a partition shape of an MTT partition, a prediction mode, an intra prediction mode, a luma intra prediction mode, a chroma intra prediction mode, an intra partition information, an inter partition information, a coding block partition information, a prediction block partition information, a transform block partition information, a reference sample line index, a reference sample filtering method, a reference sample filter tap, a reference sample filter coefficient, a prediction block filter method, a prediction block filter tap, a prediction block filter coefficient, a prediction block boundary filtering method, a prediction block boundary filter tap, a prediction block boundary filter coefficient, an inter prediction mode, motion information, MV, a motion vector difference (MV). Difference (MVD), MVD resolution, MV size, MV representation accuracy, reference picture list, reference picture, reference picture index, inter prediction direction, inter prediction indicator, prediction list utilization flag, POC, MV candidate, MV candidate index, MV candidate list, AMVP mode usage information, merge candidate, merge index, merge candidate list, merge mode usage information, motion information compensation information, skip mode usage information, intra block copy mode usage information, BV (Block Vector), Block Vector Difference (BVD), BVD resolution, BV size, BV representation accuracy, BV candidate, BV candidate index, BV candidate list, filter tap of interpolation filter, filter coefficient of interpolation filter, transformation type, transformation size, transformation selection information, primary transformation usage information,Secondary transform usage information, primary transform selection information, secondary transform selection information, residual block presence information, coded block pattern, coded block flag, QP, delta QP, quantization matrix, deblocking filter usage information, coefficients of the deblocking filter, filter taps of the deblocking filter, strength of the deblocking filter, shape / shape of the deblocking filter, adaptive sample offset usage information, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, adaptive loop filter usage information, coefficients of the adaptive loop filter, filter taps of the adaptive loop filter, shape / shape of the adaptive loop filter, binarization / debinarization method, context model, context model determination method, context model update method, regular mode usage information, bypass mode usage information, significant coefficient flag, last significant coefficient flag, coefficient group unit coding flag, last significant coefficient position, flag indicating whether the coefficient value is greater than 1, whether the coefficient value is greater than 2 A flag indicating whether the coefficient value is greater than 3, a flag indicating whether the coefficient value is greater than 3, remaining coefficient value information, sign information, context bin, bypass bin, reconstructed sample, reconstructed luma sample, reconstructed chroma sample, residual sample, residual luma sample, residual chroma sample, transform coefficient, luma transform coefficient, chroma transform coefficient, transform coefficient level, luma transform coefficient level, chroma transform coefficient level, transform coefficient level scanning method, quantized level, luma quantized level, chroma quantized level, size of MV search region on the side of the decoding device, shape of MV search region on the side of the decoding device, number of MV search on the side of the decoding device, picture type, slice identification information, slice type, slice partitioning information, tile group identification information, tile group type, tile group partitioning information, tile identification information, tile type, tile partitioning information, bit depth,It may include one or more of input sample bit depth, reconstructed sample bit depth, residual sample bit depth, transform coefficient bit depth, quantized level bit depth, mapping availability information, information about luma signal, information about chroma signal, color space of target block, color space of residual block, and temporal layer information.
[0338] In addition, the coding parameter may further include 1) a value of information that may be included in the coding parameter, 2) a combination of multiple pieces of information that may be included in the coding parameter, 3) a statistical value for information that may be included in the coding parameter, 4) information related to the coding parameter, 5) information used to calculate / derive the coding parameter, and 6) information calculated / derived using the coding parameter.
[0339] In embodiments, "X usage information" may be "information indicating whether X is used / applied / performed." Alternatively, "X usage information" may be "information indicating whether X is available." For example, "specific mode usage information" may be information indicating whether a specific mode is used. The mode information may indicate a mode used for a target block among the modes described in the embodiments. In embodiments, the specific mode usage information may be replaced with mode information, and the description of the specific mode usage information may also be applied to the mode information. "X usage information" and "X indicator" may be used interchangeably.
[0340] In embodiments, coding parameters and syntax elements may correspond to each other. For example, syntax elements of an embodiment may be used as coding parameters, and coding parameters may be signaled as syntax elements.
[0341] In embodiments, “X presence information” may be considered as “information indicating whether X exists” or “information indicating whether information indicating X exists in the bitstream.”
[0342] In embodiments, the “X selection information” may be information indicating one of the candidates or methods for X. The “X selection information” may be considered an “X index.”
[0343] In embodiments, the splitting form of a particular tree may represent one of symmetric splitting and asymmetric splitting, and may represent one of QT, BT, TT, and non-split. The splitting direction of a particular tree may represent one of horizontal and vertical directions.
[0344] In embodiments, when a coding parameter has one of multiple values, "coding parameter" may be replaced with "whether the coding parameter has a specific value among the multiple values available to the coding parameter."
[0345] In embodiments, when a coding parameter points to one of a plurality of objects, “coding parameter” may be replaced with “whether the coding parameter points to a specific object among the plurality of objects.”
[0346]
[0347] System for video coding
[0348] Figure 1 illustrates a system for video coding according to one embodiment.
[0349] The system (100) may include at least one of an encoding device (110) and a decoding device (150).
[0350] Each of the encoding device (110) and the decoding device (150) may be a computer or an electronic apparatus.
[0351]
[0352] Structure of the encoding device
[0353] The encoding device (110) may include a processor (120), storage (140), and a communicator (149).
[0354] The processor (120), storage (140), and communication device (149) can be connected via a bus.
[0355] The processor (120) may be a semiconductor device that executes instructions or computer-executable codes, such as a central processing unit (CPU). The processor (120) may be at least one hardware processor.
[0356] The processor (120) can perform generation and processing of information input to the encoding device (110), output from the encoding device (110), or used within the encoding device (110) in the embodiments, and can perform comparisons and judgments related to such information.
[0357] The processor (120) may include a plurality of components. The plurality of components may include a partitioner (122), a subtractor (124), a transformer (125), a quantizer (126), an inverse quantizer (127), an inverse transformer (128), an adder (129), a filter (130), and an entropy encoder (139).
[0358] At least some of the aforementioned components may be program modules. The program modules may be included in the encoding device (110) in the form of an operating system, applications, and other program modules. The program modules may be instructions or computer-executable codes stored in the storage (140) and executed by the processor (120).
[0359] The storage (140) may include various types of volatile storage media and non-volatile storage media. For example, the storage (140) may include memory such as ROM and RAM.
[0360] The storage (140) can store instructions and computer-executable codes used for the operation of the encoding device (110), and can store information and bitstreams described in the embodiments. The storage (140) can include a reference picture buffer (141).
[0361] The communication device (149) can perform functions related to the communication of information in the encoding device (110). For example, the communication device (149) can transmit a bitstream to the decoding device (150).
[0362] Among the names of components of the encoding device (110), “-er” or “-or” may be replaced with “-unit”. The storage (140) may also be named a storage unit.
[0363]
[0364] Operation of the encoding device
[0365] The encoding device (110) can sequentially encode one or more images of a video.
[0366] The storage (140) can store the original image. The original image can be used as a target image in the encoding device (110).
[0367] The processor (120) can generate a bitstream including encoded information by performing encoding on the target image, and can store the generated bitstream in the storage (140). The generated bitstream can be stored in a computer-readable recording medium, and can be transmitted to the communication device (189) of the decoding device (150) via a wired and / or wireless transmission medium by the communication device (149).
[0368] The segmenter (122) can determine a target block by performing segmentation on the target image.
[0369] The predictor (123) can determine the prediction mode of the target block. The predictor (123) can generate a prediction block of the target block by performing prediction according to the prediction mode.
[0370] The prediction mode of the target block may be one of the available prediction modes. For example, the available prediction modes may include intra prediction, inter prediction, and IBC prediction.
[0371] For example, if the prediction mode is intra prediction, the predictor (123) can perform intra prediction on the target block to generate a prediction block of the target block.
[0372] For example, if the prediction mode is inter prediction, the predictor (123) can perform inter prediction on the target block to generate a prediction block of the target block.
[0373] For example, when the prediction mode is IBC, the predictor (123) can perform IBC prediction on the target block to generate a prediction block of the target block.
[0374] The subtractor (124) can generate a residual block of the target block. The residual block may be the difference between the original block and the predicted block. The original block may be the area pointed to by the target block in the original image. Alternatively, the residual block may refer to a block generated by applying one or more of transformation and quantization to the difference between the original block and the predicted block.
[0375] The transformer (125) can perform a transformation on the residual block to generate transformation coefficients.
[0376] The converter (125) can perform the conversion using one of a plurality of conversion methods.
[0377] For example, the multiple transform methods may include a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), and transforms based on each transform.
[0378] Transform skip mode may be a mode for generating a reconstructed block using a reconstructed residual block and a prediction block for which transformation and inverse transformation have not been performed. When transform skip mode is applied to a target block, transformation and inverse transformation for the target block may be omitted, and only quantization and inverse quantization for the target block may be performed.
[0379] A quantizer (126) can generate quantized levels by applying quantization using quantization parameters to transform coefficients. In embodiments, the quantized levels may also be referred to as transform coefficients.
[0380] An entropy encoder (139) can generate encoded information by performing entropy encoding based on a probability distribution on information for decoding an image. The bitstream can include encoded information.
[0381] Information for decoding an image may include quantized levels and syntax elements produced by a quantizer (126).
[0382] The probability distribution can be determined based on the quantized levels and coding parameters.
[0383] The entropy encoder (139) can use scanning to change the quantized levels in the form of two-dimensional blocks into the form of one-dimensional vectors in order to perform encoding on the quantized levels. In scanning, which scan among the upper right diagonal scan, vertical scan, and horizontal scan will be used can be determined based on coding parameters such as the block size and the block intra prediction mode.
[0384] When encoding is performed on a target image / block, the predictor (123) uses a reference image / block for prediction. The encoded target image / block can be used as a reference image / block for other images / blocks to be processed later. Accordingly, the processor (120) can perform restoration on the encoded target block, and store a restored image including the restored target block generated by the restoration as a reference image in the reference picture buffer (141). Inverse quantization and inverse transformation can be performed on the encoded target block for restoration.
[0385] The dequantizer (127) can generate dequantized transform coefficients by performing dequantization on the quantized level.
[0386] The inverse transformer (128) can generate inverse quantized and inversely transformed coefficients by performing inverse transformation on the inverse quantized transform coefficients. In embodiments, the inverse quantized and / or inversely transformed coefficients may refer to coefficients to which at least one of inverse quantization and inverse transformation has been applied. The inverse quantized and inversely transformed coefficients may be a restored residual block.
[0387] The adder (129) can generate a restored block by combining a predicted block and a restored residual block.
[0388] The restoration block may pass through a filter (130). The filter (130) may apply one or more of a plurality of filters to the target. Each filter of the plurality of filters may be an in-loop filter. The target may be a restoration sample, a restoration block, or a restoration image.
[0389] The reference picture buffer (141) can store a restored block / image provided from the filter (130). The restored image may be an image including a restored block. Alternatively, the restored image may be an image composed of restored blocks.
[0390] The reference picture buffer (141) can provide the stored restored image as a reference image to the predictor (123). In terms of storing the decoded (i.e., restored) picture, the reference picture buffer (141) may also be referred to as a decoded picture buffer (DPB).
[0391]
[0392] Structure of the decryption device
[0393] The decryption device (150) may include a processor (160), a storage (180), and a communication device (189).
[0394] The description of the processor (120), storage (140), and communication device (149) related to the encoding device (110) can also be applied to the processor (160), storage (180), and communication device (189) related to the decoding device (150). Duplicate descriptions are omitted.
[0395] The processor (160) may include a plurality of components. The plurality of components may include an entropy decoder (161), a divider (162), a predictor (163), an inverse quantizer (167), an inverse transformer (168), an adder (169), and a filter (170).
[0396] The storage (180) may include a reference picture buffer (181).
[0397] The communication device (189) can perform functions related to communication of information in the decryption device (150). For example, the communication device (189) can receive a bitstream from the encoding device (110).
[0398] Among the names of components of the decryption device (150), “-er” or “-or” may be replaced with “-unit”. The storage (180) may also be named a storage unit.
[0399]
[0400] Operation of the decryption device
[0401] The communication device (149) of the encoding device (110) can transmit the bitstream generated by the encoding device (100) to the decoding device (150). Alternatively, a computer-readable recording medium storing the bitstream can transmit the bitstream generated by the encoding device (100) to the decoding device (150).
[0402] The communication device (189) can receive a bitstream from the encoding device (110) via a wired and / or wireless transmission medium. The received bitstream can be stored in the storage (180).
[0403] The processor (160) can obtain a bitstream from a storage (180) or a computer-readable recording medium.
[0404] A bitstream may contain encoded information.
[0405] An entropy decoder (161) can generate information for decoding an image by performing entropy decoding based on a probability distribution on the encoded information of a bitstream.
[0406] Information for decoding an image may include quantized levels and syntax elements.
[0407] The entropy decoder (161) can use scanning to change the quantized levels in the form of a one-dimensional vector into the form of a two-dimensional block to perform decoding on the quantized levels. In scanning, which scan among the upper right diagonal scan, vertical scan, and horizontal scan will be used can be determined based on coding parameters such as the block size and the block intra prediction mode.
[0408] The entropy decoder (161) can provide syntax elements to other components of the processor (160), such as the segmenter (162).
[0409]
[0410] A common description of the relationship between the components of the encoding device and the components of the decoding device.
[0411] The decoding device (150) performs decoding using the bitstream generated by the encoding device (110). The encoding device (110) can perform encoding on the target block using a restored image derived within the decoding device (150), rather than an original image that is not provided to the decoding device (150). Therefore, the encoding device (110) and the decoding device (150) may need to generate restored blocks / images in the same manner. In this respect, the descriptions of the divider (122), predictor (123), inverse quantizer (127), inverse transformer (128), adder (129), filter (130), and reference picture buffer (141) of the encoding device (110) disclosed in the embodiments can also be applied to the divider (162), predictor (163), inverse quantizer (167), inverse transformer (168), adder (169), filter (170), and reference picture buffer (181) of the decoding device (150). Duplicate descriptions are omitted.
[0412] Additionally, each of the divider (122), predictor (123), inverse quantizer (127), inverse transformer (128), adder (129), and filter (130) of the encoding device (110) can generate syntax element information that specifies processing for the target. Each of the divider (162), predictor (163), inverse quantizer (167), inverse transformer (168), adder (169), and filter (170) of the decoding device (150) can perform processing for the target (same as that performed in the encoding device (110)) using the syntax element information.
[0413] As described above, corresponding components of the encoding device (110) and the decoding device (150) may perform the same or corresponding functions. In embodiments, the processor may represent the processor (120) of the encoding device (110) and / or the processor (160) of the decoding device (150). For example, in the function related to prediction, the processor may represent a predictor (123), a subtractor (124), and an adder (129), and may represent a predictor (163) and an adder (169). In the function related to transformation, the processing unit may represent a transformer (125) and an inverse transformer (128), and may represent an inverse transformer (168). In the function related to quantization, the processor may represent a quantizer (126) and an inverse quantizer (127), and may represent an inverse quantizer (167). In the function related to entropy encoding / decoding, the processing unit may represent an entropy encoder (139) and / or an entropy decoder (161). In the function related to filtering, the processing unit may represent a filter (130) and / or a filter (170). The storage may represent a storage (140) of an encoding device (110) and / or a storage (180) of a decoding device (150). The reference picture buffer may represent a reference picture buffer (141) of an encoding device (110) and / or a reference picture buffer (181) of a decoding device (150). The communication unit may represent a communication unit (149) of an encoding device (110) and / or a communication unit (189) of a decoding device (150).
[0414]
[0415] Division of the units that make up the image
[0416] Figure 2 shows a segmentation structure of an image according to one embodiment.
[0417] Figure 2 can schematically represent an example in which one unit is divided into multiple sub-units.
[0418] A CU can be used as a basic unit for encoding and decoding images. In addition, a CU can be a basic unit for prediction, transformation, quantization, inverse quantization, inverse transform entropy encoding, and entropy decoding.
[0419] A CU can be used as a unit to which a prediction mode is applied. In other words, during coding, it can be determined which prediction mode among the available prediction modes will be applied to each CU. For example, available prediction modes may include intra prediction, inter prediction, and intra-block copy prediction (IBC).
[0420] The target image (200) can be sequentially divided into units of CTUs. A division structure can be determined for each CTU. The CTU can be divided into CUs according to the division structure. Alternatively, a single CTU can be used as a CU. The size of the CTU can be the maximum CU size.
[0421] Each CU can have depth information. The depth information can indicate the depth of the CU and the size of the CU. The depth of a CTU can be 0. The depth of a CU generated by splitting a CTU can be 1. When a parent CU is split into child CUs, the depth of the child CU can be 1 greater than the depth of the parent CU. The number of split CUs can be a positive integer greater than or equal to 2, including 2, 4, 8, and 16. At least one of the horizontal size and the vertical size of the child CU generated by splitting the parent CU can be smaller than at least one of the horizontal size and the vertical size of the parent CU, depending on the number of child CUs.
[0422] A partitioned CU can be recursively partitioned in the same manner up to a predefined maximum depth or a predefined minimum size. The depth of the smallest coding unit (SCU) can be the predefined maximum depth, and the size of the SCU can be the predefined minimum size. The size of the SCU can be the size of the minimum CU.
[0423] For example, the depth of a CU can range from 0 to 3. Depending on the depth of the CU, the CU can have a size from 64x64 to 8x8. A CTU with a depth of 0 can be a 64x64 block. 0 can be the minimum depth. An SCU with a depth of 3 can be an 8x8 block. 3 can be the maximum depth. A depth of 0 can represent a CTU that is a 64x64 block. A depth of 1 can represent a CU that is a 32x32 block. A depth of 2 can represent a CU that is a 16x16 block. A depth of 3 can represent an SCU that is an 8x8 block.
[0424] The partition information of a CU can indicate whether the CU is partitioned. The partition information can be a 1-bit flag. All CUs except SCUs can include partition information. For example, the partition information of a CU that is not further partitioned can be the first value, '0', and the partition information of a CU that is being partitioned can be the second value, '1'.
[0425] Quad Tree (QT) partitioning may mean that one CU is partitioned into four CUs. When a parent CU is partitioned into four child CUs, the width and height of each child CU may be half the width and half the height of the parent CU, respectively.
[0426] A binary tree (BT) split may mean that one CU is split into two CUs. For example, if a parent CU is split into two child CUs, the width or height of each child CU may be half the width or half the height of the parent CU.
[0427] A Ternary Tree (TT) partition may mean that a single CU is partitioned into three CUs. For example, if a parent CU is partitioned into three child CUs, the three child CUs can be created by partitioning the width or height of the parent CU in a ratio of 1:2:1. The width or height of the child CUs may be 1 / 4, 1 / 2, and 1 / 4 of the width or height of the parent CU, respectively.
[0428] In Fig. 2, QT type segmentation was applied to the first CTU. QT segmentation, BT segmentation, and TT segmentation were applied to the second CTU.
[0429] To partition a CTU, at least one of different types of partitions, such as QT partitioning, BT partitioning, and TT partitioning, may be applied to the CTU. Different types of partitions may be applied based on specific priorities.
[0430] For example, QT partitioning may be preferentially applied to a CTU. A CU to which QT partitioning can no longer be applied may correspond to a leaf node of QT. A CU that is a leaf node of QT may be a root node of BT and / or TT. A CU that is a leaf node of QT may be partitioned into a BT or TT form, or may not be partitioned any further. In this case, QT partitioning may not be applied again to a CU that is created by applying a BT or TT partition to a CU that is a leaf node of QT.
[0431] The partitioning of a CU corresponding to each node of QT can be signaled using QT partitioning information. The QT partitioning information can be a flag. The QT partitioning information of a unit can be information indicating whether the unit is partitioned in a QT form. A first value of the QT partitioning information, '0', can indicate that the CU is not partitioned in a QT form. The QT partitioning information having a first value can indicate a multi-type tree (MTT) partitioning. The MTT partitioning can include a BT partitioning and a TT partitioning. A second value of the QT partitioning information, '1', can indicate that the CU is partitioned in a QT form.
[0432] There may be no priority between BT and TT splits. That is, a CU corresponding to a leaf node of QT may be split into either BT or TT forms. Furthermore, a CU generated by BT or TT splits may be split again into BT or TT forms, or may not be split any further.
[0433] A CU corresponding to a leaf node of QT can become the root node of MTT. For each CU corresponding to an MTT node, the CU may further include MTT-type split direction information and split type information.
[0434] Split direction information can indicate the split direction of MTT splitting. The first value of the split direction information, '0', can indicate that the CU is split horizontally. The second value of the split direction information, '1', can indicate that the CU is split vertically.
[0435] The partition type information can indicate the partition type used for multi-type tree partitioning. The first value of the partition type information, '0', can indicate that the CU is partitioned in the TT form. The second value of the partition type information, '1', can indicate that the CU is partitioned in the BT form.
[0436] Here, each of the aforementioned split direction information and split type information may be a flag having a specific length (e.g., 1 bit).
[0437] The CU's partition information may also include QT partition information, partition direction information, and partition shape information.
[0438] CUs that are no longer split by QT splitting, BT splitting, and TT splitting can be used as units for specific processing, such as prediction, transformation, quantization, inverse quantization, inverse transform, entropy encoding, and entropy decoding. That is, for specific processing, CUs may no longer be split. Therefore, splitting information for splitting such CUs into PUs and / or TUs, etc., may not exist in the bitstream.
[0439] On the other hand, if the size of a CU is larger than the maximum TU size, the CU can be recursively split until the size of the CU becomes smaller than or equal to the maximum TU size. For example, if the size of a CU is 64x64 and the maximum TU size is 32x32, the CU can be split into four 32x32 TUs for transformation. For example, if the size of a CU is 32x64 and the maximum TU size is 32x32, the CU can be split into two 32x32 TUs for transformation.
[0440] In such cases, information regarding whether a CU is split for transformation may not be separately signaled. Whether a CU is split may be determined by comparing the size of the CU (width / height) with the maximum TU size (width / height), without signaling. For example, if the width of the CU is greater than the width of the maximum TU size, the CU may be split into two vertically. Additionally, if the height of the CU is greater than the height of the maximum TU size, the CU may be split into two horizontally.
[0441] For example, the minimum size of a CU may be 4x4. For example, the maximum size of a transform block may be 64x64. For example, the minimum size of a transform block may be 4x4. The QT minimum size may be the minimum size of a CU corresponding to a leaf node of the QT. The MTT maximum depth may be the maximum depth of the path from the root node to the leaf node of the MTT.
[0442] The BT maximum size may represent the maximum size of the CU corresponding to each node of the BT, and the TT maximum size may represent the maximum size of the CU corresponding to each node of the TT. The BT minimum size and / or the TT minimum size may be set to the minimum size of the CU.
[0443] If the depth within the MTT of a CU corresponding to a node of the MTT is equal to the maximum depth of the MTT, the CU may not be split into BT shape and / or TT shape.
[0444] Based on the various sizes and depths of the CUs described above, each piece of information described in the embodiments may or may not be present in the bitstream.
[0445] Information about the maximum or minimum size described in the embodiments may be signaled at a higher level of the CU. In the embodiments, the higher level of the CU may include a video level, a sequence level, a picture level, a subpicture level, a tile group level, a tile level, and a slice level.
[0446] The information described in the embodiments may be signaled separately for different types of slices. The different types of slices may include intra-slices and inter-slices.
[0447]
[0448] Processing blocks according to their properties
[0449] Whether a specific process described in the embodiments is applied / performed may be determined based on the properties of a block related to the specific process. Whether a specific process described in the embodiments is applied / performed may be determined based on whether the properties of a block related to the specific process satisfy a specific condition. For example, a block may include a target block, a neighboring block, and a reference block. A block may include other blocks described in the embodiments. A block may be one of the blocks and units described in the embodiments.
[0450] The blocks to which the specific processing described in the examples is applied may have a square shape or a non-square shape.
[0451] In one embodiment, the block's attributes may include the block's size. Certain processing described in the embodiments may be applied / performed when certain conditions regarding the block's size are met.
[0452] In one embodiment, the specific conditions may include a minimum block size condition and a maximum block size condition. The blocks to which the minimum block size condition applies and the blocks to which the maximum block size condition applies may be different.
[0453] In one embodiment, a minimum block size and / or a maximum block size for a particular process may be predefined.
[0454] In one embodiment, the processing of the embodiment may be applied / performed when the size of the block is greater than or equal to the minimum block size and / or when the size of the block is less than or equal to the maximum block size. Alternatively, in one embodiment, the processing of the embodiment may be applied / performed when the size of the block is greater than the minimum block size and / or when the size of the block is less than the maximum block size.
[0455] In one embodiment, the processing of the embodiment may be applied / performed only when the block size is greater than or equal to the minimum block size and less than or equal to the maximum block size. Alternatively, the processing of the embodiment may be applied / performed only when the block size is greater than or equal to the minimum block size and less than or equal to the maximum block size. Alternatively, the processing of the embodiment may be applied / performed only when the block size is greater than or equal to the minimum block size and less than or equal to the maximum block size. The processing of the embodiment may be applied / performed only when the block size is greater than or equal to the minimum block size and less than or equal to the maximum block size.
[0456] In one embodiment, the processing of the embodiment may be applied / performed only when the block size is a predefined block size.
[0457] In embodiments, the size of a block may be determined in various ways. For example, the size of a block may refer to the width or height of the block. The size of a block may refer to both the width and height of the block. The size of a block may refer to the area of the block. The size of a block may refer to 1) a result value of a known formula using the width and height of the block, 2) a result value of a formula of the embodiment, or 3) a statistical value.
[0458] Additionally, for the first size, the processing of the first embodiment among the embodiments may be applied / performed, and for the second size, the processing of the second embodiment among the embodiments may be applied / performed.
[0459] In embodiments, the block size may be 2x2, 4x4, 8x8, 16x16, 32x32, 64x64 or 128x128, etc. Alternatively, in embodiments, the block size may be (2*SIZE X )x(2*SIZE Y ) etc. SIZE X can be one of the integers greater than or equal to 1. SIZE Y can be one of the integers greater than or equal to 1.
[0460]
[0461] Predictive information for prediction
[0462] Prediction information can be used to generate a prediction block for the target block.
[0463] The encoding device (110) can generate prediction information required for prediction and can generate a bitstream including the prediction information. The prediction information can be signaled from the encoding device (110) to the decoding device (150) via the bitstream. The decoding device (150) can obtain the prediction information from the bitstream and perform prediction on the target block using the prediction information, thereby generating a prediction block.
[0464] Prediction information may include intra-prediction information, inter-prediction information, and IBC prediction information. In embodiments, prediction information may be replaced with intra-prediction information, inter-prediction information, and / or IBC information. Intra-prediction information may include information used for intra-prediction as described in embodiments. Inter-prediction information may include information used for inter-prediction as described in embodiments. IBC information may include information used for IBC prediction as described in embodiments.
[0465]
[0466] Intra prediction
[0467] Figure 3 illustrates the structure of intra prediction according to one embodiment.
[0468] Intra prediction can be performed using reference samples and coding parameters of the target block. The reference sample can be a (restored) sample within the (restored) reference block. Alternatively, an intermediate prediction sample can be generated using a sample described in the embodiment, such as a reconstructed sample, and a reference sample can be generated again using the intermediate prediction sample. Processing described in the embodiment, such as filtering, can be applied when generating the reference sample.
[0469] A reference block may be a (spatial) neighboring block of the target block. The coding parameters may be coding parameters for the target block and / or coding parameters for the reference block. In intra prediction, a reference sample may mean a neighboring sample.
[0470] A prediction block can be generated by performing intra prediction on a target block according to an intra prediction mode based on reference samples within a target image and information related to the reference samples. The size of the target block and the size of the prediction block can be the same.
[0471] In embodiments, the prediction block may be a PU. Alternatively, the prediction block may correspond to a CU or TU described in the embodiments. The prediction block may have a square or rectangular shape.
[0472] An intra prediction mode can be expressed by at least one of a mode number, a mode value, a mode angle, and a mode direction. The prediction directions of a plurality of intra prediction modes for a target block are illustrated in the lower right corner of Fig. 3. Among the plurality of intra prediction modes, the remaining intra prediction modes excluding the DC and planar modes may be directional modes. A directional mode may be an intra prediction mode having a specific direction or a specific angle. The intra prediction mode for the target block may be selected from among directional modes and non-directional modes.
[0473] In the lower right rectangle representing the target block, the number '0' may represent the planar mode, which is a non-directional intra prediction mode. The number '1' may represent the DC mode, which is a non-directional intra prediction mode. In the lower right rectangle representing the target block, the arrows from the center to the periphery of the rectangle may represent the prediction directions of the directional intra prediction modes. In addition, the number indicated close to the arrow may represent an example of the mode value assigned to the intra prediction mode or the prediction direction of the intra prediction mode.
[0474] Intra prediction can be performed based on an intra prediction mode for the target block. One of the available intra prediction modes for the target block can be used as the intra prediction mode for the target block.
[0475] The number of intra prediction modes available to a target block may be a predefined value. Alternatively, the number of intra prediction modes available to a target block may be determined based on the properties of the prediction block. For example, the properties of the prediction block may include coding parameters such as shape, size, and color components.
[0476] For example, in Figure 3, the directional modes depicted by the dotted lines (i.e., the directional modes numbered between -14 and -1, or between 67 and 80) can only be applied to predictions for non-square blocks. Therefore, the number of intra prediction modes available for predictions for square blocks can be 67 (planar mode, DC mode, and 65 directional modes).
[0477] For example, the number of available intra prediction modes may vary depending on whether the color component of the block is a luma signal or a chroma signal. The number of available intra prediction modes for a block containing a luma component may be greater than the number of available intra prediction modes for a block containing a chroma component.
[0478] Intra prediction modes may include horizontal-below mode, horizontal mode, vertical mode, and vertical-right mode. The horizontal-below mode may be an intra prediction mode located below the horizontal mode. The vertical-right mode may be a mode located to the right of the vertical mode. For example, in FIG. 3, the mode value of the horizontal mode may be 18. The mode value of the vertical mode may be 50. Intra prediction modes whose mode values are one of 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, and 66 may be vertical-right modes. Intra prediction modes whose mode value is one of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, and 17 can be horizontal bottom modes.
[0479] The number of intra prediction modes and the mode number of each intra prediction mode described above may be merely exemplary. The number of intra prediction modes and the mode number of each intra prediction mode described above may be defined differently depending on the embodiment, implementation, and / or needs.
[0480] When the intra prediction mode is the planar mode, when generating a prediction block of a target block, a sample value of the prediction sample can be generated using a weighted sum (weighted sum) of an upper reference sample of the target sample, a left reference sample of the target sample, an upper right reference sample of the target block, and a lower left reference sample of the target block, depending on the position of the prediction sample within the prediction block.
[0481] When the intra prediction mode is DC mode, a prediction block may be generated based on an average of sample values of a plurality of reference samples. The plurality of reference samples may include upper reference samples and left reference samples of the target block. The value of the prediction sample of the prediction block may be determined based on an average of the sample values of the plurality of reference samples. In addition, filtering using the values of the reference samples may be performed for specific rows and / or specific columns within the target block. The specific rows may be one or more upper rows adjacent to the upper reference samples. The specific columns may be one or more left columns adjacent to the left reference samples.
[0482] When the intra prediction mode is a directional mode, a prediction block can be generated using the top reference sample, the left reference sample, the top right reference sample, and / or the bottom left reference sample of the target block.
[0483] The intra prediction mode of the target block may be determined based on the intra prediction mode of a neighboring block of the target block. Information for determining the intra prediction mode of the target block may be signaled.
[0484] For example, if the intra prediction modes of the target block and the neighboring block are the same, an indicator indicating that the intra prediction modes of the target block and the neighboring block are the same can be signaled.
[0485] For example, an indicator may be signaled that indicates an intra prediction mode that is the same as the intra prediction mode of the target block among the intra prediction modes of multiple neighboring blocks.
[0486] For example, if the intra prediction modes of the target block and neighboring blocks are different, an indicator indicating the intra prediction mode of the target block may be signaled. Alternatively, information used to derive the intra prediction mode of the target block based on the intra prediction mode of the neighboring block may be signaled.
[0487] Reference samples used for intra prediction for the target block may include lower left reference samples, left reference samples, upper left reference samples, upper reference samples, and upper right reference samples.
[0488] For example, the left reference samples may be reconstructed reference samples adjacent to the left side of the target block. The top reference samples may be reconstructed reference samples adjacent to the top side of the target block. The top left reference sample may be a reconstructed reference sample diagonally adjacent to the top left side of the target block. The bottom left reference samples may be reference samples located below the left reference samples among samples located on the same line as the left sample line composed of the left reference samples. The top right reference samples may be reference samples located on the right side of the top reference samples among samples located on the same line as the top sample line composed of the top reference samples.
[0489] Reference samples used for intra prediction for a target block can be determined based on the intra prediction mode of the target block. One or more reference samples can be used to determine the sample value of a prediction sample of a prediction block. In FIG. 3, the direction of the intra prediction mode indicated by the arrow can represent the direction from the prediction sample to the reference sample. The direction of the intra prediction mode can represent the dependency relationship between the reference samples and the prediction samples. For example, depending on the intra prediction mode, the sample value of a specific reference sample can be used as the sample value of at least one sample of the prediction block. Here, the specific reference sample and the at least one sample of the prediction block can be samples designated by a straight line in the direction of the intra prediction mode. In other words, the sample value of the specific reference sample can be copied as the sample value of the prediction sample located in the reverse direction of the direction of the intra prediction mode. Alternatively, the sample value of the prediction sample of the prediction block can be the sample value of the reference sample located in the direction of the intra prediction mode based on the position of the prediction sample.
[0490] The reference samples used for intra prediction may not be limited to samples immediately adjacent to the target block. As illustrated in FIG. 3, at least one of reference sample lines 0 to 3 may be used for intra prediction of the target block.
[0491] Each reference sample line of FIG. 3 may include one or more reference samples. A smaller number of a reference sample line may be a line of reference samples closer to the target block. Reference sample line 0 may be a line of reference samples immediately adjacent to the target block. When the upper left coordinates of the target block are (X, Y), the horizontal length is W, and the vertical length is H, the reference samples of reference sample line 0 may be samples whose x-coordinate is X-1 or whose y-coordinate is Y-1. Here, the y-coordinates of the reference samples whose x-coordinate is X-1 may be Y-1 to Y+2H. The x-coordinates of the reference samples whose y-coordinate is Y-1 may be X-1 to X+2W. The reference samples of reference sample line A may be samples whose x-coordinate is XA-1 or whose y-coordinate is YA-1. Here, the y-coordinates of the reference samples whose x-coordinate is XA-1 may be YA-1 to Y+2H+A. The x-coordinates of reference samples whose y-coordinate is YA-1 can be XA-1 to X+2W+A. A can be 1, 2, or 3.
[0492] Instead of obtaining samples from the reconstructed neighboring blocks, samples of segment A and segment F can be derived using padding using the nearest samples from segment B and segment E, respectively.
[0493] A reference sample line index may indicate a reference sample line among multiple reference sample lines used for intra prediction of a target block. For example, the reference sample line index may have a value of one of 0 to 3. The reference sample line index may be signaled.
[0494] When inter-color component intra prediction is used for a target block, a prediction block of a second color component can be generated based on a reconstructed block of a first color component for the target block. For example, the first color component can be a luma component, and the second color component can be a chroma component.
[0495] For intra prediction between color components, parameters between the first color component and the second color component can be derived based on a template. For example, the parameters can be parameters of a linear model.
[0496] For example, the template may include a top reference sample and / or a left reference sample of the target block, and may include a top reference sample and / or a left reference sample of the restoration block of the first color component corresponding to these reference samples.
[0497] Once the parameters are derived, a prediction block of a second color component for the target block can be generated by applying the reconstructed block of the first color component to a linear model. Depending on the image format or the type of intra prediction between color components, subsampling / downsampling can be performed on the surrounding samples of the reconstructed block of the first color component and the reconstructed block of the first color component. When subsampling is performed, the derivation of the parameters and intra prediction between color components can be performed using the corresponding samples derived by the subsampling.
[0498] Intra Sub-Partitions (ISP) prediction may refer to sequential intra prediction for multiple sub-blocks generated by dividing a target block. In ISP prediction, a target block may be divided into two or four sub-blocks in the horizontal and / or vertical directions. The divided sub-blocks may be sequentially reconstructed. As intra prediction is performed on a sub-block, a sub-prediction block for the sub-block may be generated. Additionally, as inverse quantization and / or inverse transformation is performed on the sub-block, a sub-residual block for the sub-block may be generated. A reconstructed sub-block may be generated by adding the sub-prediction block to the sub-residual block. The reconstructed sub-block may be used as a reference sample for intra predictions for other sub-blocks to be processed subsequently.
[0499] In performing prediction on a target block, it can be determined whether samples included in a reconstructed neighboring block can be used as reference samples of the target block. If there is an unavailable sample among the samples of the neighboring block that cannot be used as a reference sample of the target block, a value generated by copying and / or interpolating using the sample value of at least one sample among the samples included in the reconstructed neighboring block can replace the sample value of the unavailable sample. If the value generated by copying and / or interpolating is replaced with the sample value of the sample, the sample can be used as a reference sample of the target block.
[0500] In intra prediction, the sample value of a prediction sample of a prediction block can be determined by the sample value of a reference sample. The position of the reference sample can be specified by the position of the prediction sample and the direction of the intra prediction mode. If the position specified by the position of the prediction sample and the direction of the intra prediction mode is an integer position, the sample value of one reference sample pointed to by the integer position can be used to determine the sample value of the prediction sample of the prediction block. If the position specified by the position of the prediction sample and the direction of the intra prediction mode is not an integer position, an interpolated reference sample can be generated based on two reference samples closest to the specified position. The sample value of the interpolated reference sample can be used to determine the sample value of the prediction sample. That is, when the position specified by the position of the prediction sample and the direction of the intra prediction mode represents a space between two reference samples, an interpolated sample value can be generated based on the sample values of the two samples.
[0501]
[0502]
[0503] Inter prediction
[0504] Figure 4 shows the structure of inter prediction to explain the inter prediction process according to one embodiment.
[0505] The rectangle illustrated in Fig. 4 can represent an image. Additionally, the arrow in Fig. 4 can represent a prediction direction.
[0506] Each picture composing a video can be classified into an I picture (i.e., an intra picture), a P picture (i.e., a uni-prediction picture), and a B picture (i.e., a bi-prediction picture) according to its coding type. Coding can be performed for each picture according to its coding type.
[0507] If the target picture is an I-picture, coding for the target picture can be performed using information within the target picture without inter prediction referring to other images. For example, coding for the I-picture can be performed using intra prediction and / or IBC prediction.
[0508] Coding for P pictures and B pictures can be performed by at least one of intra prediction, IBC prediction, and inter prediction using a reference picture.
[0509] If the target picture is a P picture, coding for the target picture can be performed using unidirectional inter prediction using one reference picture list.
[0510] When the target picture is a B picture, coding for the target picture can be performed using unidirectional inter prediction or bidirectional inter prediction using two reference picture lists.
[0511] Below, inter prediction for a target block in inter mode according to an embodiment is specifically described.
[0512] When the prediction mode of the target block is inter mode, inter prediction can be performed on the target block. The target block can be a prediction block or a split prediction block.
[0513] Inter prediction can be performed using reference images and motion information. In inter prediction, a reference image can be selected using a reference image index, and a reference block corresponding to a target block within the reference image can be determined using motion information. A prediction block for the target block can be generated using the determined reference block.
[0514] Motion information can be derived using coding parameters, etc. For example, motion information can be derived using motion information of a reconstructed neighboring block, motion information of a call block, and / or motion information of a block adjacent to a call block.
[0515] In embodiments, a candidate list may be used for inter prediction. The candidate list may include multiple candidates. An index indicating a candidate used for inter prediction for a target block among the candidates in the candidate list may be signaled. The candidate list may be derived in the same manner based on the same information in the encoding device (110) and the decoding device (150). Here, the same information may include a restored image and a restored block. Furthermore, in order to specify a candidate by index, the order of the candidates within the candidate list may need to be consistent.
[0516] In one embodiment, prediction of a target block can be performed by using motion information of a spatial candidate or a temporal candidate as motion information of the target block. The motion information of the spatial candidate may be referred to as spatial motion information. The motion information of the temporal candidate may be referred to as temporal motion information.
[0517] A spatial candidate may be a restored spatial neighboring block that is spatially adjacent to the target block.
[0518] A spatial candidate may be a block that 1) exists within the target image, 2) has already been restored through decryption, and 3) is adjacent to the target block.
[0519] Spatial candidates may include the left block, the top block, the bottom left block, the top right block, and the top left block of the target block.
[0520] A temporal candidate may be a restored temporal neighboring block corresponding to a target block in a restored COL image.
[0521] In embodiments, the motion information of a spatial candidate may be motion information of a block containing the spatial candidate. The motion information of a temporal candidate may be motion information of a block containing the temporal candidate.
[0522] In inter prediction, a call (COL) block for a target block can be identified. The area of the target block within the target image and the area of the call block within the call image may be identical. In other words, a call block may be a block occupying a specific area within the call image. The specific area may correspond to the area of the target block within the call image.
[0523] A temporal candidate may be a location inside and / or outside a call block within a call image.
[0524] For example, a call block may include a first call block and a second call block. When the upper left coordinates of a call block are (xP, yP) and the size of the call block is (nPSW, nPSH), the first call block may be a block occupying coordinates (xP + nPSW, yP + nPSH). The second call block may be a block occupying coordinates (xP + (nPSW >> 1), yP + (nPSH >> 1)). The second call block may be optionally used as a call block when the first call block is unavailable.
[0525] The MV of the target block can be determined based on the MV of the call block. Scaling can be performed on the MV of the call block. The scaled MV of the call block can be used as the MV of the target block or as the predicted MV. Alternatively, the MV of the temporal candidate stored in the candidate list associated with inter prediction can be a scaled MV.
[0526] The ratio of the scaled MV and the MV of the call block may be equal to the ratio of the first temporal distance and the second temporal distance. The first temporal distance may be the distance between the reference image and the target image of the target block. The second temporal distance may be the distance between the reference image and the call image of the call block.
[0527] The method by which motion information is derived can be determined by the inter prediction mode of the target block. For example, as the inter prediction mode, AMVP mode, merge mode, skip mode, merge mode with MVD, subblock merge mode, GPM, Combined Inter Intra Prediction (CIIP) mode, and affine inter mode can be used. In the embodiments below, each of the inter prediction modes is described.
[0528]
[0529] AMVP mode
[0530] When the AMVP mode is used as a prediction mode, an MV candidate list including one or more MV candidates can be generated using the MV of the spatial candidate, the MV of the temporal candidate, the history-based MV candidate, and the zero vector. At least one of the MV of the spatial candidate, the MV of the temporal candidate, and the zero vector can be determined and used as an MV candidate.
[0531] A spatial candidate may include a reconstructed spatial neighboring block. The MV of the reconstructed spatial neighboring block may be referred to as a spatial MV candidate (spatial motion vector candidate). A temporal candidate may include a called block and a block adjacent to the called block. The MV of the called block or the MV of a block adjacent to the called block may be referred to as a temporal MV candidate (temporal motion vector candidate). A history-based MV candidate may be an MV in a list including MVs of other blocks that were previously encoded / decoded prior to encoding / decoding of the target block.
[0532] The encoding device (110) can use the MV candidate list to determine an MV to be used for encoding the target block within the search range. The maximum number of MV candidates in the MV candidate list can be predefined. N can represent a predefined maximum number. For example, N can be 2. Alternatively, the maximum number of candidates can be signaled from the encoding device to the decoding device or derived from the decoding device. The encoding device (110) can determine an MV candidate to be used as a prediction MV of the target block among the MV candidates in the MV candidate list. The MV to be used for encoding the target block can be an MV that can be encoded at the minimum cost. The encoding device (110) can determine whether to use the AMVP mode in encoding the target block, and can generate AMVP mode usage information indicating whether the AMVP mode is used.
[0533] Inter prediction information may include 1) AMVP mode usage information, 2) MV candidate index, 3) MVD, 4) MVD resolution information, 5) reference direction, and 6) reference image index, and may include a residual block. The inter prediction information may be signaled from the encoding device (110) to the decoding device (150) in the form of a bitstream.
[0534] The decoding device (150) can obtain AMVP mode usage information from the bitstream. If the AMVP mode usage information indicates that the AMVP mode is used, the decoding device (150) can obtain an MV candidate index, an MVD, MVD resolution information, a reference direction, and a reference image index from the bitstream. Among the MV candidates included in the MV candidate list, an MV candidate indicated by the MV candidate index can be selected as the prediction MV of the target block.
[0535] The MVD may represent the difference between the MV to be actually used for inter prediction of the target block and the predicted MV. The encoding device (110) may derive a predicted MV that is close to the MV to be actually used for inter prediction of the target block in order to use an MVD with a size as small as possible. The decoding device (150) may derive the MV of the target block by combining the MVD and the predicted MV. In other words, the MV of the target block derived by the decoding device (150) may be the sum of the MVD and the predicted MV candidate.
[0536] Additionally, the encoding device (110) can generate MVD resolution information. The MVD resolution information may be information used to adjust the resolution of the MVD. The decoding device (150) can adjust the resolution of the MVD using the MVD resolution information.
[0537] Meanwhile, the encoding device (110) can calculate the MVD based on the affine model. The affine control point MV of the target block can be derived based on the sum of the affine control point MV candidate and the MVD. The MV of each subblock within the target block can be derived using the affine control point MV.
[0538]
[0539] Merge mode
[0540] When merge mode is used, a merge candidate list including multiple merge candidates can be generated using motion information of spatial candidates and motion information of temporal candidates. The motion information can include 1) MV, 2) reference image index, and 3) reference direction. The merge candidate can be motion information.
[0541] Merge candidates may include 1) a spatial merge candidate generated based on a spatial candidate, 2) a temporal merge candidate generated based on a temporal candidate, 3) a history-based merge candidate, 4) an average merge candidate, and 5) a zero merge candidate.
[0542] A history-based merge candidate may be motion information within a list that includes motion information of other blocks that were previously encoded / decoded prior to encoding / decoding of the target block.
[0543] An average merge candidate may be a merge candidate generated based on the average of two merge candidates in the merge candidate list.
[0544] A zero merge candidate may be zero vector motion information. Zero vector motion information may be motion information whose MV is a zero vector.
[0545] Merge candidates can be added to the merge candidate list according to a predefined method and a predefined order so that the merge candidate list has a set number of merge candidates. The same merge candidate list can be constructed in the encoding device (110) and the decoding device (150) through the predefined method and the predefined order.
[0546] The encoding device (110) can select a merge candidate to be used for encoding a target block from among the merge candidates in the merge candidate list. The encoding device (110) can determine whether to use a merge mode in encoding the target block, and can generate merge mode usage information indicating whether the merge mode is used.
[0547] Inter prediction information may include 1) merge mode usage information, 2) merge index, and 3) correction information, and may include a residual block. Inter prediction information may be signaled from an encoding device (110) in bitstream form to a decoding device (150) in bitstream form.
[0548] The decoding device (150) can obtain merge mode usage information from the bitstream. If the merge mode usage information indicates that the merge mode is being used, the decoding device (150) can obtain information related to the merge mode, such as a merge index, from the bitstream.
[0549] The encoding device (110) can select an optimal merge candidate from among the merge candidates included in the merge candidate list, and set the value of the merge index to point to the selected merge candidate.
[0550] Correction information may be information used for correcting an MV. The encoding device (110) may generate the correction information. The decoding device (150) may perform correction on the MV of the merge candidate selected by the merge index based on the correction information, thereby deriving a corrected MV. The corrected MV may be used as the MV of the target block.
[0551] In one embodiment, the correction information may include an MVD. The correction information may include one or more of correction usage information, correction direction information, and correction size information. The correction usage information may indicate whether correction is used for the MV. A merge mode that performs correction for the MV based on the correction information may be referred to as a merge mode with an MVD.
[0552] In merge mode, prediction for a target block can be performed using a merge candidate pointed to by a merge index among the merge candidates included in the merge candidate list.
[0553] Motion information of the target block can be derived from 1) MV, 2) reference image index, and 3) reference direction of the merge candidate pointed to by the merge index.
[0554] In one embodiment, the merge candidates in the merge candidate list may be specific modes that derive inter-prediction information. The merge candidate may be information indicating a specific mode that derives inter-prediction information. Inter-prediction information of the target block may be derived according to the specific mode indicated by the merge candidate. From this perspective, a specific mode may be considered a specific inter-prediction information derivation mode or a specific motion information derivation mode. A specific mode may include a series of processes that derive inter-prediction information.
[0555] Inter prediction information of a target block can be derived based on a specific mode indicated by a merge candidate selected by a merge index among the merge candidates in the merge candidate list. For example, the specific modes may include a subblock-level motion information derivation mode and an affine motion information derivation mode, and may include other modes for deriving motion information described in the embodiments.
[0556] Skip mode may be a mode that does not use residual blocks. That is, when skip mode is used, the reconstructed block may be identical to the predicted block. The description of merge mode in the embodiments may also apply to skip mode. The difference between merge mode and skip mode may be whether or not residual blocks are signaled and used. That is, skip mode may be similar to merge mode except that residual blocks are not transmitted / used, and the description of merge mode may also apply to skip mode.
[0557] The subblock merge mode may be a mode in which motion information of a target subblock is derived for a target subblock within a target block. When the subblock merge mode is applied, a list of subblock merge candidates may be generated using affine control point motion vector merge candidates and / or subblock-based temporal merge candidates. The subblock-based temporal merge candidates may be motion information of a call subblock of the target subblock.
[0558] In GPM, a first prediction block and a second prediction block can be generated using two pieces of motion information for a target block. For each coordinate of the target block, a final prediction sample of a final prediction block can be generated using a weighted sum of the first prediction sample of the first prediction block and the second prediction sample of the second prediction block.
[0559] Here, the first weight for the weighted consensus first prediction sample and the second weight for the weighted consensus second prediction sample can be determined based on the boundary of the GPM. The boundary can represent a dividing line that divides the target block. Based on the boundary, the target block can be divided into a first divided region and a second divided region.
[0560] If the distance between the final prediction sample and the boundary is less than or equal to a reference value, the value of the final prediction sample of the final prediction block may be determined using a weighted sum of the first prediction sample of the first prediction block and the second prediction sample of the second prediction block. If the distance between the final prediction sample and the boundary is greater than the reference value, one of the first weight and the second weight may be 1, and the other may be 0.
[0561] Combined Inter-Intra Prediction (CIIP) mode may be a mode that derives a prediction sample of a target block using a weighted sum of prediction samples generated by inter prediction and prediction samples generated by intra prediction.
[0562] In the aforementioned modes, self-improvement of the derived motion information can be performed, and the improved motion information can be used as motion information for the target block. For example, blocks within a specific region determined based on the derived motion information can be searched, and the motion information of the block with the smallest sum of absolute differences (SAD) value among the searched blocks can be used as the improved motion information for the target block. The specific region can be a square region within a reference image specified by the motion information. The point indicated by the motion information can be the center of the specific region.
[0563] In the aforementioned modes, compensation for prediction samples derived through inter prediction can be performed using optical flow.
[0564]
[0565] Figure 5 shows the order in which spatial candidates are added to the candidate list according to one embodiment.
[0566] In Fig. 5, the locations of spatial candidates are shown.
[0567] The large block in the center can represent the target block. The five smaller blocks adjacent to the target block can represent spatial candidates.
[0568] The coordinates of the target block can be (xP, yP), and the size of the target block can be (nPSW, nPSH).
[0569] A spatial candidate A0 may be a block adjacent to the lower left of the target block. A0 may be a block that occupies samples at coordinates (xP - 1, yP + nPSH).
[0570] A spatial candidate A1 may be a block adjacent to the left of the target block. A1 may be the bottommost block among the blocks adjacent to the left of the target block. Alternatively, A1 may be a block adjacent to the top of A0. A1 may be a block that occupies a sample at coordinates (xP - 1, yP + nPSH - 1).
[0571] A spatial candidate B0 may be a block adjacent to the upper right of the target block. B0 may be a block that occupies a sample at coordinates (xP + nPSW, yP - 1).
[0572] A spatial candidate B1 may be a block adjacent to the top of the target block. B1 may be the rightmost block among the blocks adjacent to the top of the target block. Alternatively, B1 may be a block adjacent to the left of B0. B1 may be a block that occupies a sample at coordinates (xP + nPSW - 1, yP - 1).
[0573] A spatial candidate B2 may be a block adjacent to the upper left of the target block. B2 may be a block that occupies a sample at coordinates (xP - 1, yP - 1).
[0574] As shown in Figure 5, when adding spatial candidates to the candidate list, B1, A1, The order of B0, A0 and B2 can be used, i.e. B1, A1, Available spatial candidates can be added to the candidate list in the order of B0, A0, and B2. The order in which the spatial candidates are added to the merge candidate list illustrated in Fig. 5 may be merely an example.
[0575] The above candidate list may include a motion information candidate list, a merge candidate list, an MV candidate list, a BV candidate list, and an MPM list.
[0576] To include a spatial or temporal candidate in the candidate list, its availability can be determined. If the candidate block is outside the boundaries of an image, slice, or tile, the candidate block's availability can be set to false. The phrase "availability is set to false" can mean "it is set to non-availability."
[0577] The maximum number of candidates in a candidate list can be set. N can represent the set maximum number. The set maximum number can be signaled through a parameter set or header, etc. For example, the maximum number of candidates in the candidate list for a target block within a slice can be set by the slice header. For example, the default value of N can be 5.
[0578]
[0579] IBC mode
[0580] IBC mode may be an intra-block copy prediction mode that generates prediction blocks for target blocks by referencing already-restored regions within the target image. In this respect, IBC mode may also be referred to as a current image reference mode. A block vector (BV) may be used to specify the already-restored region.
[0581] Whether the target block is encoded / decoded in IBC mode can be determined using IBC mode usage information. The encoding device (110) can determine whether to use IBC mode in encoding the target block and can generate IBC mode usage information indicating whether IBC mode is used. The decoding device (150) can obtain IBC mode usage information from the bitstream.
[0582] In IBC mode, a prediction block of a target block can be generated based on a block vector (BV). The BV can specify a reference block. The BV can indicate displacement between the target block and the reference block. The reference block can be a block within the target image. The description of the MV in the embodiments can also be applied to the BV.
[0583] The IBC mode may include skip mode, merge mode, and AMVP mode. The description of the AMVP mode, merge mode, and skip mode of the embodiments may also be similarly applied to the AMVP mode, merge mode, and skip mode of the IBC mode.
[0584] In skip mode or merge mode, a merge candidate list can be constructed, and a merge index can specify one merge candidate among the merge candidates in the merge candidate list. The BV of the specified merge candidate can be used as the BV of the target block.
[0585] In AMVP mode, BVD can be used. The description of MVD in the embodiments can also be applied to BVD.
[0586] The reference block in IBC mode may be limited to a block within an already reconstructed region of the target image. Alternatively, the reference block may be contained within at least one of the target CTU or the left CTUs. For example, the value of BV may be limited so that the reference block is located within a specific region. The specific region may be an area of three blocks of a specific size that are encoded / decoded before the block of a specific size that contains the target block. The specific size may be 64x64.
[0587]
[0588] Transformation and quantization
[0589] A quantized level can be generated by performing transformation and / or quantization on a residual block. The residual block can represent the difference between the original block and the predicted block. A reconstructed residual block can be generated by performing inverse quantization and / or inverse transformation on the quantized level. The reconstructed residual block can represent the difference between the reconstructed block and the predicted block.
[0590] When a transformation or inverse transformation is performed, a separable transformation or a 2-dimensional (2D) non-separable transformation can be performed on the residual block. A separable transformation can be a transformation that performs 1-dimensional (1D) transformations on the residual block in each of the horizontal and vertical directions.
[0591] The transform kernels used for the transformation may include various DCT kernels such as DCT type 2 (DCT-II), 1) DST kernels, and 3) kernels induced by training. For 1D transform, the DCT type and DST type may include DCT-V, DCT-VIII, DST-I, and DST-VII in addition to DCT-II.
[0592] A set of transforms may be used to determine the DCT type, DST type, or learning-induced kernel to be used for the transformation. Each transform set may include multiple transform candidates. Each transform candidate may be a DCT type, a DST type, or a learning-induced kernel.
[0593] The encoding device (110) can perform transformation and inverse transformation using transformation candidates included in the transformation set. The decoding device (150) can perform inverse transformation using transformation candidates included in the transformation set. Transform selection information indicating which transformation candidate among a plurality of transformation candidates included in the transformation set applied to the residual block is used can be signaled. The transformation selection information can include vertical transformation selection information and horizontal transformation selection information. The vertical transformation selection information can indicate which transformation among the transformations included in the transformation set is used for vertical transformation. The horizontal transformation selection information can indicate which transformation among the transformations included in the transformation set is used for horizontal transformation.
[0594] The transform may include at least one of a primary transform and a secondary transform. A primary transform coefficient may be generated by performing a primary transform on a residual block, and a secondary transform coefficient may be generated by performing a secondary transform on the transform coefficient. Here, the transform coefficient may include a primary transform coefficient and a secondary transform coefficient.
[0595] The primary transformation may mean Multiple Transform Selection (MTS), which applies different transformations for each of the 1D directions (i.e., vertical and horizontal directions).
[0596] A secondary transform may be a transform for improving the energy concentration of the transform coefficients generated by the primary transform. The secondary transform may be 1) a separable transform like the primary transform, or 2) a 2D non-separable transform. The 2D non-separable transform may refer to a low frequency non-separable transform (LFNST) or a non-separable primary transform (NSPT).
[0597] NSPT can be applied to specific block sizes such as 4x4, 4x8, 8x4, 4x16, 16x4, 8x8, 8x16, and 16x8 for intra coding.
[0598] The primary transform can be performed using at least one of a plurality of predefined transform methods. For example, the plurality of predefined transform methods can include DCT, DST, and KLT. In addition, the primary transform can be a transform having various transform types according to a transform kernel function defining DCT and DST. For example, the primary transform can include a plurality of transforms such as DCT-2, DCT-4, DCT-5, DCT-7, DCT-8, DST-1, DST-2, DST-4, DST-7, and DST-8 according to a plurality of transform kernels.
[0599] In one embodiment, the transform type may be determined based on coding parameters associated with the target block. For example, the transform type may be determined based on one or more of: 1) a prediction mode of the target block (e.g., one of intra prediction and inter prediction), 2) a size of the target block, 3) a shape of the target block, 4) an intra prediction mode of the target block, 5) a component of the target block (e.g., one of a luma component and a chroma component), and 6) a split type applied to the target block (e.g., one of QT, BT, TT, and non-split).
[0600] As with the first-order transformation, a set of transformations can also be defined for the second-order transformation. The methods for deriving and / or determining the set of transformations of the embodiments can be applied to both the first-order transformation and the second-order transformation.
[0601] In one embodiment, a primary transformation and / or a secondary transformation may be determined for a specific target. The transformation selection information may include transformation target information. The transformation target information may indicate the target to which the primary transformation and / or the secondary transformation is applied.
[0602] For example, a first-order transform and / or a second-order transform may be applied to one or more of the signal components, including the luma component and the chroma component.
[0603] In one embodiment, the transform selection information may include primary transform usage information and secondary transform usage information. The primary transform usage information may indicate whether the primary transform is applied to the residual block of the target block. The secondary transform usage information may indicate whether the secondary transform is applied to the residual block of the target block.
[0604] In one embodiment, whether a primary transform and / or a secondary transform is applied may be determined based on coding parameters for the target / neighboring blocks, such as the size and shape of the target / neighboring blocks.
[0605] In one embodiment, the transform selection information may include primary transform selection information and secondary transform selection information. The primary transform selection information may indicate a transform method to be applied to a residual block among a plurality of transform methods that may be used in the primary transform. The primary transform selection information may be a primary transform index. The secondary transform selection information may indicate a transform method to be applied to a transform coefficient among a plurality of transform methods that may be used in the secondary transform. The secondary transform selection information may be a secondary transform index.
[0606] In one embodiment, the transformation methods of the first and second transformations may each be derived based on specific information such as coding parameters. For example, the coding parameters may include coding parameters for target / neighboring blocks.
[0607] In embodiments, information related to transformation, such as transformation selection information, and sub-information of the transformation selection information may be signaled for a specific target. For example, the specific target may be a CU.
[0608] Information related to transformation, such as transformation selection information, and sub-information of transformation selection information can be derived for a specific target. For example, the specific target may be a CU.
[0609] Quantized levels can be generated by performing quantization on the result or residual block generated by performing the first transform and / or the second transform.
[0610] The description of the transformation described above can also be applied to the inverse transformation. In this application, the reverse processing of the processing described for the transformation can be performed in the inverse transformation. The term "transformation" in the name related to the transformation can be changed to "inverse transformation." Furthermore, the input of the transformation can be considered the output of the inverse transformation. The output of the transformation can be considered the input of the inverse transformation. The decoding device (150) can obtain information related to the transformation, such as transformation selection information, and can perform the reverse processing of the processing related to the transformation indicated by the information related to the transformation using the information related to the transformation.
[0611] A target block may include multiple subblocks. Each subblock may be defined according to a minimum block size or a minimum block shape. The target block may be divided into multiple subblocks, and each subblock may include coefficients of sizes such as 4x4, 2x8, and 8x2. The target block may be a transform block. Transform coefficients or quantized levels may be expressed in the form of a block. The transform coefficients may be quantized transform coefficients.
[0612] The transform coefficients or quantized levels can be scanned according to at least one of the scanning types, such as diagonal scanning, vertical scanning, and horizontal scanning. The diagonal scanning can be right-upper diagonal scanning or left-lower diagonal scanning.
[0613] For example, coefficients can be transformed or arranged into a one-dimensional vector by scanning the coefficients of a block using diagonal scanning. Vertical scanning can scan coefficients in the form of two-dimensional blocks in the column direction. Horizontal scanning can scan coefficients in the form of two-dimensional blocks in the row direction.
[0614] The scanning type for coefficients can be determined based on coding parameters such as the intra prediction mode, block size, and block shape. For example, whether diagonal scanning, vertical scanning, or horizontal scanning is used can be determined based on coding parameters such as the intra prediction mode, block size, and block shape. A block can be a transform unit.
[0615] Scanning for each scanning type can start at a specific starting point and end at a specific ending point.
[0616] In scanning, a scanning order based on the scanning type may first be applied between subblocks. Next, a scanning order based on the scanning type may be applied to transform coefficients or quantized levels within the subblock.
[0617] The encoding device (110) can perform entropy encoding on transform coefficients or quantized levels to generate a bitstream including entropy-encoded transform coefficients or entropy-encoded quantized levels.
[0618] The decoding device (150) can obtain entropy-encoded transform coefficients or entropy-encoded quantized levels from a bitstream and perform entropy decoding to generate transform coefficients or quantized levels. The coefficients can be arranged in the form of two-dimensional blocks through inverse scanning. The arrangement of the inverse scanning can be a rearrangement opposite to the arrangement of the scanning.
[0619] Inverse scanning of coefficients can generate inversely scanned transform coefficients or inversely scanned quantized levels. At this time, the inverse scanning types of the inverse scanning can include diagonal scanning, vertical scanning, and horizontal scanning, and the inverse scanning type of the inverse transformation corresponding to the scanning type of the transformation can be selected.
[0620] In the decoding device (150), inverse quantization can be performed on (inversely scanned) coefficients. Depending on whether a second inverse transform is performed, a second inverse transform can be performed on the result generated by performing inverse quantization. In addition, depending on whether a first inverse transform is performed, a first inverse transform can be performed on the result generated by performing the second inverse transform. A restored residual block can be generated by selectively performing the second inverse transform and the first inverse transform on the coefficients.
[0621]
[0622] Filtering
[0623] To improve the image quality, filtering may be performed on blocks. The values of target samples may be determined or updated through filtering.
[0624] The target sample may be one of the samples described in the embodiments. For example, the target sample may be one or more of the samples described in the embodiments, such as a prediction sample, a reference sample, a residual sample, a restored sample, and a restored sample with filtering applied.
[0625] The target sample may be a sample within one or more of a target picture, a target slice, a target CTB, a target block, a reference sample line, and a template. The target block may be one of the blocks described in the embodiments. For example, the target block may be one or more of the blocks described in the embodiments, such as a transform block, a prediction block, a reference block, a residual block, and a reconstruction block.
[0626] In embodiments, the filtering process described as being applied to one object may also be applied to other objects. For example, the filtering process described in a specific in-loop filtering may also be applied to transform blocks, prediction blocks, reference blocks, and residual blocks.
[0627] A specific type of filtering may be used for the filters of the embodiments. The type of filtering may include filter taps (or filter tap lengths), filter shapes, filter strengths, filter coefficients (or weights), and offsets.
[0628] The filter tab may indicate the number of input samples used for the filter. The input samples may include the target sample. Alternatively, the input samples may include a specific value determined for the target sample. The input samples may include one or more reference samples. The one or more reference samples may be determined based on an attribute of the target block described in the embodiments. The attribute may include a coding parameter. For example, an attribute of the target sample may include a position of the target sample. One or more reference samples may be specified based on a relative position with respect to the position of the target sample.
[0629] A filter shape can represent the shape formed by input samples. A specific value determined for a target sample can be considered a target sample. In other words, if a specific value determined for a target sample is used as an input sample of a filter, the target sample can also be considered to form a filter shape.
[0630] The number of samples whose values are determined by filtering may be multiple. The filter strength may indicate the range of samples whose values are determined by filtering. The filter strength may be either a strong filtering strength or a weak filtering strength. The number of samples whose values are determined by a strong filtering strength may be greater than the number of samples whose values are determined by a weak filtering strength. Alternatively, the filter strength may indicate the range of values that are changed by filtering. The range of sample values that are changed by a strong filtering strength may be wider than the range of sample values that are changed by a weak filtering strength.
[0631] The filter coefficients can be coefficients or weights of the input samples.
[0632] An offset can be a specific value that is added to the result calculated using the values and coefficients of the input samples, such as a weighted sum.
[0633] Filtering, interpolation, and sampling may have in common that they update the values of samples. Therefore, the description of any one of filtering, interpolation, and sampling in the embodiments may also apply to any other of filtering, interpolation, and sampling. Here, sampling may include at least one of upsampling, downsampling, and subsampling.
[0634] Filtering may include filtering performed by predictor (123) and predictor (163), etc.
[0635] In encoding a target block, a prediction error may exist between the original samples of the original block and the prediction samples of the prediction block. To reduce the prediction error, filtering may be performed on at least one of the prediction samples of the prediction block and the reference samples referenced for prediction.
[0636] For example, in intra prediction, the reference sample may include one or more of the upper left reference sample, the upper reference sample, the upper right reference sample, the left reference sample, and the lower left reference sample. Filtering on the predicted sample may be performed by applying specific weights to the predicted sample, the left reference sample, the upper reference sample, and / or the upper left reference sample, respectively.
[0637] Filtering of at least one of the prediction sample and the reference sample may be performed based on the attributes of the target block and the attributes of the prediction sample. For example, whether filtering is performed, the type of filter, the area to which the filtering is applied, the filtering weights, the reference sample, the range of the reference sample, and the location of the reference sample may each be determined based on the attributes of the target block and the attributes of the prediction sample.
[0638] For example, the properties of the target block may include information related to the target block described in the embodiments, such as 1) size of the target block, 2) prediction mode, 3) intra prediction mode, 4) reference sample line, 5) sample value, and 6) coding parameter.
[0639] For example, the attributes of a prediction sample may include information related to the prediction sample described in the embodiments, such as 1) a sample value of the prediction sample and 2) a location within a target block, and may include coding parameters related to the prediction sample.
[0640] Filtering may include in-loop filtering performed by filter (130) and filter (170), etc.
[0641]
[0642] Figure 6 illustrates multiple in-loop filters according to an example.
[0643] The plurality of in-loop filters of the in-loop filtering may include one or more of Luma Mapping with Chroma Scaling (LMCS), a deblocking filter, a Sample Adaptive Offset (SAO), and an Adaptive Loop Filter (ALF).
[0644] Multiple in-loop filters can be connected sequentially. For example, the multiple in-loop filters can be connected in the order of LMCS, deblocking filter, SAO, and ALF. Furthermore, the multiple in-loop filters can be connected in any order among all available permutations of the multiple in-loop filters. The output from one of the multiple in-loop filters can be used as the input to the next filter.
[0645] As illustrated in FIG. 6, an input image may be input to the first filter. The input image may be a block described in the embodiments. For example, the input image may be a reconstructed block generated by an adder (129) or an adder (169). The output from one filter may be input to the next filter. An output image may be generated by the last filter. The output image may be a filtered block described in the embodiments. For example, the output image may be a filtered reconstructed image generated by a filter (130) or a filter (170).
[0646] The target block can represent an image input to the filter. The filtered target block can represent an image output from the filter.
[0647] LMCS may include luma signal mapping to a luma signal of a target block and chroma signal scaling to a chroma signal of the target block.
[0648] Luma signal mapping can perform codeword redistribution for the luma signal.
[0649] Luma signal mapping can include forward mapping and reverse mapping. In forward mapping, the existing dynamic range can be divided into multiple intervals. The mapped dynamic range can be determined by performing codeword redistribution on the input image using a linear model for each interval. In reverse mapping, reverse mapping is performed from the mapped dynamic range to the existing dynamic range.
[0650] Chroma scaling can correct chroma signals based on the correlation between a luma signal and a corresponding chroma signal.
[0651] Forward mapping can be performed between inter prediction for a luma signal and reconstruction for the luma signal, and between inter prediction for the luma signal and chroma scaling. Backward mapping can be performed between reconstruction for the luma signal and in-loop filtering for the luma signal. Chroma scaling can be performed between inverse transformation and reconstruction for the chroma signal.
[0652] According to this structure, inverse quantizations for luma and chroma signals, inverse transformations for luma and chroma signals, prediction for luma signals, and restoration for luma signals can be performed within the mapped dynamic range. In-loop filterings for luma and chroma signals, inter predictions for luma and chroma signals, intra prediction for chroma signals, and restoration for chroma signals can be performed within the existing dynamic range.
[0653] A deblocking filter can remove block distortion occurring at boundaries between blocks within a restored image. For example, the blocks may be transform blocks. Furthermore, the blocks may be subblocks of a specific block described in the embodiments. Here, the boundaries between blocks may refer to samples adjacent to the boundaries between blocks.
[0654] Deblocking filters can be applied to vertical and horizontal boundaries between blocks. After filtering the vertical boundaries of blocks, filtering can be performed again on the horizontal boundaries of the filtered blocks.
[0655] A deblocking filter may be applied selectively. Whether to apply a deblocking filter to a target block may be determined based on at least one of the sample(s) contained within a specific number of columns or rows within the target block and the sample(s) contained within a specific number of columns or rows within a neighboring block adjacent to a specific boundary.
[0656] When a deblocking filter is applied to a target block, the filter to be applied may be determined based on the strength of the required deblocking filtering. In other words, among multiple other filters, a filter determined based on the strength of the deblocking filtering may be applied to the target block. The multiple filters may include one of a long-tap filter, a strong filter, a weak filter, and a Gaussian filter.
[0657] The maximum length of the deblocking filter can be determined based on the properties of the target block, such as the size of the target block, the components of the target block, and the coding parameters.
[0658] SAO can compensate for distortion between the original and reconstructed images on a sample-by-sample basis. To compensate, SAO can apply an appropriate offset to the sample values of each sample. That is, the offset can be added to the sample values.
[0659] An offset can be determined for the target block. For example, an offset can be determined for each component of the CTB. The determined offset can be applied to samples within a specific component of the CTB.
[0660] SAO may include SAO using Edge Offset (EO) and SAO using Band Offset (BO). Depending on the characteristics of samples within a specific block, such as a CTU, whether SAO using EO or SAO using BO may be performed may be determined.
[0661] In SAO using EO, distortion correction of samples can be performed based on the direction of the edge within the target block. Pattern classes of EO can include horizontal patterns, vertical patterns, 135 degree diagonal patterns, and 45 degree diagonal patterns. For a target block, information indicating a pattern class applied to the target block and multiple offsets of the pattern class can be signaled. There can be four offsets. For a target sample within the target block, adjacent samples of the target sample can be determined based on the direction of the pattern class. An offset to be applied to the target sample can be determined based on the pattern of the adjacent samples.
[0662] In an offset using BO, distortion of a sample can be corrected by classifying the brightness values of samples within a target block into specific bands. The bit depth of an input image can be divided into m sections. For example, m can be 32. The specific bands can be n consecutive sections among the m sections. For example, n can be 4. N offsets for the n sections can be signaled. Additionally, information indicating a first section selected as one of the n sections among the m sections can be signaled. The offset of the section to which the target sample corresponds can be added to the sample value of the target sample of the target unit.
[0663] ALF can compensate for distortion between the restored image and the original image.
[0664] The filter coefficients of ALF can be signaled via the bitstream.
[0665] The filter shape of ALF can be determined by the components of the target block. For example, a 7x7 diamond-shaped filter can be used for the luma component. A 5x5 diamond-shaped filter can be used for the chroma component.
[0666] In ALF, the characteristics of a specific block can be determined for a specific block, and the class of the specific block can be determined based on the characteristics. In other words, the determination of characteristics and class of ALF can be performed in units of 4x4 blocks. Filter coefficients can be calculated based on the class. A specific block can be a 4x4 block.
[0667] One of 25 classes can be determined as the class of a specific block based on the direction and activity determined using the gradient of the specific block. Rotation, vertical symmetry, and / or diagonal symmetry transformations can be applied to the filter based on the gradient of the specific block.
[0668] Information regarding whether ALF applies can be signaled for specific units, such as CTB.
[0669] An index indicating a filter to be applied to a specific unit among available filters may be signaled. Here, the available filters may include fixed filters and filters configured using a parameter set. For example, the parameter set may be an adaptive parameter set (APS). The fixed filters may be identically predefined in the encoding device (110) and the decoding device (150). The filter coefficients of the filters configured using the parameter set may be determined based on coding parameters.
[0670]
[0671] Entropy encoding and entropy decoding
[0672] Figure 7 illustrates entropy encoding and entropy decoding according to an example.
[0673] The processes of entropy encoding by the entropy encoder (139) are illustrated at the top of Fig. 7.
[0674] The entropy encoder (139) may include a context modeler, a binarization unit, and an entropy encoder. The context modeler may include a context selection unit and a context memory.
[0675] The binarization unit can generate bins for syntactic elements by performing binarization on the syntactic elements of the target block. Binarization may be a process of converting syntactic elements into the form of bins.
[0676] Information about syntactic elements and bins can be provided from the binarization unit to the context selection unit.
[0677] A context modeler can perform context updates.
[0678] Context can mean occurrence probability information for each bin for syntactic elements that have already been encoded.
[0679] The context modeler can update the context to apply current probability information to the entropy encoding of the bins of the syntactic elements of the target block. The updated context can be stored in the context memory. At this time, the updated context corresponding to the syntactic elements of the target block (or bins within the syntactic elements of the target block) can be derived by the context modeler.
[0680] The context selector can select a context corresponding to a bin of a syntactic element of a target block. The selected context can be loaded from the context memory and used as an updated context for entropy encoding of the bins of the syntactic element of the target block.
[0681] The updated context can be used for entropy encoding of syntactic elements of the target block.
[0682] The entropy encoding unit can generate encoded information about syntactic elements of a target block by performing entropy encoding using the generated bins and the updated context, and can generate a bitstream including the encoded information. The entropy encoding unit can use at least one of an arithmetic encoding method and a bypass encoding method.
[0683] The processes of entropy decryption by the entropy decoder (161) are shown at the bottom of Fig. 7.
[0684] The entropy decoder (161) may include a context modeler, an entropy decoder, and an inverse binarizer. The context modeler may include a context selection unit and a context memory.
[0685] A context modeler can perform context updates.
[0686] Context can mean the occurrence probability information of each bin for syntactic elements that have already been decoded.
[0687] The context modeler can update the context to apply the currently decoded probability information to entropy decoding for the bins of the syntactic elements of the target block. The updated context can be stored in the context memory. At this time, the updated context corresponding to the syntactic elements of the target block (or the bins within the syntactic elements of the target block) can be derived by the context modeler.
[0688] The context selector can select a context corresponding to a blank of a syntactic element of a target block. The selected context can be loaded from the context memory and used as an updated context for entropy decoding of the syntactic element of the target block.
[0689] The updated context can be used for entropy decoding of syntactic elements of the target block.
[0690] The entropy decoding unit can generate bins for the delimiting elements of the target block by performing entropy decoding on the encoded information of the bitstream based on the updated context. The entropy decoding unit can use at least one of an arithmetic decoding method and a bypass decoding method.
[0691] The debinarization unit can obtain a syntactic element of the target block by performing debinarization on at least one of the generated bins. The debinarization may be a process of converting at least one of the bins into a form of a syntactic element.
[0692] Information about syntactic elements and bins can be provided from the de-binarization unit to the context selection unit.
[0693] A syntax element may be one of the coding parameters described in the embodiments.
[0694]
[0695] Methods for binarization, debinarization, entropy encoding, and entropy decoding
[0696] In embodiments, one or more of the binarization methods, inverse binarization methods, entropy encoding methods and entropy decoding methods listed below may be used to perform signaling for specific information.
[0697] - Signed 0-th order Exponential Golomb binarization / debinarization method (abbreviated as se(v))
[0698] - k-order exponential-Golomb binarization / inverse binarization method with sign (abbreviated as sek(v))
[0699] - 0-order exponent-Golomb binarization / inverse binarization method for unsigned positive integers (abbreviated as ue(v))
[0700] - k-order exponential-Golomb binarization / inverse binarization method for unsigned positive integers (abbreviated as uek(v))
[0701] - Fixed-length binarization / debinarization method (abbreviated as f(n))
[0702] - Truncated Rice binarization / debinarization method or truncated unary binarization / debinarization method (abbreviated as tu(v))
[0703] - Truncated binary binarization / debinarization method (abbreviated as tb(v))
[0704] - Context-adaptive arithmetic encoding / decoding method (abbreviated as ae(v))
[0705] - bit string in bytes (abbreviated as b(8))
[0706] - Signed integer binarization / debinarization method (abbreviated as i(n))
[0707] - Unsigned positive integer binarization / debinarization method (abbreviated as u(n)) ('u(n)' can also mean fixed-length binarization / debinarization method.)
[0708] - Unary binarization / inverse binarization method
[0709]
[0710] Adaptive loop filter
[0711] Adaptive loop filters can directly reduce the error between samples of the original image and samples of the reconstructed image.
[0712] For example, the error may be an error described in embodiments, such as the Mean of Squared Error (MSE). Alternatively, the adaptive loop filter may reduce the error between samples of the original image and samples of another image described in embodiments.
[0713] Adaptive loop filters can effectively reduce these errors, but they may have the following disadvantages: 1) the complexity of calculating filter coefficients for filtering is high, 2) a lot of memory is consumed for calculation, and 3) a large number of bits are used for signaling the filter coefficients.
[0714] In the embodiments, an in-loop filtering method utilizing subsample-based block classification may be described. Such an in-loop filtering method may improve image encoding efficiency, reduce computational complexity, and reduce memory access bandwidth.
[0715] In the embodiments, an in-loop filtering method utilizing multiple filter shapes may be described. Such an in-loop filtering method may improve image encoding efficiency, reduce computational complexity, and reduce memory access bandwidth.
[0716]
[0717] Filtering in Examples
[0718] In embodiments, the in-loop filtering may include at least one of deblocking filtering, sample adaptive offset (SAO), bilateral filtering, and adaptive in-loop filtering. Furthermore, the description of filtering for a specific object described in one embodiment may also apply to other filtering for other objects in other embodiments. For example, a restored image may be replaced with another image or block in the embodiment.
[0719] A reconstructed block is generated by combining a prediction block generated by inter / intra prediction and a reconstructed residual block, and a reconstructed image can be constructed using the generated reconstructed block. At least one of deblocking filtering and sample adaptive offset can be performed on the reconstructed image. By performing these operations, blocking artifacts and ringing artifacts present in the reconstructed image can be effectively reduced.
[0720] Deblocking filters can reduce blocking artifacts that occur at the boundaries between blocks by performing vertical and horizontal filtering on the block boundaries. However, deblocking filters may not minimize distortion between the original and reconstructed images due to this block boundary filtering.
[0721] Additionally, sample-adaptive offset can add an offset to a specific pixel or to pixels whose pixel values fall within a certain range by comparing adjacent pixels on a pixel-by-pixel basis to reduce ringing. Sample-adaptive offset can partially minimize distortion between the original and reconstructed images by utilizing rate-distortion optimization. However, sample-adaptive offset may have limitations in terms of distortion minimization when the difference in distortion between the original and reconstructed images is large.
[0722] In bilateral filtering, filter coefficients can be determined based on 1) the distances between the center sample and other samples within the region where filtering is performed, and 2) the differences between the values of the samples. At this time, the filter coefficients can be determined using other coding parameters of the embodiment. At this time, the distance between samples can be replaced with other calculation formulas for the samples described in the embodiments, and can be used together with other calculation formulas described in the embodiments.
[0723] Adaptive in-loop filtering can apply filters to restored images that minimize distortion between the original and restored images. By applying these filters, distortion between the original and restored images can be minimized.
[0724] In embodiments, in-loop filtering may mean adaptive in-loop filtering.
[0725]
[0726] Target of filtering
[0727] In embodiments, filtering may mean that a filter is applied to at least one of the units described in the embodiments. For example, the units may include a sample, a block, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a coding tree block (CTB), a slice, a tile, an image, and a sequence.
[0728] In embodiments, filtering may include at least one of 1) classification of blocks, 2) application of filters, and 3) encoding / decoding of filter information, as described in the embodiments. Furthermore, filtering for a specific object in the embodiments may mean applying a specific process described in relation to filtering to the specific object.
[0729] In-loop filtering may refer to applying multiple filters to a reconstructed image in a specific order. A decoded image may be generated by the filters in this specific order. For example, the specific order may be the order of bilateral filtering, deblocking filtering, sample adaptive offset, and adaptive in-loop filtering. Alternatively, the specific order may be the order in which specific filters (included in in-loop filtering) are described in one embodiment. Furthermore, the order between specific filters may be changed, and a decoded image may be generated by applying the specific filters in the changed order to the reconstructed image. Alternatively, multiple filters may be performed in parallel, and then the results of each operation may be combined to derive the combined result. The result may refer to a reconstructed / decoded image.
[0730] For example, in-loop filtering can be applied to the restored image in the following order: deblocking filtering, sample adaptive offset, and adaptive in-loop filtering.
[0731] For example, in-loop filtering can be applied to the restored image in the following order: bidirectional filtering, adaptive in-loop filtering, deblocking filtering, and sample adaptive offset.
[0732] For example, in-loop filtering can be applied to the restored image according to adaptive in-loop filtering, deblocking filtering, and sample adaptive offset order.
[0733] For example, in-loop filtering can be applied to the restored image in the following order: adaptive in-loop filtering, sample adaptive offset, and deblocking filtering.
[0734] A reconstructed block may be generated by adding a reconstructed residual block to a prediction block generated by intra-prediction, inter-prediction, and / or other prediction of the embodiment. A decoded image may refer to a result image generated by performing in-loop filtering or post-processing filtering on a reconstructed image composed of at least one reconstructed block.
[0735] A decoded image can be generated by performing adaptive in-loop filtering on a restored image. Alternatively, the adaptive in-loop filtering can be performed on a restored image to which at least one of deblocking filtering, sample adaptive offset, and bilateral filtering has been applied. Furthermore, the adaptive in-loop filtering can be performed on the restored image to which the adaptive in-loop filtering has been applied. In this case, the adaptive in-loop filtering can be repeatedly applied to the restored image N times. N can be one of positive integers.
[0736] In-loop filtering may be performed on a restored image to which at least one filtering method among the in-loop filtering methods of the embodiments has been applied. For example, when at least one other filtering method among the in-loop filtering methods is performed on a restored image to which at least one of the in-loop filtering methods of the embodiments has been applied, a parameter used in at least one of the other filtering methods may be changed, and in-loop filtering with the changed parameter may be performed on the restored image.
[0737]
[0738] Filter parameters
[0739] In embodiments, the parameters of the filtering may include one or more of the coding parameters, filter coefficients, number of filter taps (i.e., filter length), filter shape, filter type, number of times filtering is performed, filter strength, and threshold value described in the embodiments.
[0740] The filter coefficient may refer to a coefficient that constitutes a filter, and may refer to a value corresponding to a specific mask position within the form of a mask that is multiplied to a restored sample.
[0741] The number of filter taps can indicate the length of the filter.
[0742] A filter may have symmetrical properties with respect to a specific direction. If a filter has symmetrical properties with respect to a specific direction, the length of the filter may be reduced by half.
[0743] The filter tab can refer to the width of the filter in the horizontal direction and / or the height of the filter in the vertical direction.
[0744] A filter can have symmetrical characteristics with respect to two specific directions.
[0745] When the filter has a mask shape, the filter shape may represent a geometric shape that can be expressed in two dimensions or a combination of such geometric shapes.
[0746] For example, geometric shapes may include rectangles, squares, trapezoids, triangles, pentagons, hexagons, octagons, decagons, dodecagons, diagonal snowflakes, sharps, and clovers.
[0747] Additionally, the filter shape may be a shape formed by projecting a three-dimensional filter into a two-dimensional shape, and may include a shape formed by projection.
[0748] Filter types may include a Wiener filter, a low-pass filter, a high-pass filter, a linear filter, a non-linear filter, and a bilateral filter. The type of the filter of the embodiment may be at least one of the above filter types.
[0749] As a filter type for adaptive in-loop filtering, a Wiener filter can be used. The Wiener filter is an optimal linear filter that can effectively remove noise, blurring, and distortion within an image, thereby improving encoding efficiency. The Wiener filter can be designed to minimize distortion between the original and restored images.
[0750]
[0751] Filtering in encoding / decoding
[0752] At least one of the filtering methods in the embodiments may be performed within the encoding process or the decoding process.
[0753] The encoding process or decoding process may mean that encoding or decoding is performed on a specific unit described in the embodiments, such as a slice, tile, picture, or sequence.
[0754] For example, a Wiener filter can be performed within the encoding process or the decoding process in the form of adaptive in-loop filtering.
[0755] That is, in adaptive in-loop filtering, “in-loop” can mean that filtering is performed within the encoding process or the decoding process.
[0756] When adaptive in-loop filtering is performed, the restored image on which adaptive in-loop filtering has been performed can be used as a reference image for an image to be encoded / decoded later. At this time, in the image to be encoded / decoded later, inter prediction and motion compensation, etc. can be performed with reference to the restored image on which adaptive in-loop filtering has been performed. Therefore, adaptive in-loop filtering can improve not only the encoding efficiency of the restored image on which adaptive in-loop filtering has been performed, but also the encoding efficiency of the image to be encoded / decoded later.
[0757] At least one of the filtering methods of the embodiments may be performed within an encoding process or a decoding process for a specific unit described in the embodiments, such as a CTU and a block.
[0758] For example, a Wiener filter may be performed within the encoding process or decoding process for a specific unit, such as a CTU or block, as described in the embodiments in the form of adaptive in-loop filtering.
[0759] That is, in adaptive in-loop filtering, "in-loop" may mean that filtering is performed within an encoding process or a decoding process for a specific unit described in embodiments such as a CTU and a block. When adaptive in-loop filtering is performed on a specific unit described in embodiments such as a CTU and a block, the restoration unit on which adaptive in-loop filtering is performed may be used as a reference unit for a unit to be encoded / decoded later. At this time, intra prediction, etc. may be performed on the unit to be encoded / decoded later with reference to the restoration unit on which adaptive in-loop filtering is performed. Therefore, adaptive in-loop filtering can improve not only the encoding efficiency of the restoration unit on which adaptive in-loop filtering is performed, but also the encoding efficiency of the unit to be encoded / decoded later.
[0760] Additionally, at least one of the filtering methods of the embodiments may be performed after the decryption process in the form of a post-processing filter.
[0761] A post-processing filter may be a filter that is not used when a particular block is stored as a reference image, but is applied before a particular block is output / displayed.
[0762] For example, a Wiener filter may be performed after the decoding process as a post-processing filter. If the Wiener filter is performed after the decoding process, after a decoded block is generated by decoding and the decoded block is stored as a reference image, the Wiener filter may be performed on the restored image before the output / display of the restored image. If post-processing filtering is performed, the restored image on which the post-processing filtering has been performed may not be used as a reference image for an image to be encoded / decoded later.
[0763] Adaptive in-loop filtering may utilize block-based filter adaptation. In embodiments, block-based filter adaptation may mean that, for each block, the filter to be used for the block is adaptively selected from among multiple filters.
[0764]
[0765] Filtering based on block classification
[0766] Figure 8 is a flowchart of an encoding method according to one embodiment.
[0767] The encoding method may include a block classification step (810) for performing classification on a block; a filtering step (820) for performing filtering on a block based on the classification; a filter information encoding step (830) for performing encoding on filtering information indicating information on filtering; and a bitstream transmission step (840) for transmitting a bitstream.
[0768] The bitstream may contain encoded filter information or filter information.
[0769] The encoding device (110), the processor (120) of the encoding device (110) and / or the filter (130) of the encoding device (110) can perform steps (810, 820 and 830).
[0770] The encoding device (110) and / or the communication device (149) of the encoding device (110) can perform step (840).
[0771]
[0772] Figure 9 is a flowchart of a decryption method according to one embodiment.
[0773] The decoding method may include a bitstream receiving step (910) of receiving a bitstream; a filter information decoding step (920) of performing decoding on encoded filter information within the bitstream; a block classification step (930) of performing classification on a block based on the filter information; and a filtering performing step (940) of performing filtering on a block based on the classification.
[0774] The decryption device (150) and / or the communication device (189) of the decryption device (150) can perform step (910).
[0775] The bitstream may contain encoded filter information or filter information.
[0776] The decryption device (150), the processor (160) of the decryption device (150) and / or the filter (170) of the decryption device (150) can perform steps (920, 930 and 940).
[0777] To avoid redundant description, in the embodiments, the steps performed in the encoding device (110) and the decoding device (150) may be described using a single name as follows.
[0778] - The encoding device or encoder may mean a decoding device (110).
[0779] - A decryption device or decoder may mean a decryption device (150).
[0780] - The block classification step may refer to step (810) and / or step (930). The description related to block classification in the embodiments may be considered to be performed in the block classification step.
[0781] - The filtering execution step may refer to step (820) and / or step (940). The description related to the performance of filtering in the embodiments may be considered to be performed in the filtering execution step.
[0782] - The filter information encoding / decoding step may refer to step (830) and / or step (920). The description related to the encoding / decoding of filter information in the embodiments may be considered to be performed in the filter information encoding / decoding step.
[0783] - The bitstream transmission / reception step may refer to step (840) and / or step (910). The description related to bitstream transmission / reception in the embodiments may be considered to be performed in the bitstream transmission / reception step.
[0784] Adaptive in-loop filtering may include a block classification step; a filtering performance step; and a filtering information encoding / decoding step.
[0785] Adaptive in-loop filtering may further include a filter coefficient derivation step; a filtering performance determination step; and a filter shape determination step. For example, the filter coefficient derivation step; the filtering performance determination step; and the filter shape determination step may be performed before the filtering performance step described in the embodiments is performed. Alternatively, when the method described in the embodiments includes the filtering performance step, the method may further include a filter coefficient derivation step; a filtering performance determination step; and a filter shape determination step.
[0786] The description related to the derivation of filter coefficients in the embodiments may be considered to be performed in the filter coefficient derivation step.
[0787] The description related to the decision on whether to perform filtering in the embodiments may be considered to be performed in the filtering performance decision step.
[0788] The description related to the determination of the filter shape in the embodiments may be considered to be performed in the filter shape determination step.
[0789] Below, the filter coefficient derivation step; the filtering performance decision step; and the filter shape decision step are described, respectively.
[0790]
[0791] Filter coefficient derivation step
[0792] In the filter coefficient derivation step, Wiener filter coefficients can be derived from the perspective of minimizing distortion between the original image and the filtered image.
[0793] Wiener filter coefficients can be derived for each block classification index.
[0794] Additionally, the Wiener filter coefficients may be derived based on at least one piece of filter-related information. For example, the filter-related information may include filter taps and filter shapes.
[0795] When the Wiener filter coefficients are derived, an auto-correlation function for the reconstructed sample; a cross-correlation function for the original sample and the reconstructed sample; an auto-correlation matrix; and a cross-correlation matrix can be derived.
[0796] The Wiener-Hopf equation can be derived using the derived autocorrelation matrix and cross-correlation matrix, and the filter coefficients can be calculated using the Wiener-Hopf equation. Here, the filter coefficients can be calculated using the Gaussian elimination method or the Cholesky decomposition method in the Wiener-Hopf equation.
[0797]
[0798] Decision stage for performing filtering
[0799] In the filtering execution decision step, it may be determined whether adaptive in-loop filtering is to be applied to a specific unit as described in the embodiments. Whether adaptive in-loop filtering is to be applied to a specific unit may be determined from a rate-distortion optimization perspective. Alternatively, whether adaptive in-loop filtering is to be applied to a specific unit may be determined based on filter information.
[0800] For example, a particular unit may include one or more of a slice, a picture, a block, and a CU.
[0801] In determining the rate, filter information to be encoded may be included.
[0802] Distortion can be a value for the difference between the original image and the reconstructed image, or the difference between the original image and the filtered reconstructed image. To calculate the distortion, formulas for deriving the difference between objects described in the embodiments, such as the Mean of Squared Error (MSE), the Sum of Squared Error (SSE), and the Sum of Absolute Difference (SAD), can be used.
[0803] In the filtering performance decision step, it can be determined whether filtering is performed on the chroma component as well as the luma component.
[0804]
[0805] Filter shape determination step
[0806] In the filter shape determination step, when applying adaptive in-loop filtering, the filter shape and number of filter taps of the filter can be determined.
[0807] For example, in the filter type determination step, a filter type to be used can be determined from among a plurality of available filter types, and the number of filter taps to be used for filtering can be determined from among a plurality of available filter tap numbers.
[0808] Here, the filter shape and the number of filter taps can be determined from a rate-distortion optimization perspective. Alternatively, the filter shape and the number of filter taps can be indicated by filter information.
[0809] Below, the operations performed at each step in the encoding device and / or decoding device are described in more detail.
[0810]
[0811] Block classification step
[0812] By classifying the blocks, a block classification index C can be assigned to each block within the image.
[0813] Here, the image may be one or more of the images described in the embodiments, such as a restored image and a difference image, and may mean samples within the image.
[0814] Block classification indices C can be assigned to each unit of NxM blocks within an image. In other words, a unit to which a block classification index C is assigned can be an NxM block within an image.
[0815] In embodiments, a "block classification unit" may represent a unit to which a block classification index C is assigned. For example, a "block classification unit" may be an NxM block within an image.
[0816] Blocks can be classified into L types. The block classification index C of each block can have one of L values.
[0817] Here, each of N, M and L can be a positive integer.
[0818] For example, each of N and M can be one of 2, 4, 8, 16, and 32. L can be one of 4, 8, 16, 20, 24, 25, and 32.
[0819] If N is 1 and M is 1, sample classification can be performed on a sample-by-sample basis rather than block classification on a block-by-block basis.
[0820] The values of N and M can be different. For example, an NxM block can have a non-square shape. Alternatively, N and M can be the same.
[0821] For example, in a restored image, a total of 25 block classification indices can be assigned to each of the 2x2 block units.
[0822] For example, in a restored image, a total of 25 block classification indices can be assigned to each of the 4x4 block units.
[0823] The value of the block classification index C can range from 0 to L-1, or from 1 to L.
[0824] The blocks may have other geometric shapes than those described in the embodiments, other than rectangular ones.
[0825] A block may be determined by the size of its subunits, which are determined by the division of a particular unit as described in the embodiments.
[0826] The block classification index C can be classified into units of NxN blocks.
[0827] The block classification index C can be derived based on at least one of the sum of the absolute values of the frequency coefficients calculated by a 2D transform of KxK; the median value; and the mean value.
[0828] Alternatively, the block classification index C of a block may be determined based on the statistical values of coefficients derived by the transformation for the block of the embodiments. Here, the statistical value may be one of the statistical values described in the embodiments, such as the sum of absolute values; the median; and the mean.
[0829] In the embodiments, the mean may mean a value determined by a calculation formula using multiple values described in the embodiments, such as the arithmetic mean; the geometric mean; the harmonic mean; and the median.
[0830] Here, N can be a positive integer. For example, N can be either 2 or 4.
[0831] Here, K can be any positive integer. For example, K can be one of 2, 4, 6, and 8.
[0832] In embodiments, the transform may include one or more of the transforms described in the embodiments, such as the Discrete Cosine Transform (DCT), the Discrete Sine Transform (DST), and the Hadamard Transform, or may be or use other transforms.
[0833] In embodiments, the transformation may mean a frequency transformation. Alternatively, “frequency transformation” may be abbreviated as “transformation.”
[0834] [Formula 1] below is the frequency conversion coefficient k ij A first example of the basic coefficients used for derivation can be presented.
[0835] [Formula 1]
[0836] a0= {84, 47, 55, 29},
[0837] a1= {74, 0, -74, -74},
[0838] a2= {55, -74, -29, 84},
[0839] a3= {29, -74, 84, -55}
[0840] [Formula 2] below is the frequency conversion coefficient k ij A second example of the basic coefficients used for derivation can be presented.
[0841] [Formula 2]
[0842] a0= {29, 55, 74, 84},
[0843] a1= {74, 74, 0, -74},
[0844] a2= {84, -29, -74, 55},
[0845] a3= {55, -84, 74, -29}
[0846] [Formula 3] below is the frequency conversion coefficient k ij A third example of the basic coefficients used for derivation can be presented.
[0847] [Formula 3]
[0848] a0= {64, 64, 64, 64},
[0849] a1= {83, 36, -36, -83},
[0850] a2= {64, -64, -64, 64},
[0851] a3= {36, -83, 83, -36}
[0852] [Formula 4] below is the frequency conversion coefficient k ij A fourth example of the basic coefficients used for derivation can be presented.
[0853] [Formula 4]
[0854] a0= {1, 1, 1, 1},
[0855] a1= {1, 1, -1, -1},
[0856] a2= {1, -1, -1, 1},
[0857] a3= {1, -1, 1, -1}
[0858] Frequency transform coefficient k ij can be derived by [Formula 5] below using the basic coefficients of [Formula 1] to [Formula 4].
[0859] [Formula 5]
[0860] k ij = a i × a j
[0861] In [Formula 5], “×” can represent a vector cross product.
[0862] The coefficients of [Formula 1] to [Formula 4] are merely exemplary, and some of the coefficients illustrated in [Formula 1] to [Formula 4] may be used.
[0863] The derived transformation coefficients k ij can be expressed in the form of a KxK matrix. Here, a i may be a vector array for deriving the transformation coefficients. a i [j] can represent the jth element of the vector array of the ith transform coefficient.
[0864] For example, a 4x4 DCT8 transform coefficient can be derived using the coefficients of [Formula 1].
[0865] For example, a 4x4 DST7 transform coefficient can be derived using the coefficients of [Formula 2].
[0866] For example, a 4x4 DCT2 transform coefficient can be derived using the coefficients of [Formula 3].
[0867] For example, a 4x4 Hadamard transform coefficient can be derived using the coefficients of [Equation 4].
[0868]
[0869] Figure 10 can represent DCT8 transform coefficients according to an example.
[0870] For example, the 4x4 2D-DCT8 transform coefficients may be the same as the coefficients illustrated in Fig. 10.
[0871]
[0872] Figure 11 can represent Hadamard transform coefficients according to an example.
[0873] For example, the 4x4 2D-Hadamard transform coefficients may be the same as the coefficients illustrated in Fig. 11.
[0874] The coefficients in FIGS. 10 and 11 are merely exemplary, and some of the coefficients shown in FIGS. 10 and 11 may also be used.
[0875]
[0876] Figure 12 illustrates a 2x2 2D transform for 2x2 block classification according to an example.
[0877] Figure 13 shows the 2x2 2D transformation result coefficients according to an example.
[0878] For example, in 2x2 block classification, as illustrated in Fig. 12, to derive the block classification index C of the 2x2 block, a KxK 2D transform in the same range as the 2x2 block is performed on the surrounding samples of the block, thereby obtaining the 2D transform result coefficient x i,j can be derived.
[0879] Next, T0, T1, T2, and T3 can be calculated according to [Formula 6], [Formula 7], [Formula 8], and [Formula 9] below, respectively. Each of T0, T1, T2, and T3 can be the sum of the absolute values of these coefficients.
[0880] [Formula 6]
[0881]
[0882] [Formula 7]
[0883]
[0884] [Formula 8]
[0885]
[0886] [Formula 9]
[0887]
[0888] In the embodiments, the sum of the absolute values of specific coefficients, the sum of specific coefficients, the average of specific coefficients, and the statistical value of specific coefficients may be substituted for each other. For example, if in the embodiments, specific information is described as being calculated by one of the sum of the absolute values of specific coefficients, the sum of specific coefficients, the average of specific coefficients, and the statistical value of specific coefficients, it may also be understood that it is calculated by the other.
[0889] Here, for each 2x2 block, we use KxK 2D transform to get the 2D transform result coefficient x i,j Each of them can be calculated.
[0890] In the drawings of embodiments such as FIG. 12, an entire rectangle including small rectangles indicated by thin solid lines may represent a block. A small rectangle indicated by thin solid lines may represent a single restored sample or the location of a single restored sample. A rectangle indicated by thick solid lines may represent a range to which a specific process is applied.
[0891] The position of a small square within the entire square may correspond to the position / coordinate within the block of samples corresponding to the small square. Here, the x-coordinate may be one of the values 0 to w-1. The y-coordinate may be one of the values 0 to h-1. w may be the width of the block. h may be the height of the block. Alternatively, the x-coordinate may be one of the values 1 to w. The y-coordinate may be one of the values 1 to h.
[0892] The relative position of a small square within the entire square may correspond to the relative position of the sample corresponding to the small square within the block.
[0893] A full rectangle can represent the entire block. For example, the size of a block can correspond to the number of small rectangles depicted in the drawing. That is, the width of a block can be the number of small rectangles in the horizontal direction. The height of a block can be the number of small rectangles in the vertical direction.
[0894] Alternatively, the entire rectangle may represent a portion of a block. Here, the portion of the block may be a portion of the upper left corner of the block. For example, in the drawing, the samples of the upper left corner of the block may be depicted as small rectangles, while the samples on the right and the samples on the bottom may not be depicted. Alternatively, the portion of the block may be a portion of the upper right corner; the lower left corner; or the lower right corner of the block. If the samples of a portion of the block are depicted and the samples of other portions are omitted, the width and height of the block may each be an integer greater than or equal to 2. Alternatively, the width and height of the block may be considered to be the width and height of another block described in the embodiments.
[0895] In Fig. 12, the rectangle indicated by the bold solid line can represent the range to which the KxK 2D transformation is applied.
[0896] Here, the position corresponding to (i, j) = (0, 0) may mean the upper left position within the range where the KxK 2D transformation is applied.
[0897] After the calculation of the 2D transformation, a specific coefficient x i,j can be changed to 0, and 0 is a specific coefficient x i,j can be used for calculations instead.
[0898] In embodiments, a particular coefficient may mean a portion of the coefficients of the whole.
[0899]
[0900] Figure 14 shows a 4x4 transform for 2x2 block classification according to an example.
[0901] Figure 15 shows the 4x4 2D transformation result coefficients according to an example.
[0902] Figure 16 shows an 8x8 2D transform for 2x2 block classification according to an example.
[0903] Figure 17 shows the 8x8 2D transformation result coefficients according to an example.
[0904] For example, in 2x2 block classification, as shown in FIGS. 14 and 16, a KxK 2D transform with a larger range than the 2x2 block is performed on the surrounding samples of the block to derive the block classification index C of the 2x2 block, thereby obtaining the 2D transform result coefficient x i,j can be derived.
[0905] Next, T0, T1, T2, and T3 can be calculated according to the aforementioned [Formula 6], [Formula 7], [Formula 8], and [Formula 9], respectively. Each of T0, T1, T2, and T3 can be the sum of the absolute values of these coefficients.
[0906] Here, for each 2x2 block, we use KxK 2D transform to get the 2D transform result coefficient x i,j Each of them can be calculated.
[0907] In Figures 14 and 16, the entire rectangle including the small rectangles indicated by the thin solid lines may represent a block. The small rectangles indicated by the thin solid lines may represent one restored sample or the location of one restored sample.
[0908] In Figures 14 and 16, the rectangles indicated by thick solid lines can represent the range to which the KxK 2D transformation is applied.
[0909] Here, the position corresponding to (i, j) = (0, 0) may mean the upper left position within the range where the KxK 2D transformation is applied.
[0910] After the calculation of the 2D transformation, a specific coefficient x i,j can be changed to 0, and 0 is a specific coefficient x i,j can be used for calculations instead.
[0911]
[0912] Figure 18 illustrates a 4x4 2D transform for 4x4 block classification according to an example.
[0913] An example of the 4x4 2D transform result coefficients described above with reference to FIG. 15 can also be applied to the 4x4 2D transform for 4x4 block classification of FIG. 18.
[0914] For example, in 4x4 block classification, as illustrated in Fig. 18, to derive the block classification index C of the 4x4 block, a KxK 2D transform in the same range as the 4x4 block is performed on the surrounding samples of the block, thereby obtaining the 2D transform result coefficient x i,j can be derived.
[0915] Next, T0, T1, T2, and T3 can be calculated according to the aforementioned [Formula 6], [Formula 7], [Formula 8], and [Formula 9], respectively. Each of T0, T1, T2, and T3 can be the sum of the absolute values of these coefficients.
[0916] Here, for each 4x4 block, we use the KxK 2D transform to obtain the 2D transform result coefficient x i,j Each of them can be calculated.
[0917] In Fig. 18, the entire rectangle including the small rectangles indicated by the thin solid lines may represent a block. The small rectangles indicated by the thin solid lines may represent one restored sample or the location of one restored sample.
[0918] In Fig. 18, the rectangle indicated by a thick solid line can represent the range to which the KxK 2D transformation is applied.
[0919] Here, the position corresponding to (i, j) = (0, 0) may mean the upper left position within the range where the KxK 2D transformation is applied.
[0920] After the calculation of the 2D transformation, a specific coefficient x i,j can be changed to 0, and 0 is a specific coefficient x i,j can be used for calculations instead.
[0921]
[0922] Figure 19 shows an 8x8 2D transform for 4x4 block classification according to an example.
[0923] An example of the 8x8 2D transform result coefficients described above with reference to FIG. 17 can also be applied to the 8x8 2D transform for 4x4 block classification of FIG. 19.
[0924] For example, in 4x4 block classification, as illustrated in Fig. 18, in order to derive the block classification index C of the 4x4 block, a KxK 2D transform is performed on the surrounding samples of the block in a larger range than the 4x4 block, thereby obtaining the 2D transform result coefficient x i,j can be derived.
[0925] Next, T0, T1, T2, and T3 can be calculated according to the aforementioned [Formula 6], [Formula 7], [Formula 8], and [Formula 9], respectively. Each of T0, T1, T2, and T3 can be the sum of the absolute values of these coefficients.
[0926] Here, for each 4x4 block, we use the KxK 2D transform to obtain the 2D transform result coefficient x i,j Each of them can be calculated.
[0927] In Fig. 19, the entire rectangle including the small rectangles indicated by the thin solid lines may represent a block. The small rectangles indicated by the thin solid lines may represent one restored sample or the location of one restored sample.
[0928] In Fig. 19, the rectangle indicated by the bold solid line can represent the range to which the KxK 2D transformation is applied.
[0929] Here, the position corresponding to (i, j) = (0, 0) may mean the upper left position within the range where the KxK 2D transformation is applied.
[0930] After the calculation of the 2D transformation, a specific coefficient x i,j can be changed to 0, and 0 is a specific coefficient x i,j can be used for calculations instead.
[0931]
[0932] Figure 20 illustrates 4x4 2D transform 2x2 windows according to an example.
[0933] Figure 21 illustrates specific parts and sums of specific parts of a 4x4 2D transform according to an example.
[0934] Figure 22 shows the excluded region of a 4x4 2D transformation according to an example.
[0935] Figure 23 illustrates 8x8 2D transform 2x2 windows according to an example.
[0936] Figure 24 shows the sum of specific parts of specific parts of an 8x8 2D transform according to an example.
[0937] Figure 25 shows the excluded region of an 8x8 2D transformation according to an example.
[0938] In FIGS. 20 to 25, samples indicated with the same hatching may represent the ranges in which the operations described in the embodiments are performed.
[0939] For example, in FIGS. 20 to 25, the coefficients of the samples indicated by the same hatching can be used to derive the sum described in the embodiments. That is, the sum of the coefficients described in the embodiments; or the absolute values of the coefficients; can be calculated using the coefficients of the samples indicated by the same hatching in at least one of FIGS. 20 to 25.
[0940] In each drawing of the embodiments, hatching for a sample may indicate that samples indicated by the same hatching are used together in a particular operation and may mean that they are included in a particular group.
[0941] The areas marked in white in FIGS. 22 and 25 may represent coefficients / samples that are excluded from the calculations described in the examples among the total coefficients / samples.
[0942] In the process of performing a 2D transformation operation and comparing values to derive a sum of 2D transformation coefficients, at least one of the ranges in which the operation illustrated in FIGS. 20 to 25 is performed may be used.
[0943] In each of FIGS. 20 to 25, four regions are depicted. T0, T1, T2, and T3 can be derived for each of the four regions. T0, T1, T2, and T3 can be derived using the sum of the 2D transformation coefficients corresponding to each of the four regions.
[0944] The block classification index C can be classified into units of NxM blocks.
[0945] The block classification index C can be derived according to the order of the average of the sizes of surrounding samples having a range of KxL or LxK.
[0946] Here, the surrounding sample may be a restored sample or a differential sample of a differential block. Alternatively, the surrounding sample may be a sample of a specific block or a specific signal as described in the embodiments.
[0947] Each of N and M can be a positive integer. For example, each of N and M can be one of 2 and 4.
[0948] K can be any positive integer. For example, K can be one of 2, 3, 4, 5, 6, 7, and 8.
[0949] L can be any positive integer. For example, L can be one of 2, 3, 4, 5, 6, 7, and 8.
[0950]
[0951] Figure 26 shows a case where the range in which a sample is calculated according to an example has a square shape.
[0952] Figure 27 shows a case where the range in which samples are calculated according to an example has a rectangular shape.
[0953] For example, to derive a 2x2 block classification index, the averages S0, S1, S2 and S3 of samples in the KxL or LxK range of the 2x2 surrounding blocks can be computed using at least one of [Equation 10], [Equation 11], [Equation 12] and [Equation 13] below.
[0954] [Formula 10]
[0955]
[0956] [Formula 11]
[0957]
[0958] [Formula 12]
[0959]
[0960] [Formula 13]
[0961]
[0962] Here, α i , β i , γ i and δ i Each of them can be a weight that is multiplied to the sample.
[0963] As illustrated in FIG. 26, the range in which the samples are calculated may have a square shape. Alternatively, as illustrated in FIG. 27, the range in which the samples are calculated may have a rectangular shape.
[0964] In Figures 26 and 27, a square indicated by a thin solid line may represent a single restored sample or the location of a single restored sample. A square indicated by a thick solid line may represent a range over which an average is calculated.
[0965]
[0966] Figure 28 shows a case where the range in which a sample is calculated according to an example has a square shape.
[0967] Figure 29 shows a case where the range in which samples are calculated according to an example has a rectangular shape.
[0968] For example, to derive a 4x4 block classification index, the averages S0, S1, S2 and S3 of samples in the KxL or LxK range of the 4x4 surrounding blocks can be calculated using at least one of the aforementioned [Equation 10], [Equation 11], [Equation 12] and [Equation 13].
[0969] Here, α i , β i , γ i and δ i Each of them can be a weight that is multiplied to the sample.
[0970] As illustrated in FIG. 28, the range in which the samples are calculated may have a square shape. Alternatively, as illustrated in FIG. 29, the range in which the samples are calculated may have a rectangular shape.
[0971] In Figures 28 and 29, a square indicated by a thin solid line may represent a single restored sample or the location of a single restored sample. A square indicated by a thick solid line may represent a range over which an average is calculated.
[0972]
[0973] Calculation of block classification index
[0974] In the description below, P0, P1, P2 and P3 may represent T0, T1, T2 and T3, which are the sum or average of the transform coefficients after the KxK 2D transform is calculated; or S0, S1, S2 and S3, which are the averages of the surrounding pixels of the block.
[0975] The block classification index C of a block can be determined based on the transformation coefficients; and the pixel values of the surrounding pixels of the block.
[0976] The block classification index C of a block can be determined based on the statistical values of the pixel values of the surrounding pixels. Here, the statistical values can be one of the statistical values described in the embodiments, such as the sum of absolute values; the median value; and the mean value.
[0977] In the embodiments, the mean may mean a value determined by a calculation formula using multiple values described in the embodiments, such as the arithmetic mean; the geometric mean; the harmonic mean; and the median.
[0978] An NxM block classification index C can be derived by comparing the values of P0, P1, P2 and P3 with each other.
[0979] Each of a, b, c, and d can represent one of 0, 1, 2, and 3 according to [Formula 14] below. a can be the index of the largest value among P0, P1, P2, and P3. b can be the index of the second largest value among P0, P1, P2, and P3. c can be the index of the third largest value among P0, P1, P2, and P3. d can be the index of the smallest value among P0, P1, P2, and P3.
[0980] [Formula 14]
[0981] P a ≥ P b ≥ Pc ≥ P d
[0982] Given a, b, c, and d according to [Formula 14], the block classification index C can be derived by [Formula 15] below. Here, the block classification index C according to [Formula 15] can be one of the values greater than or equal to 0 and less than or equal to 23.
[0983] [Formula 15]
[0984] C = 6 * a + 2 * ((b > c) ? (b - 1) : b) + ((c > d) ? 1:0)
[0985] Additionally, the value of the block classification index C for the case where Equation 16 below holds may be additionally used. That is, when P0, P1, P2, and P3 are all equal, the block classification index C may have a specific value indicating that P0, P1, P2, and P3 are equal. Alternatively, the block classification index C of the specific value above may indicate that P0, P1, P2, and P3 are equal.
[0986] [Formula 16]
[0987] P a = P b = P c = P d
[0988] Alternatively, given a, b, c, and d according to [Formula 14], the block classification index C can be derived by [Formula 17] below. Here, the block classification index C according to [Formula 17] can be one of values greater than or equal to 0 and less than or equal to 11.
[0989] [Formula 17]
[0990] C = 3 * a + ((b > a) ? (b - 1) : b)
[0991] Additionally, the value of the block classification index C for the case where Equation 16 holds may be additionally used. That is, when P0, P1, P2, and P3 are all equal, the block classification index may have a specific value indicating that P0, P1, P2, and P3 are equal. Alternatively, the block classification index C of the above specific value may indicate that P0, P1, P2, and P3 are equal.
[0992] Alternatively, given a, b, c, and d according to [Formula 14], the block classification index C can be derived by [Formula 18] below. Here, the block classification index C according to [Formula 18] can be one of values greater than or equal to 0 and less than or equal to 3.
[0993] [Formula 18]
[0994] C = a
[0995] Additionally, to indicate the case where [Formula 16] holds, one more value of the block classification index C can be additionally used. That is, when P0, P1, P2, and P3 are all equal, the block classification index C can have a specific value indicating that P0, P1, P2, and P3 are equal. Alternatively, the block classification index C of the above specific value can indicate that P0, P1, P2, and P3 are equal.
[0996]
[0997] Figure 30 illustrates extracting samples in a vertical direction from surrounding samples according to an example.
[0998] Figure 31 illustrates extracting samples in a horizontal direction from surrounding samples according to an example.
[0999] Figure 32 illustrates extracting samples in the form of a 45 degree diagonal direction from surrounding samples according to an example.
[1000] Figure 33 illustrates extracting samples in the form of a -45 degree diagonal direction from surrounding samples according to an example.
[1001] The block classification index C can be classified into units of NxM blocks.
[1002] The block classification index C can be computed by the aligned KxL samples extracted at specific locations.
[1003] KxL samples can be partitioned into Kx1 vectors. By partitioning KxL samples, L Kx1 vectors can be determined.
[1004] The block classification index C can be determined based on L Kx1 vectors.
[1005] The block classification index C can be determined based on L values obtained by applying a specific operation to each of L Kx1 vectors.
[1006] The block classification index C can be determined based on a value obtained by applying a 1D transformation to each Kx1 vector among the L Kx1 vectors. L values can be obtained by applying a 1D transformation to each of the L Kx1 vectors, and the block classification index can be determined based on the L values.
[1007] Here, the samples may be restored samples; or differential samples of differential blocks; or, the samples may be samples of a specific block or a specific signal as described in the embodiments.
[1008] Here, each of N and M can be a positive integer. For example, each of N and M can be one of 2 and 4.
[1009] Here, K can be any positive integer. For example, K can be one of 2, 4, 8, and 10.
[1010] Here, L can be a positive integer. For example, L can be either 2 or 4.
[1011] In embodiments, the transform may include one or more of the transforms described in the embodiments, such as the Discrete Cosine Transform (DCT), the Discrete Sine Transform (DST), and the Hadamard Transform, or may be or use other transforms.
[1012] As illustrated in Figure 30, for 2x2 block classification, samples can be extracted in a vertical direction from surrounding samples.
[1013] As illustrated in Figure 31, for 2x2 block classification, samples can be extracted in a horizontal direction from surrounding samples.
[1014] As illustrated in Figure 32, for 2x2 block classification, samples can be extracted in the form of a 45 degree diagonal direction from surrounding samples.
[1015] As illustrated in Figure 33, for 2x2 block classification, samples can be extracted in the form of a -45 degree diagonal direction from surrounding samples.
[1016] In Figures 30 to 33, hatched areas may represent classified blocks. Rectangles marked with thin solid lines may represent a single restored sample or the location of a single restored sample. Rectangles marked with thick solid lines may represent the location of an extracted sample. Symbols within the thick solid lines may represent extracted samples.
[1017]
[1018] Figure 34 shows a first vertical sample extraction form in 2x2 block classification according to an example.
[1019] Figure 35 shows a second vertical sample extraction form in 2x2 block classification according to an example.
[1020] Figure 36 shows a third vertical sample extraction form in 2x2 block classification according to an example.
[1021] Figure 37 shows a fourth vertical sample extraction form in 2x2 block classification according to an example.
[1022] Figure 38 shows a fifth vertical sample extraction form in 2x2 block classification according to an example.
[1023] As illustrated in FIGS. 34 to 38, in 2x2 block classification, the vertical sample extraction form may be one of the first vertical sample extraction form, the second vertical sample extraction form, the third vertical sample extraction form, the fourth vertical sample extraction form, and the fifth vertical sample extraction form.
[1024] In Figs. 34 to 38, v l,i can represent a sample extracted by a vertical method.
[1025] A set of samples {v l,0 , v l,1 , ... , v l,k-1} is the lth sample vector V l It could be.
[1026] The lth vector V l In , at least n samples can be extracted, and a 1D transformation can be applied to a vector of extracted samples.
[1027] For example, at least n samples can be extracted from each sample vector of the sample vectors, and 1D transformations can be applied to each vector of the extracted samples. Here, as exemplified in FIGS. 34 to 38, samples corresponding to specific locations can be extracted when generating a vector.
[1028] Here, n can be a positive integer smaller than k.
[1029] For example, from the vector samples illustrated in FIG. 34, two vectors composed of 4x1 vector samples of [Equation 19] below can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1030] [Formula 19]
[1031] {v 0,3 , v 0,4 , v 0,5 , v 0,6}, {v 1,3 , v 1,4 , v 1,5 , v 1,6}
[1032] For example, even in the vector samples illustrated in FIGS. 35, 37, and 38, two vectors composed of 4x1 vector samples of [Equation 19] can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1033] For example, from the vector samples illustrated in FIG. 34, two vectors composed of 8x1 vector samples of [Formula 20] below can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1034] [Formula 20]
[1035] {v 0,0 , v 0,1 , v 0,2 , v 0,3 , v 0,4 , v 0,5 , v 0,6 , v 0,7}, {v 1,0 , v 1,1 , v 1,2 , v 1,3 , v 1,4 , v 1,5 , v 1,6 , v 1,7}
[1036] For example, even in the vector samples illustrated in FIGS. 35, 37, and 38, two vectors composed of 8x1 vector samples of [Formula 20] can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1037] For example, from the vector samples illustrated in FIG. 36, four vectors composed of 4x1 vector samples of [Equation 21] below can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1038] [Formula 21]
[1039] {v 0,3 , v 0,4 , v 0,5 , v 0,6}, {v 1,3 , v 1,4 , v 1,5 , v 1,6}, {v 2,3 , v 2,4 , v 2,5 , v 2,6}, {v 3,3 , v 3,4 , v 3,5 , v 3,6}
[1040] For example, from the vector samples illustrated in FIG. 36, four vectors composed of 8x1 vector samples of [Equation 22] below can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1041] [Formula 22]
[1042] {v 0,0 , v 0,1 , v 0,2 , v 0,3 , v 0,4 , v 0,5 , v 0,6 , v 0,7}, {v 1,0 , v 1,1 , v 1,2 , v 1,3 , v 1,4 , v 1,5 , v 1,6 , v1,7}, {v 2,0 , v 2,1 , v 2,2 , v 2,3 , v 2,4 , v 2,5 , v 2,6 , v 2,7}, {v 3,0 , v 3,1 , v 3,2 , v 3,3 , v 3,4 , v 3,5 , v 3,6 , v 3,7}
[1043]
[1044] Figure 39 shows a first horizontal sample extraction form in 2x2 block classification according to an example.
[1045] Figure 40 shows a second horizontal sample extraction form in 2x2 block classification according to an example.
[1046] Figure 41 shows a third horizontal sample extraction form in 2x2 block classification according to an example.
[1047] Figure 42 shows a fourth horizontal sample extraction form in 2x2 block classification according to an example.
[1048] Figure 43 shows a fifth horizontal sample extraction form in 2x2 block classification according to an example.
[1049] As illustrated in FIGS. 39 to 43, in 2x2 block classification, the horizontal sample extraction form may be one of the first horizontal sample extraction form, the second horizontal sample extraction form, the third horizontal sample extraction form, the fourth horizontal sample extraction form, and the fifth horizontal sample extraction form.
[1050] In Figs. 34 to 38, h l,i can represent a sample extracted by a horizontal method.
[1051] A set of samples {h l,0 , h l,1, ... , h l,k-1} is the lth sample vector H l It could be.
[1052] The lth vector H l In , at least n samples can be extracted, and a 1D transformation can be applied to a vector of extracted samples.
[1053] For example, at least n samples can be extracted from each sample vector of the sample vectors, and 1D transformations can be applied to each vector of the extracted samples. Here, as exemplified in FIGS. 34 to 38, samples corresponding to specific locations can be extracted when generating a vector.
[1054] Here, n can be a positive integer smaller than k.
[1055] For example, from the vector samples illustrated in FIG. 39, two vectors composed of 4x1 vector samples of [Equation 23] below can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1056] [Formula 23]
[1057] {h 0,3 , h 0,4 , h 0,5 , h 0,6}, {h 1,3 , h 1,4 , h 1,5 , h 1,6}
[1058] For example, even in the vector samples illustrated in FIGS. 40, 42, and 43, two vectors composed of 4x1 vector samples of [Equation 23] can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1059] For example, from the vector samples illustrated in FIG. 40, two vectors composed of 8x1 vector samples of [Equation 24] below can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1060] [Formula 24]
[1061] {h 0,0 , h 0,1 , h 0,2 , h 0,3 , h 0,4 , h 0,5 , h 0,6 , h 0,7}, {h 1,0 , h 1,1 , h 1,2 , h 1,3 , h 1,4 , h 1,5 , h 1,6 , h 1,7}
[1062] For example, even in the vector samples illustrated in FIGS. 40, 42, and 43, two vectors composed of 8x1 vector samples of [Equation 24] can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1063] For example, from the vector samples illustrated in FIG. 41, four vectors composed of 4x1 vector samples of [Equation 25] below can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1064] [Formula 25]
[1065] {h 0,3 , h 0,4 , h 0,5 , h 0,6}, {h 1,3 , h 1,4 , h 1,5 , h 1,6}, {h 2,3 , h 2,4 , h 2,5 , h 2,6}, {h 3,3 , h 3,4 , h 3,5 , h 3,6}
[1066] For example, from the vector samples illustrated in FIG. 41, four vectors composed of 8x1 vector samples of [Equation 26] below can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1067] [Formula 26]
[1068] {h 0,0 , h 0,1 , h 0,2 , h 0,3 , h 0,4 , h 0,5 , h 0,6 , h 0,7}, {h 1,0 , h 1,1 , h 1,2 , h 1,3 , h 1,4 , h 1,5 , h 1,6 , h 1,7}, {h 2,0 , h 2,1 , h 2,2 , h 2,3 , h 2,4 , h 2,5 , h 2,6 , h 2,7}, {h 3,0 , h 3,1 , h 3,2 , h 3,3 , h 3,4 , h 3,5 , h 3,6 , h 3,7}
[1069]
[1070] Figure 44 shows the first -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[1071] Figure 45 shows the second -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[1072] Figure 46 shows the third -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[1073] Figure 47 shows the fourth -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[1074] Figure 48 shows the fifth -45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[1075] As illustrated in FIGS. 44 to 48, in 2x2 block classification, the -45 degree diagonal sample extraction form may be one of the first -45 degree diagonal sample extraction form, the second -45 degree diagonal sample extraction form, the third -45 degree diagonal sample extraction form, the fourth -45 degree diagonal sample extraction form, and the fifth -45 degree diagonal sample extraction form.
[1076] In Figs. 44 to 48, d l,i can represent a sample extracted by the -45 degree diagonal method.
[1077] A set of samples {d l,0 , d l,1 , ... , d l,k-1} is the lth sample vector D l It could be.
[1078] lth vector D l In , at least n samples can be extracted, and a 1D transformation can be applied to a vector of extracted samples.
[1079] For example, at least n samples can be extracted from each sample vector of the sample vectors, and 1D transformations can be applied to each vector of the extracted samples. Here, as exemplified in FIGS. 44 to 48, samples corresponding to specific locations can be extracted when generating a vector.
[1080] Here, n can be a positive integer smaller than k.
[1081] For example, from the vector samples illustrated in FIG. 44, two vectors composed of 4x1 vector samples of [Equation 27] below can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1082] [Formula 27]
[1083] {d 0,3 , d 0,4 , d 0,5 , d 0,6}, {d 1,3 , d 1,4 , d 1,5 , d 1,6}
[1084] For example, even in the vector samples illustrated in FIGS. 45, 47, and 48, two vectors composed of 4x1 vector samples of [Equation 27] can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1085] For example, from the vector samples illustrated in FIG. 44, two vectors composed of 8x1 vector samples of [Equation 28] below can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1086] [Formula 28]
[1087] {d 0,0 , d 0,1 , d 0,2 , d 0,3 , d 0,4 , d 0,5 , d 0,6 , d 0,7}, {d 1,0 , d 1,1 , d 1,2 , d 1,3 , d 1,4 , d 1,5 , d 1,6 , d 1,7}
[1088] For example, even in the vector samples illustrated in FIGS. 45, 47, and 48, two vectors composed of 8x1 vector samples of [Equation 28] can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1089] For example, from the vector samples illustrated in FIG. 46, four vectors composed of 4x1 vector samples of [Equation 29] below can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1090] [Formula 29]
[1091] {d 0,3 , d 0,4 , d 0,5 , d 0,6}, {d 1,3 , d 1,4 , d 1,5 , d 1,6}, {d 2,3 , d 2,4 , d 2,5 , d 2,6}, {d 3,3 , d 3,4 , d 3,5 , d 3,6}
[1092] For example, from the vector samples illustrated in FIG. 46, four vectors composed of 8x1 vector samples of [Formula 30] below can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1093] [Formula 30]
[1094] {d 0,0 , d 0,1 , d 0,2 , d 0,3 , d 0,4 , d 0,5 , d 0,6 , d 0,7}, {d 1,0 , d 1,1 , d 1,2 , d 1,3 , d 1,4 , d 1,5 , d 1,6 , d1,7}, {d 2,0 , d 2,1 , d 2,2 , d 2,3 , d 2,4 , d 2,5 , d 2,6 , d 2,7}, {d 3,0 , d 3,1 , d 3,2 , d 3,3 , d 3,4 , d 3,5 , d 3,6 , d 3,7}
[1095]
[1096] Figure 49 shows the first 45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[1097] Figure 50 shows the second 45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[1098] Figure 51 shows the third 45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[1099] Figure 52 shows the fourth 45-degree diagonal sample extraction form in 2x2 block classification according to an example.
[1100] Figure 53 shows the fifth 45 degree diagonal sample extraction form in 2x2 block classification according to an example.
[1101] As illustrated in FIGS. 49 to 53, in 2x2 block classification, the 45-degree diagonal sample extraction form may be one of the first 45-degree diagonal sample extraction form, the second 45-degree diagonal sample extraction form, the third 45-degree diagonal sample extraction form, the fourth 45-degree diagonal sample extraction form, and the fifth 45-degree diagonal sample extraction form.
[1102] In Figs. 49 to 53, e l,i can represent a sample extracted by a 45 degree diagonal method.
[1103] A set of samples {e l,0 , e l,1 , ... , e l,k-1} is the lth sample vector E l It could be.
[1104] The lth vector E l In , at least n samples can be extracted, and a 1D transformation can be applied to a vector of extracted samples.
[1105] For example, at least n samples can be extracted from each sample vector of the sample vectors, and 1D transformations can be applied to each vector of the extracted samples. Here, as illustrated in FIGS. 49 to 53, when generating a vector, samples corresponding to specific locations can be extracted.
[1106] Here, n can be a positive integer smaller than k.
[1107] For example, from the vector samples illustrated in FIG. 49, two vectors composed of 4x1 vector samples of [Equation 31] below can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1108] [Formula 31]
[1109] {e 0,3 , e 0,4 , e 0,5 , e 0,6}, {e 1,3 , e 1,4 , e 1,5 , e 1,6}
[1110] For example, even in the vector samples illustrated in FIGS. 50, 52, and 53, two vectors composed of 4x1 vector samples of [Equation 19] can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1111] For example, from the vector samples illustrated in FIG. 49, two vectors composed of 8x1 vector samples of [Equation 32] below can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1112] [Formula 32]
[1113] {e 0,0 , e 0,1 , e 0,2 , e 0,3 , e 0,4 , e 0,5 , e 0,6 , e 0,7}, {e 1,0 , e 1,1 , e 1,2 , e 1,3 , e 1,4 , e 1,5 , e 1,6 , e 1,7}
[1114] For example, even in the vector samples illustrated in FIGS. 50, 52, and 53, two vectors composed of 8x1 vector samples of [Equation 32] can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1115] For example, from the vector samples illustrated in FIG. 51, four vectors composed of 4x1 vector samples of [Equation 33] below can be extracted, and a 4x1 1D transformation can be applied to each vector.
[1116] [Formula 33]
[1117] {e 0,3 , e 0,4 , e 0,5 , e 0,6}, {e 1,3 , e 1,4 , e 1,5 , e 1,6}, {e 2,3 , e 2,4 , e 2,5 , e 2,6}, {e 3,3 , e 3,4 , e 3,5 , e3,6}
[1118] For example, from the vector samples illustrated in FIG. 51, four vectors composed of 8x1 vector samples of [Equation 34] below can be extracted, and an 8x1 1D transformation can be applied to each vector.
[1119] [Formula 34]
[1120] {e 0,0 , e 0,1 , e 0,2 , e 0,3 , e 0,4 , e 0,5 , e 0,6 , e 0,7}, {e 1,0 , e 1,1 , e 1,2 , e 1,3 , e 1,4 , e 1,5 , e 1,6 , e 1,7}, {e 2,0 , e 2,1 , e 2,2 , e 2,3 , e 2,4 , e 2,5 , e 2,6 , e 2,7}, {e 3,0 , e 3,1 , e 3,2 , e 3,3 , e 3,4 , e 3,5 , e 3,6 , e 3,7}
[1121]
[1122] The horizontal vector samples, vertical vector samples, -45 degree diagonal vector samples and 45 degree diagonal vector samples extracted by the sample extraction method in the above-described embodiment can be represented as V, H, D and E, respectively.
[1123] For these vector samples, the 1D transformation kernel K of [Equation 35] below xBy applying the weighted sum using the kernel K as in [Equation 36] to [Equation 39] below, x Component values of I v,x , I h,x , I d,x and I e,x Each can be calculated.
[1124] [Formula 35]
[1125] K x = {k x,0 , k x,1 , ... , k x,n-1}
[1126] [Formula 36]
[1127]
[1128] [Formula 37]
[1129]
[1130] [Formula 38]
[1131]
[1132] [Formula 39]
[1133]
[1134] Here, the number of 1D transform kernels can be X. X can be a positive integer. For example, X can be 2, 3, or 4. x can be an index for the 1D transform kernels. x can be one of the values 0 to X-1.
[1135] Here, the number of sample vectors can be L. L can be a positive integer. For example, L can be 1, 2, 3, or 4. l can be an index for the sample vectors. l can be one of the values 0 to L-1.
[1136] Here, n can be the number of elements contained in the sample vector. n can be a positive integer. For example, n can be 2, 4, 6, or 8.
[1137] Kernel K x A block classification index C can be derived using the component values I of . For example, kernel K x The component value I can be used as the block classification index C.
[1138] For example, if the number of kernels x is 4, 16 component values can be derived as in [Formula 40] below.
[1139] [Formula 40]
[1140] I v,0 , I v,1 , I v,2 , I v,3 , I h,0 , I h,1 , I h,2 , I h,3 , I d,0 , I d,1 , I d,2 , I d,3 , I e,0 , I e,1 , I e,2 , I e,3
[1141] In one embodiment, the component values may be sorted in order of size. A block classification index C may be derived using the component value having the largest value among the component values. For example, the component value having the largest value among the component values may be used as the block classification index C.
[1142] In this case, the block classification index C can be one of the values 0 to 15.
[1143] Here, to indicate the case where all component values are the same, one more value of the block classification index C can be additionally used. That is, in the case where all I's are the same, the block classification index C can have a specific value indicating that all I's are the same. Alternatively, the block classification index C of the above specific value can indicate that all I's are the same.
[1144] In one embodiment, the component values may be sorted in order of the magnitude of their absolute values. A block classification index C may be derived using the component value with the largest absolute value among the component values. For example, the component value with the largest absolute value among the component values may be used as the block classification index C.
[1145] In addition to the absolute values, other statistical values described in the examples can be used for the derivation of the above sort and the above block classification index C.
[1146] In this case, the block classification index C can be one of the values 0 to 15.
[1147] Here, to indicate the case where all component values are the same, one more value of the block classification index C can be additionally used. That is, in the case where all I's are the same, the block classification index C can have a specific value indicating that all I's are the same. Alternatively, the block classification index C of the above specific value can indicate that all I's are the same.
[1148] In one embodiment, a block classification index C may be derived using a particular component value I having the largest absolute value among the component values and a sign of said particular component value I. The sign may be one of "+" and "-". For example, if a particular component value having the largest absolute value among the component values is identified, a value determined by a combination of the absolute value of said particular component value and the sign of said particular component value may be used as the block classification index C.
[1149] In this case, the block classification index C can be one of the values greater than or equal to 0 and less than or equal to 31.
[1150] Here, to indicate the case where all component values are the same, one more value of the block classification index C can be additionally used. That is, in the case where all I's are the same, the block classification index C can have a specific value indicating that all I's are the same. Alternatively, the block classification index C of the above specific value can indicate that all I's are the same.
[1151]
[1152] Filtering execution steps
[1153] Once the block classification index is determined, filtering can be performed on the restored image or restored block using a filter corresponding to the determined block classification index.
[1154] When such filtering is performed, one filter corresponding to the block classification index may be selected from among L available filters. The types of the L filters may be different from each other. In embodiments, the L filters may mean L types of filters.
[1155] For example, for each block classification unit, one of L filters can be selected, and filtering using the selected filter can be performed on the restored image for each restored sample.
[1156] For example, for each block classification unit, one filter among L filters can be selected, and filtering using the selected filter can be performed on the restored image block by block. Here, the block on which filtering is performed can be a block classification unit. That is, filter selection and filtering can be performed for each block classification unit.
[1157] For example, for each block classification unit, one of L filters may be selected, and filtering using the selected filter may be performed on a specific unit-by-unit basis in the restored image as described in the embodiments. For example, for each block classification unit, one of L filters may be selected, and filtering using the selected filter may be performed on a CU-by-CU basis in the restored image.
[1158] For example, for each block classification unit, U filters can be selected from L filters, and filtering using the selected filters can be performed on the restored image for each restored sample. U can be a positive integer.
[1159] For example, for each block classification unit, U filters can be selected from L filters, and filtering using the selected filters can be performed on the restored image block by block. U can be a positive integer. Here, the block on which filtering is performed can be a block classification unit. That is, filter selection and filtering can be performed for each block classification unit.
[1160] For example, for each block classification unit, U filters may be selected from L filters, and filtering using the selected filters may be performed on the restored image for each specific unit described in the embodiments. For example, for each block classification unit, U filters may be selected from L filters, and filtering using the selected filters may be performed on the restored image for each CU. U may be a positive integer.
[1161] The L filters described above can be called a filter set.
[1162] The L filters may be different from each other. The fact that the filters are different may mean that the filters are different from each other in at least one of a plurality of filter properties. In embodiments, the filter properties may include one or more of a filter coefficient, a number of filter taps (i.e., a filter length), a filter shape, and a filter type. In other words, two filters among the L filters may be different from each other in at least one of the plurality of filter properties. Alternatively, there may not be any filters among the L filters that have all of the plurality of filter properties the same.
[1163] At least one filter property among the plurality of filter properties of the L filters may be the same for at least one specific unit among the units described in the embodiments. For example, the specific units may include CU, PU, TU, CTU, CTB, slice, tile, image, and sequence, etc. The L filters may be the same for at least one of the plurality of filter properties applied to the specific unit.
[1164] The L filters may differ from each other for at least one specific unit among the units described in the embodiments. For example, the specific units may include CUs, PUs, TUs, CTUs, CTBs, slices, tiles, images, and sequences. The L filters may differ from each other in at least one of a plurality of filter properties applied to the specific unit. Alternatively, two filters among the L filters may differ from each other in at least one of a plurality of filter properties applied to the specific unit. Alternatively, there may not be filters among the L filters that have all the same filter properties applied to a specific unit.
[1165] Filtering may be performed using the same filter or different filters for at least one specific unit among the specific units described in the embodiments. For example, the specific units may include CUs, PUs, TUs, CTUs, CTBs, slices, tiles, images, and sequences.
[1166] Filtering may be performed based on information indicating whether filtering is performed on at least one specific unit among the specific units described in the embodiments. The information indicating whether filtering is performed on a specific unit may be information indicating whether filtering is to be performed on the specific unit. Information on whether filtering is to be performed on a specific unit may be signaled via a bitstream.
[1167]
[1168] Figure 54 shows a 9x9 diamond-shaped filter according to an example.
[1169] Figure 55 shows a 7x7 diamond-shaped filter according to an example.
[1170] Figure 56 shows a 5x5 diamond-shaped filter according to an example.
[1171] As a filter for filtering the embodiments, multiple filters having a rhombus (or diamond) filter shape and having different numbers of filter taps can be used.
[1172] The plurality of filters may be three filters. For example, the plurality of filters may be a diamond-shaped filter with a number of filter taps of 5x5, a diamond-shaped filter with a number of filter taps of 7x7, and a diamond-shaped filter with a number of filter taps of 9x9, as illustrated in FIGS. 54, 55, and 56.
[1173] Information about which filter among multiple filters is used can be signaled from the encoding device (110) to the decoding device (150).
[1174] Information about which filter among the plurality of diamond-shaped filters is used, such as the exceptions in FIGS. 54 to 56, can be signaled from the encoding device (110) to the decoding device (150).
[1175] Information about which filter is used may be a filter index or a coded filter index for a particular unit as described in the embodiments.
[1176] For this signaling, entropy encoding / decoding may be performed on the filter index. An encoded filter index may be generated by entropy encoding on the filter index. A filter index may be generated by entropy decoding on the encoded filter index.
[1177] For example, a specific unit may be one or more of a picture, a tile, a slice, and a sequence. That is, the filter index may be entropy encoded / decoded in the sequence parameter set, the picture parameter set, the slice header, and the slice data within the bitstream.
[1178] At least one of the plurality of rhombus-shaped filters can be used in filtering restored samples of at least one of the luma component and the chroma component.
[1179] For example, for a luma component restoration sample, at least one of the three diamond-shaped filters of FIGS. 54 to 56 can be used in filtering the restoration sample.
[1180] For example, for chroma component restoration samples, a 5x5 diamond-shaped filter may be used in restoration sample filtering. Alternatively, for chroma component restoration samples, filtering may be performed using a filter selected for the corresponding luma component.
[1181] In the filters illustrated in FIGS. 54 to 56, the numbers within each filter type may represent filter coefficient indexes.
[1182] The filters illustrated in FIGS. 54 to 56 may have filter coefficient indices that are symmetrical with respect to the center of the filter. That is, the filters illustrated in FIGS. 54 to 56 may be said to have a point symmetric filter form.
[1183] For a 9x9 rhombus-shaped filter, a total of 21 filter coefficients can be signaled / encoded / decoded. For a 7x7 rhombus-shaped filter, a total of 13 filter coefficients can be signaled / encoded / decoded. For a 3x3 rhombus-shaped filter, a total of 7 filter coefficients can be signaled / encoded / decoded. That is, a maximum of 21 filter coefficients may need to be signaled / encoded / decoded.
[1184] For a 9x9 rhombus filter, a total of 21 multiplications per sample may be required. For a 7x7 rhombus filter, a total of 13 multiplications per sample may be required. For a 5x5 rhombus filter, a total of 7 multiplications per sample may be required. In other words, filtering may need to be performed using a maximum of 21 multiplications per sample.
[1185] Since a 9x9 diamond-shaped filter has a size of 9x9, if a 9x9 diamond-shaped filter is implemented in hardware, a line buffer of 4 lines, which is half the filter's vertical length, may be required. That is, a line buffer of up to 4 lines may be required.
[1186]
[1187] Figure 57 shows an 11x11 cross-shaped filter according to an example.
[1188] Figure 58 shows a 9x9 cross-shaped filter according to an example.
[1189] Figure 59 shows a 7x7 cross-shaped filter according to an example.
[1190] As a filter for filtering the embodiments, multiple filters having a cross-shaped filter shape and having different numbers of filter taps can be used.
[1191] The plurality of filters may be three filters. For example, the plurality of filters may be an 11x11 cross-shaped filter, a 7x7 cross-shaped filter, and a 9x9 cross-shaped filter as illustrated in FIGS. 54, 55, and 56.
[1192] Here, various values can be used as the size of the filter shape. For example, at least one filter having filter taps of HxV, such as 11x9, 11x7, 11x5, 3x3, can be used.
[1193] Each of H and V can be a positive integer. H and V can be configured to have the same value. H and V can be configured to have different values.
[1194] At least one of H and V may be a pre-defined value in the encoding device (110) / decoding device (150), or at least one of H and V may be a value signaled from the encoding device (110) to the decoding device (150).
[1195] One of the values of H and V can be used to define the other value. Furthermore, the final values of H and / or V can be defined using the values of H and / or V.
[1196] In the filters illustrated in FIGS. 57 to 59, the numbers within each filter type may represent filter coefficient indices.
[1197] The filters illustrated in FIGS. 57 to 59 may have filter coefficient indices that are symmetrical with respect to the center of the filter. That is, the filters illustrated in FIGS. 57 to 59 may be said to have point-symmetric filter forms.
[1198] For an 11x11 cross-shaped filter, a total of 13 filter coefficients can be signaled / encoded / decoded. For a 9x9 cross-shaped filter, a total of 11 filter coefficients can be signaled / encoded / decoded. For a 7x7 diamond-shaped filter, a total of 9 filter coefficients can be signaled / encoded / decoded. That is, a maximum of 13 filter coefficients may need to be signaled / encoded / decoded.
[1199] For an 11x11 cross-shaped filter, a total of 13 multiplications per sample may be required. For a 9x9 diamond-shaped filter, a total of 9 multiplications per sample may be required. For a 7x7 diamond-shaped filter, a total of 9 multiplications per sample may be required. In other words, filtering may need to be performed using a maximum of 13 multiplications per sample.
[1200] Since an 11x11 cross-shaped filter has a size of 11x11, a line buffer of 5 lines, which is half the filter's vertical length, may be required when the 11x11 cross-shaped filter is implemented in hardware. That is, a line buffer of up to 5 lines may be required.
[1201]
[1202] Figure 60 shows an 11x11 diagonally symmetric cross-shaped filter according to an example.
[1203] Figure 61 shows a 9x9 diagonally symmetric cross-shaped filter according to an example.
[1204] Figure 62 shows a 7x7 diagonally symmetric cross-shaped filter according to an example.
[1205] Instead of the point-symmetric filter shape described in the examples, a filter shape in which the filter coefficients are symmetrical about the 45-degree diagonal, such as those illustrated in FIGS. 60 to 62, may also be used.
[1206] For example, an 11x11 cross-shaped filter with diagonally symmetric filter coefficients, such as that illustrated in FIG. 60, may be used. A 9x9 cross-shaped filter with diagonally symmetric filter coefficients, such as that illustrated in FIG. 61, may be used. A 7x7 cross-shaped filter with diagonally symmetric filter coefficients, such as that illustrated in FIG. 62, may be used.
[1207] In FIGS. 60 to 62, the numbers within each filter type may represent filter coefficient indices. The dotted lines may represent diagonal symmetry lines.
[1208] Here, various values can be used as the size of the filter shape. For example, at least one filter having filter taps of HxV, such as 11x9, 11x7, 11x5, 3x3, can be used.
[1209] Each of H and V can be a positive integer. H and V can be configured to have the same value. H and V can be configured to have different values.
[1210] At least one of H and V may be a pre-defined value in the encoding device (110) / decoding device (150), or at least one of H and V may be a value signaled from the encoding device (110) to the decoding device (150).
[1211] One of the values of H and V can be used to define the other value. Furthermore, the final values of H and / or V can be defined using the values of H and / or V.
[1212]
[1213] FIG. 63 may represent a 7x7 diagonally symmetric cross-shaped filter according to an example.
[1214] Fig. 64 may represent a 7x7 diagonally symmetric rhombus-shaped filter according to an example.
[1215] FIG. 65 may represent a 5x5 diagonally symmetric square shaped filter according to an example.
[1216] Figure 66 may represent a 7x7 diagonally symmetric plus shape filter according to an example.
[1217] Figure 67 may represent a 7x7 diagonally symmetric thin cross-shaped filter according to an example.
[1218] Figure 68 may represent a 7x7 diagonally symmetric octagonal filter according to an example.
[1219] The 45 degree diagonal symmetric filter may have a filter shape other than the point-symmetric cross shape, as illustrated in FIGS. 63 to 68.
[1220] For example, in FIGS. 63 to 68, rhombus, square and other filter shapes are illustrated.
[1221] In FIGS. 63 to 68, the filter shapes of the filters may be designed based on a filter length of 5 or 7. In addition to those illustrated in FIGS. 63 to 68, a filter having a filter shape designed based on a filter length of M may be used for filtering.
[1222] Here, M can be a positive integer.
[1223] In FIGS. 63 to 68, the dotted lines may represent diagonal lines of symmetry.
[1224]
[1225] Figure 69 shows a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[1226] Figure 70 illustrates a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[1227] Figure 71 shows a 5x5 diagonally symmetric rectangular shape filter according to an example.
[1228] Figure 72 shows a 7x7 diagonally symmetric thick rectangular shape filter according to an example.
[1229] A diagonally symmetric filter can have a diagonally symmetric filter form other than a point symmetric form.
[1230] For example, in FIGS. 69 through 72, cropped cross, right angle and other filter shapes are illustrated.
[1231] In FIGS. 69 to 72, the filter shapes of the filters may be designed based on a filter length of 5 or 7. In addition to those illustrated in FIGS. 69 to 72, a filter having a filter shape designed based on a filter length of M may be used for filtering.
[1232] Here, M can be a positive integer.
[1233] In Figures 69 to 72, the dotted lines may represent diagonal lines of symmetry.
[1234]
[1235] Figure 73 illustrates a case where no transformation is performed for a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[1236] Figure 74 illustrates a case where no transformation is performed for a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[1237] Figure 75 illustrates a case where no transformation is performed for a 5x5 diagonally symmetric rectangular shape filter according to an example.
[1238] Figure 76 illustrates a case where a 90 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[1239] Figure 77 illustrates a case where a 90 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[1240] Figure 78 shows a case where a 90 degree rotation transformation is performed on a 5x5 diagonally symmetric rectangular shape filter according to an example.
[1241] Figure 79 illustrates a case where a 180 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[1242] FIG. 80 illustrates a case where a 180-degree rotation transformation is performed on a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[1243] Figure 81 shows a case where a 180-degree rotation transformation is performed on a 5x5 diagonally symmetric rectangular shape filter according to an example.
[1244] Figure 82 illustrates a case where a 270 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated cross-shaped filter according to an example.
[1245] Figure 83 shows a case where a 270 degree rotation transformation is performed on a 7x7 diagonally symmetric truncated thin cross-shaped filter according to an example.
[1246] Figure 84 shows a case where a 270 degree rotation transformation is performed on a 5x5 diagonally symmetric rectangular shape filter according to an example.
[1247] Before filtering is performed on a block classification unit, the sum or average of the transformation coefficients calculated for the block classification unit, T0, T1, T2 and T3; or the average of the surrounding pixels, S0, S1, S2 and S 3; According to at least one of the following, a geometric transformation can be performed on the filter coefficients f(k, l).
[1248] A geometric transformation on the filter coefficients may mean performing at least one of a 90 degree rotation; a 180 degree rotation; a 270 degree rotation; a second diagonal flipping; a first diagonal flipping; a vertical flipping; a horizontal flipping; vertical and horizontal flipping; and a zoom in / out on the filter to produce a geometrically transformed filter.
[1249] After performing a geometric transformation on the filter coefficients, filtering the restored samples using the geometrically transformed filter coefficients; and after performing a geometric transformation on the restored samples to which the filter is applied, filtering the geometrically transformed restored samples using the filter coefficients can produce the same results.
[1250] As illustrated in FIGS. 73 to 84, filter coefficients generated by performing at least one of the geometric transformations described in the embodiments, such as 90-degree rotation, 180-degree rotation, and 270-degree rotation, on the filter coefficients of the three filter shapes can be used in filtering. In addition, other filter shapes and filter transformations can be used.
[1251]
[1252] Figure 85 shows a 1x1 filter shape and a 3x3 filter shape applied to prediction / residual samples of an adaptive in-loop filter according to an example.
[1253] In the filter of the embodiments, one square may represent a sample. The number within the square may represent a filter coefficient for the sample. The number "n" within the square may represent a filter coefficient "c n " can be considered.
[1254] The adaptive in-loop filter can use as input the reconstructed samples that are after the Sample Adaptive Offset (SAO) and before the DeBlocking Filter (DBF).
[1255] In one embodiment, for the luma component, two fixed 13x13 diamond filters can be applied to derive two intermediate samples. After this application, an adaptive filter can be signaled and applied to these two intermediate samples, as shown in Equation 41 below.
[1256] The two intermediate samples may be post-reconstruction samples of SAO; and pre-reconstruction samples of the deblocking filter.
[1257] [Formula 41]
[1258]
[1259] Here, R(x, y) can refer to the current sample. f i,j may refer to the clipped difference between the subsequent neighboring samples of SAO; and R(x, y);
[1260] g i can be the clipped difference between the intermediate sample and R(x, y).
[1261] h i,j is the previous neighboring sample of DBF; and may be the clipped difference between R(x, y).
[1262] Filter coefficients c i can be signaled. i can have values from 0 to 24.
[1263] may be a filtered sample.
[1264] To further improve the performance of the Adaptive Loop Filter (ALF), prediction samples and residual samples can be utilized as additional inputs to the ALF for the luma component.
[1265] In particular, under the assumption that prediction samples are derived, the filtered sample of the embodiment can be derived as in [Formula 42] below.
[1266] [Formula 42]
[1267]
[1268] Here, p i,j can be the clipped difference between the neighboring predicted sample; and the current sample R(x, y);
[1269] Alternatively, when residual samples are used, the filtered samples can be derived as in [Equation 43] below.
[1270] [Formula 43]
[1271]
[1272] Here, r i,j may be a clipped neighboring residual sample value.
[1273] In an embodiment, when either the prediction samples or the residual samples are applied, two filter shapes may be used. The two filter shapes may include a 1x1 diamond shape and a 3x3 diamond shape, as illustrated in FIG. 85.
[1274] In specific information described in embodiments, such as an adaptation parameter set, information indicating whether prediction samples / residual samples are used for ALF may be signaled. Here, the signaled information may be specific information described in embodiments, such as a flag. The signaling information may be set by the encoding device (110).
[1275]
[1276] Figure 86 shows filter forms of an online derived adaptive in-loop filter according to an example.
[1277] Figure 87 illustrates different filter forms of an online derived adaptive in-loop filter according to an example.
[1278] In the ALF process, the luma CTB can reference a filter set within an adaptation parameter set. A Laplacian classifier or a band-based classifier can be applied to each 2x2 reconstructed luma block within the CTB. For each filter set, information such as a flag can be signaled to indicate that either the Laplacian-based classifier or the band-based classifier is applied to the 2x2 reconstructed luma block. For each class, a specific filter among different filters within the filter set can be applied.
[1279] The filter form of the online-derived ALF can include three types of taps. The three types can be classified as follows, as illustrated in Figure 86:
[1280] - Taps applied to neighboring restoration samples, such as taps #0 to #19.
[1281] - Tabs that apply to offline-filtered results, such as tabs #20 to #27.
[1282] - Taps applied to previous samples of the deblocking filter, such as taps #28 to #30.
[1283] In embodiments, other classifiers may be applied to residual samples. In another aspect, the filter shape of the online-derived ALF may be modified.
[1284] To extend the two classifiers based on ALF, a third classifier applied to residual samples can be added to ALF. The usage of the classifier can be signaled for each luma filter set within the APS.
[1285] In this classifier, for each 2x2 block, the sum of the absolute values of the residual samples within neighboring windows can be computed. The class index can be derived as shown in [Equation 44] below:
[1286] [Formula 44]
[1287] classIdx = sum >> (sample bit depth - 2)
[1288] The value of classIdx can be in the range of 0 to 24.
[1289] In another aspect, the filter shape of the online-induced ALF can be modified as shown in Fig. 87.
[1290] The number of taps applied to neighboring restored samples (e.g., taps #0 to #9) can be reduced, and the number of taps applied to offline-filtered results (e.g., taps #10 to #29) can be increased.
[1291] When ALF is applied to a sample, two fixed filters with corresponding Laplacian-based classifiers can be first applied to the ALF input. Next, the outputs of the fixed filters can be used to derive the ALF output.
[1292] In embodiments, fixed filters can be extended to be applied to samples prior to deblocking filters. Additionally, the classifiers of the fixed filters can be extended using median values and differences between sample values. Furthermore, the outputs of the first fixed filter can be used as inputs to the second fixed filter.
[1293] In one embodiment, two Laplacian-based classifiers can be applied to a 2x2 block. For each classifier, activity and directionality values can be derived based on horizontal gradients, vertical gradients, and diagonal gradients. Next, a class index can be determined based on the activity and directionality values. Using the derived two class indices, two fixed filters can be selected from two filter sets. Next, the two fixed filters can be applied to ALF input samples. The two fixed filters can be f0 and f1. Next, the signaled filter can be applied to the ALF input samples; samples prior to the deblocking filter; outputs of the two fixed filters; and residual samples.
[1294] In one embodiment, the classifiers of fixed filters can be extended. First, for each 2x2 block, the median values of the surrounding window can be computed. Next, for each sample in this window, the difference between the sample value and the median value can be computed. A scaling factor can be determined based on the activity value derived from the Laplacian classifier. The square root of the sum of the squared differences can be further quantized into C' by the scaling factor. The value of C' can be an integer greater than or equal to 0 and less than or equal to 7.
[1295] When i is 0, C i can represent the classifier of the i-th fixed filter. At this time, the class index C i ' can be derived as shown in [Formula 45] below:
[1296] [Formula 45]
[1297] C i ' = C' * 896 + C i
[1298] In embodiments, the total number of fixed filters may not change.
[1299] As with signaled filters, fixed filters can be applied to samples prior to the DBF.
[1300] A fixed filter f1 can be applied to the output of f0; and to the samples before DBF.
[1301] Within the ALF, two fixed filters based on the Laplacian and distribution classifiers can be applied to the input of the DBF and the input of the ALF. Next, the outputs of the two fixed filters can be used to derive the ALF output.
[1302] In one embodiment, two classifiers can be applied to a 2x2 luma block based on Laplacian values and variance. For each classifier, a class index can be determined based on activity, directionality, and scaled variance. Here, the activity value can be a scaled value of the sum of the horizontal Laplacian value and the vertical Laplacian value. The derived class indices can be used to select two fixed filters from two filter sets. Next, both fixed filters can be applied to the DBF input samples and the ALF input samples. Finally, a signaled filter can be applied to the ALF input samples, the DBF input samples, the outputs of the two fixed filters, the output of the Gaussian filter, and the residual sample. When ALF is applied to the chroma samples, a 9x9 diamond-shaped filter can be applied to the ALF input samples.
[1303] In one embodiment, fixed filters based on Laplacian and variance classifiers can be applied to chroma components. The fixed filters can be extended to chroma components. Similar to classifiers for the luma component, classifiers based on Laplacian values and variance can be applied to chroma components. Compared to the luma classifier of the fixed filter, the chroma classifier can be applied at the sample level.
[1304] Additionally, when calculating the activity value, the sum of the chroma horizontal Laplacian value and the chroma vertical Laplacian value may be multiplied by 2 prior to scaling. Similarly, the chroma variance may be multiplied by 2 prior to scaling.
[1305] Next, the derived class indices can be used to select fixed filters from the set of chroma filters.
[1306] In an embodiment, a chroma-locked filter may be applied to ALF input samples in a 13x13 diamond shape and to DBF input samples in a 7x7 diamond shape.
[1307] Finally, in the signaled chroma filter, one additional tap can be introduced. The additional tap can be applied to the fixed filter output of the current sample.
[1308]
[1309] Figure 88 illustrates filter shapes within an adaptive loop filter according to an example.
[1310] Figure 89 shows an added filter form for an adaptive loop filter according to an example.
[1311] In the ALF design, online-trained filters can include three types of taps. The three types can include spatial taps, reconstruction-before-DBF taps, and fixed-filter-output based taps, as illustrated in FIG. 88.
[1312] Following the spatial taps positioned in a cross shape (e.g., taps #0 to #19), there may be three DBF-pre-restored based taps (e.g., taps #26, #27, #30) and offline-filtered taps (e.g., taps #20 to #25, #28, #20), for a total of 25 such taps.
[1313] In one embodiment, for ALF, additional fixed-filter-output based taps may be introduced to provide additional texture information.
[1314] Fig. 89 illustrates a method according to one embodiment. In Fig. 89, spatial taps (e.g., taps #0 to #19), DBF-before-reconstruction based taps (e.g., taps #26, #27, #34) and fixed-filter-output based taps (e.g., taps #20 to #25, #32, #33) may remain the same as in the embodiment described above with reference to Fig. 88, and a plurality of extended taps (e.g., #28 to #31, #35) may be introduced to the luma online-trained filters. The extended taps may take as input sources the outputs generated by DBF-before-reconstruction and additional fixed filters.
[1315] Here, the additional fixed filter may be a diamond 7x7 filter, which may have fixed coefficients in the encoding device (110) and the decoding device (150).
[1316] The method by which the coefficients are signaled may be the same as in other embodiments. The method of the embodiment may be always enabled without filter type switching.
[1317]
[1318] Figure 90 shows a diamond 9x9 filter shape of a chroma adaptive loop filter according to an example.
[1319] Figure 91 shows a cross 9x9 filter form of a component-by-component adaptive loop filter according to an example.
[1320] Figure 92 illustrates a diamond 9x9 chroma filter shape and luma residual-based taps introduced into a chroma adaptive loop filter according to an example.
[1321] Figure 93 shows a cross 9x9 component-wise adaptive loop filter shape and a luma residual-based tap introduced into the component-wise adaptive loop filter according to an example.
[1322] Luma residual values can be stored and used within a luma-ALF. In an embodiment, luma residual values can be introduced into a chroma-ALF and into a cross-component adaptive loop filter (CCALF).
[1323] The chroma-ALF filter can use a diamond 9x9 filter shape with a total of 20 taps, as illustrated in Fig. 90. All these taps can take spatial chroma restoration samples as input.
[1324] In CCALF, the online-trained CCALF filter can use a cross-shaped filter shape with a total of 25 taps, as illustrated in Figure 91. All these taps can take spatial luma reconstruction samples as input.
[1325] In one embodiment, all spatial-based taps of chroma-ALF and CCALF may remain unchanged, as illustrated in FIGS. 92 and 93.
[1326] For chroma-ALF, one luma residual tap can be added. The luma residual tap can take down-sampled luma residual samples as input.
[1327] For CCALF, five luma residual taps in a 3x3 cross shape can be added. The established taps can take co-located luma residual samples and neighboring luma residual samples as input.
[1328] The method by which the coefficients are signaled may be the same as in other embodiments. The method of the embodiment may be always enabled without filter type switching.
[1329] In CCALF, adaptive coefficient precision can be used instead of the fixed 7-bit precision.
[1330] In CCALF, the coefficient precision used in training and filtering can be a fixed value. For example, the precision value for the quanta Cb and Cr can be set to 7 bits.
[1331] During the training process, the coefficients of each filter can be computed by solving the corresponding Weiner-Hopf equation. These coefficients can be floating-point numbers during the calculation. Next, these coefficients can be mapped to {0, ±1, ±2, ±4, ±8, ±16, ±32, ±64} through a multiplication and rounding process with 128.
[1332] In the filtering process, since the coefficients were previously multiplied by 128, the filter outputs can be right-shifted by 7 bits to obtain the final correct results.
[1333] One embodiment can be applied to adaptive coefficient precision within CCALF.
[1334] During the training process, there can be four coefficient precision candidates. The candidate list can be set to {7, 8, 9, 10}.
[1335] First, each frame can traverse these precision candidates to obtain the corresponding filter coefficients.
[1336] Next, the corresponding distortion and bit cost for each precision can be calculated. The coefficient precision corresponding to the minimum RD cost can be selected as the optimal coefficient precision bestFilterCoeffPrec for the current frame. Integerized filter coefficients can be calculated using the selected bestFilterCoeffPrec. For the filtering process, the CTU filter index and the number of filters can be used, and thus, the CTU filter index and the number of filters can be included in the bitstream.
[1337] All filters for the same chroma component within a frame can share the same coefficient precision bestFilterCoeffPrec.
[1338] In the filtering process, instead of a fixed 7-bit precision value, the coefficient precision bestFilterCoeffPrec can be used. Position (x c , y c ), the output of the filter can be calculated as shown in [Equation 46] below.
[1339] [Formula 46]
[1340]
[1341] Here, bestFilterCoeffPrec may be the optimal coefficient precision selected in the training process.
[1342] can be integerized coefficients computed with bestFilterCoeffPrec.
[1343] may be the corresponding luma pixel.
[1344] Filtered pixels after CCALF can be calculated as shown in [Formula 47] below.
[1345] [Formula 47]
[1346]
[1347] Here, is the previous location of CCALF (x c , y c ) can be chroma pixel values.
[1348] There can be two bestFilterCoeffPrec syntax elements for each frame. One of the two bestFilterCoeffPrecs can be for Cr, and one can be for Cb. Furthermore, the two bestFilterCoeffPrecs can be encoded / decoded / signaled separately within the APS.
[1349] The luma coefficients of ALF can be represented by a fixed-point 8-bit decimal number. The fractional part can be represented by a 7-bit number; and 1 bit is used for the sign.
[1350] The real value of the coefficient can be in the range [-1, 1].
[1351] Similarly, the CCALF coefficient can be represented by a fixed-point 8-bit decimal number. The fractional part can be represented by a 7-bit number; and 1 bit for the sign.
[1352] CCALF coefficients can be values in the form of powers of 2.
[1353] In one embodiment, the number of bits used to represent the fractional part of a luma coefficient can be adaptively varied from 5 to 8. Each luma filter set can include up to 25 filters. For each luma filter set, a 2-bit syntax element can be signaled for each set of luma filters to indicate the number of bits for the coefficients within that set. The real value range of the coefficients can remain unchanged.
[1354] In one embodiment, the number of bits used to represent the fractional part of a CCALF can be adaptively varied from 7 to 10. Within an ALF Adaptation Parameter Set (APS), for each chroma component, a 2-bit syntax element can be encoded / decoded / signaled to indicate the number of bits used for the CCALF coefficients for that component. Additionally, the power-of-two constraint can be removed.
[1355]
[1356] Figure 94 shows the architecture of a low complexity operating point model according to an example.
[1357] Figure 95 illustrates processing of a backbone block in a low complexity operating point model according to an example.
[1358] Figure 96 shows parameters for a head block according to an example.
[1359] Figure 97 shows parameters for a backbone according to an example.
[1360] In Figure 94, the architecture of the low-complexity operation point (LOP) model is illustrated.
[1361] In the LOP model, inputs may include reconstructed samples (Rec), predicted samples (Pred), boundary strength (BS), base quantization parameter (QPbase), slice quantization parameter (QPslice), and block prediction information (IPB).
[1362] The chroma input can be upsampled by a factor of 2 to align it with the spatial dimension of the luma input.
[1363] There can be two branches within a network: one for luma and one for chroma.
[1364] On the output side, the luma branch may have a pixel shuffling process to revert the luma samples to their original spatial dimensions to compensate for downsampling that occurred within the head.
[1365]
[1366] Figure 98 shows the dimensions for a reconstruction input according to an example.
[1367] Figure 99 shows the alternative and input dimensions of DCT-II according to an example.
[1368] In an embodiment, a 2x2 DCT-II transform and reshaping can be applied to the input, and an inverse transform and inverse reshaping can be applied to the output.
[1369] Under complexity constraints, training a model to improve the transform coefficients (rather than directly peddling luma samples) may be beneficial for LOP performance.
[1370] In FIG. 98 and FIG. 99, examples of applying DCT transform to restored samples and reshaping are shown.
[1371] Figure 98 illustrates the current dimensions for the reconstruction input within the LOP.
[1372] Figure 99 illustrates the placement and input dimensions of DCT.
[1373] In Fig. 99, an example of restored samples having a size of 144x144 is shown.
[1374] An input luma with a size of 144x144x1 can be converted and reshaped to 72x72x4.
[1375] In Fig. 99, the luma restoration input can be transformed for each 2x2 sub-block. Each 2x2 sub-block can be reshaped into four channels corresponding to four DCT coefficients. A luma input of size 144x144x1 can be transformed and reshaped into 72x72x4. Here, "4" can represent four DCT frequency channels.
[1376] For chroma restoration samples, due to the upsampling by 2, each 2x2 sub-block can be constant and no transform needs to be applied. Therefore, the chroma restoration samples can be directly downsampled by 2 to have a size of 72x72x2.
[1377] Within the training data, chroma data is already upsampled and stored alongside luma, so the upsampling by a factor of 2 in chroma can be maintained for compatibility with the current training.
[1378] In more efficient implementations, both upsampling and downsampling can be omitted with identical results.
[1379]
[1380] Figure 100 illustrates an input-transformed model and dimensionality change through the above model according to an example.
[1381] Similarly, the same transformation and reshaping can follow for the inputs of "Pred" and "BS", resulting in a dimension of 72x72x6.
[1382] For "QPbase" and "QPslice", since "QPbase" and "QPslice" are constants, no transformation needs to be applied, and "QPbase" and "QPslice" can be directly reshaped into 72x72x1.
[1383] For "IPB", since only luma information exists, IPB can be converted and reshaped into 72x72x4.
[1384] Figure 100 illustrates the detailed dimensionality change through the model. The spatial dimension of the inputs can be 72x72. After downsampling by a factor of 2 within the network, the spatial dimension can become 36x36.
[1385] For the luma branch, the dimension before pixel shuffling can be 36x36x16. After pixel shuffling, the dimension can be 72x72x4.
[1386] Here, the four channels can represent four DCT frequency channels. Next, inverse transform and inverse reshaping can be applied to return to the spatial domain of dimension 144x144x1.
[1387] For each chroma branch, the previous dimension of the inverse DCT and inverse reshaping can be 36x36x(4x2), and the subsequent dimension can be 72x72x2, where 2 can represent the U and V channels.
[1388] The number of channels and backbone blocks for luma and chroma can be varied. For example, the following adjustments can be made:
[1389] - Regarding the head, d1 can be changed from 12 to 16. d2 can be 8. d3 can be 4. d4 can be 2. d5 can be 2. d6 can be changed from 24 to 48.
[1390] - About the backbone, N Y can be changed from 14 to 18. N UV can be changed from 4 to 7. C can be changed from 16 to 32. C 1Y can be changed from 64 to 96. C 1UV can be changed from 32 to 48. C 21 can be changed from 16 to 32.
[1391]
[1392] Figure 101 illustrates an input-transformed model according to another example and dimensionality change through the above model.
[1393] Figure 102 illustrates processing of a backbone block in an input-transformed model according to another example.
[1394] The description of the architecture of the model described above with reference to FIG. 100 can also be applied to the architecture of the model of FIG. 101. Duplicate descriptions are omitted.
[1395] Figure 101 illustrates the network structure of a loop filter based on LOP CNN.
[1396] Inputs to the loop filter may include reconstructed luma samples and reconstructed chroma samples (Rec), predicted luma samples and predicted chroma samples (Pred), boundary strength information for luma and chroma (BS), a base quantization parameter (QPbase), a slice quantization parameter (QPslice), and block prediction information (IPB).
[1397] Luma inputs can be transformed. For example, an input luma block with a size of WxH can be transformed by a 2x2 DCT-II for each 2x2 subblock, and reshaped into (W / 2)x(H / 2)x4. Here, "4" can represent four frequency channels.
[1398] Next, the transformed luma samples can be concatenated to the chroma U channel and chroma V channel to form a size of (W / 2)x(H / 2)x6. DCT transform and plane expansion can be applied to the luma pixels of the Rec, Pred, BS, and IPB inputs. These inputs are shown in bold in Fig. 101.
[1399] For "QPbase" and "QPslice", since "QPbase" and "QPslice" are constants, no transformation needs to be applied, and "QPbase" and "QPslice" can be directly reshaped as (W / 2)x(H / 2)x1.
[1400] For "IPB", since only luma information exists, IPB can be transformed and reshaped as (W / 2)x(H / 2)x4.
[1401] The transformed inputs can be processed by convolutional layers with kernel sizes of 3x3 or 1x1 and concatenated for fusion and transition.
[1402] In the fusion and transfer module, separable convolutions of 1x3 and 3x1 and downsampling with a factor of 2 can be used.
[1403] Next, the network can be split into two branches: one branch for luma and one branch for chroma.
[1404] Each branch may include sequential backbone blocks. Each backbone block may include a PReLU; a convolutional layer having a 1x1 kernel; separable convolutional layers having 1x3 and 3x1 kernels; and a convolutional layer having a 1x1 kernel.
[1405] On the output side, pixel shuffling can be applied to the luma output. Additionally, an inverse DCT can be applied to both luma and chroma. These applications can restore the outputs to the same spatial size as the inputs.
[1406] As illustrated in Figure 101, DCT can be applied to inputs, and inverse DCT can be applied prior to cropping the outputs.
[1407] In one embodiment, the variations of model parameters may be as follows:
[1408] - Regarding the head, d1 can be changed from 12 to 16. d2 can be 8. d3 can be 4. d4 can be 2. d5 can be 2. d6 can be changed from 24 to 64.
[1409] - About the backbone, N Y can be 14 days. N UV can be 4. C Y can be changed from 16 to 32. C UV can be changed from 16 to 32. C Y1 can be changed from 64 to 144. C UV1 can be changed from 32 to 128. C 21 can be changed from 16 to 32.
[1410]
[1411] Figure 103 shows a model in which fixed transformations are replaced with trainable components along with transformations, according to an example.
[1412] Figure 104 illustrates processing of a backbone block in a model with replacement applied according to an example.
[1413] Figure 105 shows parameters for the backbone in a model with replacement applied according to an example.
[1414] Figure 106 illustrates a trainable transformation for input adjustment according to an example.
[1415] As illustrated in Figure 103, the fixed transformation of the model of the embodiment can be replaced with trainable components with some modifications. The model of the embodiment can be identical to the models of other embodiments except for the following features.
[1416] - The DCT components with reshaping can be replaced by a trainable transformation. For example, a 1x1x4x4 convolution can be performed after a pixel unshuffle operation with a factor of 2. Here, the 1x1 convolution can perform an operation corresponding to a 2x2 spatial transformation on the original input pixel domain before the pixel arrangement with pixel unshuffle.
[1417] - The IDCT component with the reshape can be replaced with the inverse transform component, for example, by a 1x1xCxC convolution. Here, a pixel unshuffle operation with a factor of 2 can be performed after the 1x1xCxC convolution. Here, the 1x1 convolution can perform an operation corresponding to a 2x2 spatial inverse transform on the output features domain after the pixel arrangement with the pixel unshuffle. For the luma branch, C can be equal to 16. For the chroma branch, C can be equal to 8.
[1418] - Inverse trainable transformation within the luma branch can be performed prior to pixel shuffling in the LOP network.
[1419] - The entire filtering architecture can be implemented as a standalone model.
[1420] The parameters for the head block of Fig. 96 can also be applied to the model to which the replacement of this embodiment is applied.
[1421] In Fig. 106, an alternative transformation of the embodiment is illustrated.
[1422]
[1423] When a NN filter is applied to a reconstructed image, a scaling factor can be derived and signaled within the slice header for each color component. This derivation can be based on the least squares method. The difference between the input samples and the NN filtered samples (residuals) can be scaled by the scaling factors before being added to the input samples.
[1424]
[1425] Figure 107 shows a parallel synthesis of the output of a neural network loop filter and the output of a deblocking filter according to an example.
[1426] As illustrated in Figure 107, the previously restored samples of the deblocking filter can be fed to a low-complexity neural network loop filter (NNLF). Next, the final filtered samples can be generated by blending the results of the NNLF and the deblocking filter, as in [Equation 48] below.
[1427] [Formula 48]
[1428] R Refine = w × R NN + (1 - W) × R DB
[1429] R NN may be the result of NNLF. R DB can be the result of a deblocking filter. W can be a weight. R Refinecan represent the final filtered samples.
[1430] [1431...
Claims
1. A step of performing classification on blocks; and A step of performing filtering on the above block based on the above classification. A decryption method comprising:
2. In paragraph 1, A block classification index is assigned to the block by the above classification, A decryption method, wherein the filtering is performed using a filter corresponding to the above block classification index.
3. In paragraph 2, A decoding method, wherein the above block classification index is determined based on statistical values of coefficients derived by transformation for the above block.
4. In paragraph 3, A decryption method wherein the above statistical values are the sum of the values of the coefficients or the sum of the absolute values of the coefficients.
5. In paragraph 2, A decoding method, wherein the above block classification index is determined based on pixel values of surrounding pixels of the above block.
6. In paragraph 2, A decryption method, wherein a filter corresponding to the block classification index is selected from among a plurality of available filters.
7. In paragraph 2, A decoding method in which a geometric transformation is performed on the filter coefficients of the above filter.
8. Step of performing classification on blocks; and A step of performing filtering on the above block based on the above classification. An encoding method comprising:
9. In paragraph 8, A block classification index is assigned to the block by the above classification, An encoding method, wherein the filtering is performed using a filter corresponding to the block classification index.
10. In paragraph 9, An encoding method, wherein the above block classification index is determined based on statistical values of coefficients derived by transformation for the above block.
11. In paragraph 9, An encoding method wherein the above block classification index is determined based on pixel values of surrounding pixels of the above block.
12. In paragraph 9, An encoding method, wherein a filter corresponding to the block classification index is selected from among a plurality of available filters.
13. In paragraph 9, An encoding method in which a geometric transformation is performed on the filter coefficients of the above filter.
14. A computer-readable recording medium storing a bitstream generated by the encoding method of Article 8.
15. A computer-readable recording medium storing a bitstream for video decoding, wherein the bitstream comprises: Filter Information Including, Classification of blocks is performed based on the above filter information, A computer-readable recording medium in which filtering is performed on the block based on the above classification.
16. In paragraph 15, A block classification index is assigned to the block by the above classification, A computer-readable recording medium, wherein the filtering is performed using a filter corresponding to the block classification index.
17. In paragraph 16, A computer-readable recording medium, wherein the block classification index is determined based on statistical values of coefficients derived by transformation for the block.
18. In paragraph 16, A computer-readable recording medium, wherein the block classification index is determined based on pixel values of surrounding pixels of the block.
19. In Article 16, A computer-readable recording medium, wherein a filter corresponding to the block classification index is selected from among a plurality of available filters.
20. In paragraph 16, A computer-readable recording medium in which a geometric transformation is performed on filter coefficients of the above filter.
Citation Information
Patent Citations
Image processing device and method
KR101901087B1
Method and apparatus for reducing blocking artifact based on transformed coefficient correction
KR101910502B1
Moving image encoding device, moving image decoding device, moving image encoding method, moving image decoding method and storage medium
KR102024518B1
Confusion of multiple filters in adaptive loop filtering in video coding
KR102519780B1
KR20190063452A