Image encoding / decoding method and apparatus using intra-loop filtering
By combining in-loop filtering with sub-sampling block classification and multiple filter shapes, the problem of low data transmission and storage efficiency in high-resolution video is solved, reducing computational complexity and memory access bandwidth, and improving video encoding and decoding efficiency.
Patent Information
- Application Number
- CN202211663655.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-10-29
- Filing Date
- 2018-11-29
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2038-11-29
AI Technical Summary
Existing video coding technologies are inefficient in transmitting and storing high-resolution, high-quality video data, and existing filtering methods cannot effectively reduce the distortion between the original and reconstructed images.
By employing an in-loop filtering method, combining sub-sampling block classification and multiple filter shapes, filter information is determined through directionality and activity information, thereby reducing the computational complexity and memory access bandwidth of the video encoder/decoder.
It improves video encoding and decoding efficiency, reduces computational complexity and storage requirements, and lowers memory access bandwidth.
Smart Images

Figure CN115955562B_ABST
Abstract
Description
[0001] This application is a divisional application of application No. 201880086848.7 titled “Image encoding / decoding method and apparatus using in-loop filtering” filed with the China National Intellectual Property Office on November 29, 2018, which claims priority under Article 8 of the Patent Cooperation Treaty from application No. 10-2018-0154057 filed with the Korean Intellectual Property Office on November 29, 2018. TECHNICAL FIELD
[0002] The present application relates to a video encoding / decoding method, a video encoding / decoding apparatus, and a recording medium storing a bitstream. In particular, the present application relates to a video encoding / decoding method and apparatus using in-loop filtering. BACKGROUND
[0003] Currently, in various applications, there is an increasing demand for high-resolution, high-quality videos such as high-definition (HD) videos and ultra-high-definition (UHD) videos. As videos have higher resolutions and qualities, the amount of video data increases compared to existing video data. Therefore, when transmitting video data through a medium such as a wired / wireless broadband line or storing video data in an existing storage medium, the transmission or storage cost increases. To address this problem of high-resolution, high-quality video data, an efficient video encoding / decoding technique is required.
[0004] There are various video compression techniques such as an inter prediction technique for predicting pixel values in a current picture from pixel values within a previous picture or a subsequent picture, an intra prediction technique for predicting pixel values in a region of a current picture from another region of the current picture, a transform and quantization technique for compressing the energy of a residual signal, and an entropy encoding technique for assigning shorter codes to frequently occurring pixel values and longer codes to less frequently occurring pixel values. With these video compression techniques, video data can be efficiently compressed, transmitted, and stored.
[0005] Deblocking filtering is intended to reduce blocking artifacts around a block boundary by performing vertical filtering and horizontal filtering on the block boundary. However, the problem of deblocking filtering is that it cannot minimize the distortion between an original picture and a reconstructed picture when filtering is performed on the block boundary.
[0006] Sample adaptive offset (SAO) is a method for reducing ringing artifacts by adding an offset to a certain sample after comparing a pixel value of the sample with pixel values of neighboring samples on a sample-by-sample basis or adding an offset to samples whose pixel values are within a certain pixel value range. SAO has the effect of reducing the distortion between an original picture and a reconstructed picture to some extent by using rate-distortion optimization. However, there is a limit in minimizing the distortion when the difference between the original picture and the reconstructed picture is large. SUMMARY
[0007] TECHNICAL PROBLEM
[0008] An object of the present application is to provide a video encoding / decoding method and apparatus using in-loop filtering.
[0009] Another object of the present application is to provide a method and apparatus for in-loop filtering using sub-sampling based block classification to reduce computational complexity and memory access bandwidth of a video encoder / decoder.
[0010] Another object of the present application is to provide a method and apparatus for in-loop filtering using multiple filter shapes to reduce computational complexity, memory capacity requirement, and memory access bandwidth of a video encoder / decoder.
[0011] Another object of the present application is to provide a recording medium storing a bitstream generated by a video encoding / decoding method or apparatus.
[0012] Technical Solution
[0013] A video decoding method according to the present application can include decoding filter information on a coding unit, classifying samples in the coding unit into classes on a per block classification unit basis, and filtering the coding unit having the samples classified into the classes on a per block classification unit basis by using the filter information.
[0014] In the video decoding method according to the present application, the method can further include assigning a block classification index to the coding unit having the samples classified into a class on a per block classification unit basis, wherein the block classification index is determined according to directionality information and activity information.
[0015] In the video decoding method according to the present application, wherein at least one of the directionality information and the activity information is determined based on gradient values for at least one of a vertical direction, a horizontal direction, a first diagonal direction, and a second diagonal direction.
[0016] In the video decoding method according to the present application, wherein the gradient values are obtained using one-dimensional Laplacian operations for each of the block classification units.
[0017] In the video decoding method according to the present application, wherein the one-dimensional Laplacian operations are one-dimensional Laplacian operations operating on positions that are sub-sampled positions.
[0018] In the video decoding method according to the present application, wherein the gradient values are determined based on a temporal layer identifier.
[0019] In the video decoding method according to the present application, wherein the filter information includes at least one piece of information selected from information on whether to perform filtering, filter coefficient values, the number of filters, the number of filter taps (filter length), filter shape information, filter type information, information on whether to use a fixed filter for a block classification index, and filter symmetry type information.
[0020] In the video decoding method according to the present application, wherein the filter shape information includes at least one of a diamond shape, a rectangular shape, a square shape, a trapezoidal shape, a diagonal line shape, a snowflake shape, a numeral symbol shape, a four-leaf clover shape, a cross shape, a triangular shape, a pentagonal shape, a hexagonal shape, an octagonal shape, a decagonal shape, and a dodecagonal shape.
[0021] In the video decoding method according to the present application, wherein the filter coefficient values include filter coefficient values for a geometric transform of the coding unit having the samples classified into the class based on each block classification unit.
[0022] In the video decoding method according to the present application, wherein the filter symmetry type information includes at least one of point symmetry, horizontal symmetry, vertical symmetry, and diagonal symmetry.
[0023] Further, a video encoding method according to the present application can include classifying samples of a coding unit into a class based on each block classification unit, filtering the coding unit having the samples classified into the class based on each block classification unit by using filter information on the coding unit, and encoding the filter information.
[0024] In the video encoding method according to the present application, the method can further include assigning a block classification index to the coding unit having the samples classified into the class based on each block classification unit, wherein the block classification index is determined based on directionality information and activity information.
[0025] In the video encoding method according to the present application, wherein at least one of the directionality information and the activity information is determined based on gradient values for at least one of a vertical direction, a horizontal direction, a first diagonal direction, and a second diagonal direction.
[0026] In the video encoding method according to the present application, wherein the gradient values are obtained using one-dimensional Laplacian operations for each of the block classification units.
[0027] In the video encoding method according to the present application, wherein the one-dimensional Laplacian operations are one-dimensional Laplacian operations on sub-sampled positions.
[0028] In the video encoding method according to the present application, wherein the gradient value is determined based on a temporal layer identifier.
[0029] In the video encoding method according to the present application, wherein the filter information includes at least one piece of information selected from information on whether to perform filtering, filter coefficient values, the number of filters, the number of filter taps (filter length), filter shape information, filter type information, information on whether to use a fixed filter for a block classification index, and filter symmetry type information.
[0030] In the video encoding method according to the present application, wherein the filter shape information includes at least one of a diamond shape, a rectangle shape, a square shape, a trapezoid shape, a diagonal line shape, a snowflake shape, a numeral symbol shape, a shamrock shape, a cross shape, a triangle shape, a pentagon shape, a hexagon shape, an octagon shape, a decagon shape, and a dodecagon shape.
[0031] In the video encoding method according to the present application, wherein the filter coefficient values include filter coefficients of a geometric transform for each of the block classification units of the coding unit.
[0032] Further, a computer-readable recording medium according to the present application can store a bitstream generated by the video encoding method according to the present application.
[0033] Advantageous Effects
[0034] According to the present application, a video encoding / decoding method and apparatus using in-loop filtering can be provided.
[0035] In addition, according to the present application, a method and apparatus using in-loop filtering based on sub-sampled block classification to reduce the computational complexity and memory access bandwidth of a video encoder / decoder can be provided.
[0036] In addition, according to the present application, a method and apparatus using in-loop filtering using a plurality of filter shapes to reduce the computational complexity, memory capacity requirement, and memory access bandwidth of a video encoder / decoder can be provided.
[0037] In addition, according to the present application, a recording medium storing a bitstream generated by the video encoding / decoding method or apparatus can be provided.
[0038] In addition, according to the present application, video encoding and / or decoding efficiency can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 is a block diagram illustrating a configuration of an encoding apparatus to which an embodiment of the present application is applied;
[0040] Figure 2is a block diagram showing a configuration of a decoding apparatus to which an embodiment of the present application is applied;
[0041] Figure 3 is a schematic diagram showing a picture partition structure for encoding / decoding a video;
[0042] Figure 4 is a diagram showing one embodiment of intra prediction processing;
[0043] Figure 5 is a diagram showing one embodiment of inter prediction processing;
[0044] Figure 6 is a diagram for describing a transform and quantization process.
[0045] Figure 7 is a flowchart showing a video decoding method according to an embodiment of the present application;
[0046] Figure 8 is a flowchart showing a video encoding method according to an embodiment of the present application;
[0047] Figure 9 is a diagram showing an exemplary method of determining gradient values for horizontal, vertical, first diagonal, and second diagonal directions;
[0048] Figures 10 to 12 is a diagram showing a sub-sampling-based exemplary method of determining gradient values for horizontal, vertical, first diagonal, and second diagonal directions;
[0049] Figures 13 to 18 is a diagram showing a sub-sampling-based exemplary method of determining gradient values for horizontal, vertical, first diagonal, and second diagonal directions;
[0050] Figures 19 to 30 is a diagram showing an exemplary method of determining gradient values for horizontal, vertical, first diagonal, and second diagonal directions at a certain sample position according to an embodiment of the present application;
[0051] Figure 31 is a diagram showing an exemplary method of determining gradient values for horizontal, vertical, first diagonal, and second diagonal directions when a temporal layer identifier indicates a top layer;
[0052] Figure 32 is a diagram showing various computing techniques that can be used instead of one-dimensional Laplacian operation according to an embodiment of the present application;
[0053] Figure 33 is a diagram showing a diamond filter according to an embodiment of the present application;
[0054] Figure 34 is a diagram illustrating a 5x5 tap filter according to an embodiment of the present application;
[0055] Figure 35a and Figure 35b is a diagram illustrating various filter shapes according to an embodiment of the present application;
[0056] Figure 36 is a diagram illustrating horizontal and vertical symmetric filters according to an embodiment of the present application;
[0057] Figure 37 is a diagram illustrating filters generated by geometric transformation of a square filter, an octagonal filter, a snowflake filter, and a diamond filter according to an embodiment of the present application;
[0058] Figure 38 is a diagram illustrating a process of transforming a diamond filter including 9x9 coefficients into a square filter including 5x5 coefficients; and
[0059] Figures 39 to 55d is a diagram illustrating an exemplary method of determining gradient values with respect to horizontal, vertical, first diagonal, and second diagonal directions based on sub-sampling. DETAILED DESCRIPTION
[0060] Various modifications can be made to the present application, and there are various embodiments of the present application, and examples of the various embodiments will now be provided with reference to the accompanying drawings and will be described in detail. However, the present application is not limited thereto, although the exemplary embodiments can be interpreted as including all modifications, equivalents, or substitutions within the technical concept and technical scope of the present application. Like reference numerals refer to the same or similar functions throughout. In the drawings, the shapes and sizes of elements can be exaggerated for clarity. In the following detailed description of the present application, reference is made to the accompanying drawings that illustrate a specific embodiment of the present application by way of illustration. The embodiments are described in sufficient detail to enable those skilled in the art to implement the disclosure. It is understood that various embodiments of the present disclosure, although different, are not necessarily mutually exclusive. For example, specific features, structures, and characteristics described herein in connection with one embodiment can be implemented in other embodiments without departing from the spirit and scope of the present disclosure. Furthermore, it is understood that the position or arrangement of individual elements within each disclosed embodiment can be modified without departing from the spirit and scope of the present disclosure. Therefore, the following detailed description should not be interpreted in a limiting sense, and the scope of the present disclosure is limited only by the appended claims, along with the full range of equivalents that the claims are entitled to in the appropriate jurisdictions.
[0061] The terms "first", "second", etc. used in the specification can be used to describe various components, but the components are not construed as being limited by the terms. The terms are used only to distinguish one component from the other components. For example, a "first" component can be called a "second" component without departing from the scope of the present application, and a "second" component can also be similarly called a "first" component. The term "and / or" includes a combination of a plurality of items or any one of the plurality of items.
[0062] It will be understood that, in the specification, when an element is referred to as being "connected to" or "coupled to" another element, it can be "directly connected to" or "directly coupled to" the other element or connected to or coupled to the other element with other elements in between. On the contrary, it should be understood that when an element is referred to as being "directly coupled to" or "directly connected to" another element, there are no other elements interposed therebetween.
[0063] Further, constituent components shown in the embodiments of the present application are independently shown in order to present different characteristic functions from each other. Therefore, this does not mean that each of the constituent components is composed of a separate hardware or software constituent unit. In other words, for convenience, each of the constituent components includes each of the enumerated constituent components. Therefore, at least two of the constituent components in each of the constituent components can be combined to form one constituent component, or one constituent component can be divided into a plurality of constituent components for performing each function. Embodiments in which each of the constituent components is combined and embodiments in which one constituent component is divided are also included in the scope of the present application without departing from the essence of the present application.
[0064] The terms used in the specification only serve to describe specific embodiments and are not intended to limit the present application. Expressions used in the singular include the plural, unless they have a clearly different meaning in the context. In the specification, it will be understood that the terms such as "include", "has", "comprise", etc. are intended to indicate the existence of the features, numbers, steps, actions, elements, parts, or combinations thereof disclosed in the specification, and are not intended to exclude the possibility of the existence or addition of one or more other features, numbers, steps, actions, elements, parts, or combinations thereof. In other words, when a certain element is referred to as "included", no other element than the corresponding element is excluded, but additional elements can be included in the embodiments of the present application or the scope of the present application.
[0065] Furthermore, some of the constituent elements can not be essential elements to perform the functions essential to the present application, but can be optional elements to improve the performance thereof. The present application can be implemented by including only the essential constituent elements to implement the essence of the present application, excluding the constituent elements used when improving the performance. A structure including only the essential constituent elements, excluding the optional elements used when improving the performance, is also included in the scope of the present application.
[0066] Hereinafter, embodiments of the present application will be described in detail with reference to the accompanying drawings. In describing the exemplary embodiments of the present application, a well-known function or construction will not be described in detail in order not to unnecessarily obscure the understanding of the present application. The same or similar constituent elements are designated by the same reference numerals in the accompanying drawings, and repetitive description on the same elements will be omitted.
[0067] Hereinafter, an image can refer to a picture constituting a video, or can refer to a video itself. For example, "encoding or decoding or both encoding and decoding an image" can refer to "encoding or decoding or both encoding and decoding a moving picture", and can refer to "encoding or decoding or both encoding and decoding one of the pictures of a moving picture."
[0068] Hereinafter, the terms "moving picture" and "video" can be used as the same meaning and can be replaced with each other.
[0069] Hereinafter, a target image can be an encoding target image as an encoding target and / or a decoding target image as a decoding target. Furthermore, the target image can be an input image input to an encoding apparatus, and an input image input to a decoding apparatus. Here, the target image can have the same meaning as a current image.
[0070] Hereinafter, the terms "image", "picture", "frame", and "screen" can be used as the same meaning and can be replaced with each other.
[0071] Hereinafter, a target block can be an encoding target block as an encoding target and / or a decoding target block as a decoding target. Furthermore, the target block can be a current block as a target of current encoding and / or decoding. For example, the terms "target block" and "current block" can be used as the same meaning and can be replaced with each other.
[0072] Hereinafter, the terms "block" and "unit" can be used as the same meaning and can be replaced with each other. Or, "block" can mean a specific unit.
[0073] Hereinafter, the terms "region" and "segment" can be replaced with each other.
[0074] Hereinafter, a certain signal can be a signal representing a certain block. For example, an original signal can be a signal representing a target block. A prediction signal can be a signal representing a prediction block. A residual signal can be a signal representing a residual block.
[0075] In an embodiment, each of certain information, data, flag, index, element, and attribute, etc. can have a value. A value of information, data, flag, index, element, and attribute equal to "0" can represent a logical false or a first predefined value. In other words, the value "0", false, logical false, and the first predefined value can be replaced with each other. A value of information, data, flag, index, element, and attribute equal to "1" can represent a logical true or a second predefined value. In other words, the value "1", true, logical true, and the second predefined value can be replaced with each other.
[0076] When a variable i or j is used to represent a column, a row, or an index, a value of i can be an integer equal to or greater than 0, or an integer equal to or greater than 1. That is, a column, a row, an index, etc. can be counted from 0, or can be counted from 1.
[0077] Term Description
[0078] Encoder: denotes a device performing encoding. That is, denotes an encoding device.
[0079] Decoder: denotes a device performing decoding. That is, denotes a decoding device.
[0080] Block: is an array of samples of MxN. Here, M and N can denote a positive integer, and the block can denote an array of samples in a two-dimensional form. The block can refer to a unit. A current block can denote an encoding target block which becomes a target at the time of encoding, or a decoding target block which becomes a target at the time of decoding. Further, the current block can be at least one of an encoding block, a prediction block, a residual block, and a transform block.
[0081] Sample: is a basic unit constituting a block. According to a bit depth (Bd), a sample can be represented as a value from 0 to 2 Bd In the present invention, a sample can be used as a meaning of a pixel. That is, a sample, a pel, a pixel can have the same meaning as each other.
[0082] Unit: Can refer to a coding and decoding unit. When an image is coded and decoded, the unit can be a region generated by partitioning a single image. Also, when a single image is partitioned into sub-partitioned units during coding or decoding, the unit can mean a sub-partitioned unit. That is, an image can be partitioned into a plurality of units. When an image is coded and decoded, predetermined processing for each unit can be performed. A single unit can be partitioned into sub-units having sizes smaller than that of the unit. Depending on a function, the unit can mean a block, a macroblock, a coding tree unit, a coding tree block, a coding unit, a coding block, a prediction unit, a prediction block, a residual unit, a residual block, a transform unit, a transform block, etc. Also, to distinguish the unit from a block, the unit can include a luma component block, a chroma component block associated with the luma component block, and syntax elements of each color component block. The unit can have various sizes and shapes, and specifically, the shape of the unit can be a two-dimensional geometric figure such as a square, a rectangle, a trapezoid, a triangle, a pentagon, etc. Also, unit information can include at least one of a unit type indicating a coding unit, a prediction unit, a transform unit, etc., and a unit size, a unit depth, an order of coding and decoding of the unit, etc.
[0083] Coding tree unit: A single coding tree block configured with a luma component Y and two coding tree blocks related to chroma components Cb and Cr. Also, the coding tree unit can mean to include a block and syntax elements of each block. Each coding tree unit can be partitioned by using at least one of a quad-tree partitioning method, a binary-tree partitioning method, and a ternary-tree partitioning method to configure lower-level units such as coding units, prediction units, transform units, etc. The coding tree unit can be used as a term for designating a block of samples that becomes a processing unit when an image as an input image is coded / decoded. Here, the quad-tree can mean a quad-ary tree.
[0084] Coding tree block: Can be used as a term for designating any one of a Y coding tree block, a Cb coding tree block, and a Cr coding tree block.
[0085] Neighbor block: Can mean a block neighboring a current block. The block neighboring the current block can mean a block contacting a boundary of the current block, or a block located within a predetermined distance from the current block. The neighbor block can mean a block neighboring a vertex of the current block. Here, the block neighboring the vertex of the current block can mean a block vertically neighboring a neighbor block horizontally neighboring the current block, or a block horizontally neighboring a neighbor block vertically neighboring the current block.
[0086] Reconstructed neighboring block: can denote a neighboring block that is adjacent to the current block and has been spatially / temporally encoded or decoded. Here, the reconstructed neighboring block can denote a reconstructed neighboring unit. The reconstructed spatial neighboring block can be a block within the current picture and has been reconstructed by encoding or decoding or both. The reconstructed temporal neighboring block is a block or a neighboring block of the block at a position corresponding to the current block of the current picture within the reference picture.
[0087] Unit depth: can denote a degree of partitioning of a unit. In a tree structure, the highest node (root node) can correspond to a first unit that is not partitioned. Also, the highest node can have a minimum depth value. In this case, the depth of the highest node can be level 0. A node with a depth of level 1 can denote a unit generated by partitioning the first unit once. A node with a depth of level 2 can denote a unit generated by partitioning the first unit twice. A node with a depth of level n can denote a unit generated by partitioning the first unit n times. A leaf node can be a lowest node and is a node that cannot be further partitioned. The depth of the leaf node can be a maximum level. For example, a predefined value of the maximum level can be 3. The depth of the root node can be the lowest, and the depth of the leaf node can be the deepest. Also, when a unit is denoted as a tree structure, a level in which the unit exists can denote a unit depth.
[0088] Bitstream: can denote a bitstream including encoded image information.
[0089] Parameter set: corresponds to header information among configurations within a bitstream. At least one of a video parameter set, a sequence parameter set, a picture parameter set, and an adaptation parameter set can be included in the parameter set. Also, the parameter set can include slice header, parallel block group header, and parallel block header information. The term "parallel block group" denotes a group of parallel blocks and has the same meaning as a slice.
[0090] Parsing: can denote determining a value of a syntax element by performing entropy decoding, or can denote entropy decoding itself.
[0091] Symbol: can denote at least one of a syntax element, an encoding parameter, and a transform coefficient value of an encoding / decoding target unit. Also, the symbol can denote an entropy encoding target or an entropy decoding result.
[0092] Prediction mode: can be information indicating a mode encoded / decoded using intra prediction or a mode encoded / decoded using inter prediction.
[0093] Prediction unit: can mean a basic unit when performing prediction such as inter prediction, intra prediction, inter compensation, intra compensation, and motion compensation. A single prediction unit can be partitioned into multiple partitions having smaller sizes, or can be partitioned into multiple lower-level prediction units. The multiple partitions can be basic units when performing prediction or compensation. The partitions generated by partitioning the prediction unit can also be prediction units.
[0094] Prediction unit partition: can mean a shape obtained by partitioning a prediction unit.
[0095] Reference picture list: can mean a list including one or more reference pictures used for inter prediction or motion compensation. There are several types of available reference picture lists including LC (list combination), L0 (list 0), L1 (list 1), L2 (list 2), L3 (list 3).
[0096] Inter prediction indicator: can mean a direction of inter prediction (uni-prediction, bi-prediction, etc.) of a current block. Alternatively, the inter prediction indicator can mean a number of reference pictures used to generate a prediction block of the current block. Alternatively, the inter prediction indicator can mean a number of prediction blocks used when performing inter prediction or motion compensation on the current block.
[0097] Prediction list utilization flag: can mean whether at least one reference picture in a specific reference picture list is used to generate a prediction block. The prediction list utilization flag can be used to derive the inter prediction indicator, and vice versa. For example, when the prediction list utilization flag has a first value of zero (0), it means that no reference picture in the reference picture list is used to generate the prediction block. On the other hand, when the prediction list utilization flag has a second value of one (1), it means that the reference picture list is used to generate the prediction block.
[0098] Reference picture index: can mean an index indicating a specific reference picture in a reference picture list.
[0099] Reference picture: can mean a reference picture referred to by a specific block for the purpose of performing inter prediction or motion compensation on the specific block. Alternatively, the reference picture can be a picture including a reference block referred to by the current block for inter prediction or motion compensation. Hereinafter, the terms "reference picture" and "reference image" have the same meaning and can be replaced with each other.
[0100] Motion vector: can be a two-dimensional vector used for inter prediction or motion compensation. The motion vector can mean an offset between a coding / decoding target block and a reference block. For example, (mvX, mvY) can mean a motion vector. Here, mvX can mean a horizontal component, and mvY can mean a vertical component.
[0101] A search range can be a two-dimensional region searched during inter prediction for retrieving a motion vector. For example, a size of the search range can be MxN. Here, M and N are each an integer.
[0102] A motion vector candidate can refer to a prediction candidate block or a motion vector of the prediction candidate block when a motion vector is predicted. Also, the motion vector candidate can be included in a motion vector candidate list.
[0103] A motion vector candidate list can denote a list consisting of one or more motion vector candidates.
[0104] A motion vector candidate index can denote an indicator indicating a motion vector candidate in a motion vector candidate list. Alternatively, the motion vector candidate index can be an index of a motion vector predictor.
[0105] Motion information can denote information including at least one of a motion vector, a reference picture index, an inter prediction indicator, a prediction list utilization flag, reference picture list information, a reference picture, a motion vector candidate, a motion vector candidate index, a merge candidate, and a merge index.
[0106] A merge candidate list can denote a list consisting of one or more merge candidates.
[0107] A merge candidate can denote a spatial merge candidate, a temporal merge candidate, a combined merge candidate, a combined bi-predictive merge candidate, or a zero merge candidate. The merge candidate can include motion information such as an inter prediction indicator, a reference picture index for each list, a motion vector, a prediction list utilization flag, and an inter prediction indicator.
[0108] A merge index can denote an indicator indicating a merge candidate in a merge candidate list. Alternatively, the merge index can indicate a block in which a merge candidate has been derived among reconstructed blocks spatially / temporally neighboring a current block. Alternatively, the merge index can indicate at least one piece of motion information of the merge candidate.
[0109] A transform unit: can denote a basic unit when encoding / decoding a residual signal such as a transform, an inverse transform, a quantization, a dequantization, a transform coefficient encoding / decoding. A single transform unit can be partitioned into a plurality of lower-level transform units having a smaller size. Here, the transform / inverse transform can include at least one of a first transform / first inverse transform and a second transform / second inverse transform.
[0110] Scaling: can denote a process of multiplying a quantized level by a factor. A transform coefficient can be generated by scaling a quantized level. Scaling can also be referred to as dequantization.
[0111] Quantization parameter: can denote a value used when using a transform coefficient to generate a quantized level during quantization. The quantization parameter can also denote a value used when generating a transform coefficient by scaling a quantized level during inverse quantization. The quantization parameter can be a value mapped on a quantization step.
[0112] Delta quantization parameter: can denote a difference value between a predicted quantization parameter and a quantization parameter of a coding / decoding target unit.
[0113] Scan: can denote a method of ordering coefficients within a unit, a block, or a matrix. For example, changing a two-dimensional matrix of coefficients into a one-dimensional matrix can be referred to as a scan, and changing a one-dimensional matrix of coefficients into a two-dimensional matrix can be referred to as a scan or inverse scan.
[0114] Transform coefficient: can denote a coefficient value generated after performing a transform in an encoder. The transform coefficient can denote a coefficient value generated after performing at least one of entropy decoding and inverse quantization in a decoder. A quantized level obtained by quantizing a transform coefficient or a residual signal or a quantized transform coefficient level can also fall within the meaning of a transform coefficient.
[0115] Quantized level: can denote a value generated by quantizing a transform coefficient or a residual signal in an encoder. Alternatively, the quantized level can denote a value of an inverse quantization target that has undergone inverse quantization in a decoder. Similarly, a quantized transform coefficient level as a result of a transform and quantization can also fall within the meaning of a quantized level.
[0116] Non-zero transform coefficient: can denote a transform coefficient having a value other than zero, or a transform coefficient level or a quantized level having a value other than zero.
[0117] Quantization matrix: can denote a matrix used in a quantization process or an inverse quantization process performed to improve subjective image quality or objective image quality. The quantization matrix can also be referred to as a scaling list.
[0118] Quantization matrix coefficient: can denote each element within a quantization matrix. The quantization matrix coefficient can also be referred to as a matrix coefficient.
[0119] Default matrix: can denote a predetermined quantization matrix that is predefined in an encoder or a decoder.
[0120] Non-default matrix: can denote a quantization matrix that is not predefined in an encoder or a decoder but is signaled by a user.
[0121] Statistical value: a statistical value for at least one among a variable having a specific value that can be calculated, a coding parameter, a constant value, etc. can be one or more among an average value, a sum value, a weighted average value, a weighted sum value, a minimum value, a maximum value, a most frequently occurring value, a median value, an interpolated value.
[0122] Figure 1 FIG. 1 is a block diagram illustrating a configuration of an encoding apparatus according to an embodiment of the present application.
[0123] The encoding apparatus 100 can be an encoder, a video encoding apparatus, or an image encoding apparatus. The video can include at least one image. The encoding apparatus 100 can sequentially encode the at least one image.
[0124] Referring to Figure 1 The encoding apparatus 100 can include a motion prediction unit 111, a motion compensation unit 112, an intra prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy encoding unit 150, a dequantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference picture buffer 190.
[0125] The encoding apparatus 100 can perform encoding on an input image by using an intra mode or an inter mode or both the intra mode and the inter mode. Further, the encoding apparatus 100 can generate a bitstream including encoded information by encoding the input image, and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium, or can be streamed through a wired / wireless transmission medium. When the intra mode is used as a prediction mode, the switch 115 can be switched to the intra mode. Alternatively, when the inter mode is used as the prediction mode, the switch 115 can be switched to the inter mode. Here, the intra mode can denote an intra prediction mode, and the inter mode can denote an inter prediction mode. The encoding apparatus 100 can generate a prediction block for an input block of the input image. Further, the encoding apparatus 100 can encode a residual block using a residual of the input block and the prediction block after the prediction block is generated. The input image can be referred to as a current image as a current encoding target. The input block can be referred to as a current block as a current encoding target, or as an encoding target block.
[0126] When the prediction mode is the intra mode, the intra prediction unit 120 can use samples of a block that has been encoded / decoded and is adjacent to the current block as reference samples. The intra prediction unit 120 can perform spatial prediction on the current block by using the reference samples, or generate prediction samples of the input block by performing the spatial prediction. Here, the intra prediction can denote prediction within a frame.
[0127] When the prediction mode is the inter mode, the motion prediction unit 111 can retrieve a region most matching the input block from a reference picture when performing motion prediction, and derive a motion vector by using the retrieved region. In this case, the search region can be used as the region. The reference picture can be stored in the reference picture buffer 190. Here, when encoding / decoding of the reference picture is performed, the reference picture can be stored in the reference picture buffer 190.
[0128] The motion compensation unit 112 can generate a prediction block by performing motion compensation on the current block by using the motion vector. Here, the inter prediction can mean prediction between frames or motion compensation.
[0129] When a value of the motion vector is not an integer, the motion prediction unit 111 and the motion compensation unit 112 can generate a prediction block by applying an interpolation filter to a partial region of a reference picture. In order to perform inter-picture prediction or motion compensation on a coding unit, it can be determined which mode among a skip mode, a merge mode, an advanced motion vector prediction (AMVP) mode, and a current picture reference mode is used for motion prediction and motion compensation of a prediction unit included in the corresponding coding unit. Then, depending on the determined mode, inter-picture prediction or motion compensation can be differently performed.
[0130] The subtractor 125 can generate a residual block by using a residual of the input block and the prediction block. The residual block can be referred to as a residual signal. The residual signal can mean a difference between an original signal and a prediction signal. Also, the residual signal can be a signal generated by transforming or quantizing or transforming and quantizing a difference between the original signal and the prediction signal. The residual block can be a residual signal of a block unit.
[0131] The transform unit 130 can generate transform coefficients by performing a transform on the residual block, and output the generated transform coefficients. Here, the transform coefficients can be coefficient values generated by performing a transform on the residual block. When a transform skip mode is applied, the transform unit 130 can skip the transform on the residual block.
[0132] A quantized level can be generated by applying quantization to the transform coefficients or to the residual signal. Hereinafter, the quantized level can also be referred to as a transform coefficient in an embodiment.
[0133] The quantization unit 140 can generate a quantized level by quantizing the transform coefficients or the residual signal according to a parameter, and output the generated quantized level. Here, the quantization unit 140 can quantize the transform coefficients by using a quantization matrix.
[0134] The entropy encoding unit 150 can generate a bitstream by performing entropy encoding on the values calculated by the quantization unit 140 or on the encoding parameter values calculated when encoding is performed according to a probability distribution, and output the generated bitstream. The entropy encoding unit 150 can perform entropy encoding on the sample information of the image and information used to decode the image. For example, the information used to decode the image can include syntax elements.
[0135] When entropy encoding is applied, symbols are represented such that a smaller number of bits is allocated to symbols having a high generation probability, and a larger number of bits is allocated to symbols having a low generation probability, and thus, the size of a bitstream of the symbols to be encoded can be reduced. The entropy encoding unit 150 can use an encoding method for entropy encoding such as exponential Golomb, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), or the like. For example, the entropy encoding unit 150 can perform entropy encoding by using a variable length coding / code (VLC) table. Also, the entropy encoding unit 150 can derive a binarization method of a target symbol and a probability model of the target symbol / binary bit, and perform arithmetic encoding by using the derived binarization method and context model.
[0136] In order to encode the transform coefficient levels (quantized levels), the entropy encoding unit 150 can change the coefficients in a two-dimensional block form to a one-dimensional vector form by using a transform coefficient scanning method.
[0137] The coding parameters can include information such as syntax elements (flags, indices, etc.) that are coded in the encoder and signaled to the decoder, and information derived when performing encoding or decoding. The coding parameters can represent information needed when encoding or decoding an image. For example, at least one value or combination of the following can be included in the coding parameters: unit / block size, unit / block depth, unit / block partition information, unit / block shape, unit / block partition structure, whether or not partitioning in a quad-tree form is performed, whether or not partitioning in a binary tree form is performed, partitioning direction in a binary tree form (horizontal direction or vertical direction), partitioning form in a binary tree form (symmetric partitioning or asymmetric partitioning), whether or not the current coding unit is partitioned by triple tree partitioning, triple tree partitioning direction (horizontal direction or vertical direction), triple tree partitioning type (symmetric type or asymmetric type), whether or not the current coding unit is partitioned by multi-type tree partitioning, multi-type tree partitioning direction (horizontal direction or vertical direction), multi-type tree partitioning type (symmetric type or asymmetric type), and multi-type tree partitioning tree (binary tree or triple tree) structure, prediction mode (intra prediction or inter prediction), luma intra prediction mode / direction, chroma intra prediction mode / direction, intra partition information, inter partition information, coding block partition flag, prediction block partition flag, transform block partition flag, reference sample filtering method, reference sample filter tap, reference sample filter coefficient, prediction block filtering method, prediction block filter tap, prediction block filter coefficient, prediction block boundary filtering method, prediction block boundary filter tap, prediction block boundary filter coefficient, intra prediction mode, inter prediction mode, motion information, motion vector, motion vector difference, reference picture index, inter prediction angle, inter prediction indicator, prediction list utilization flag, reference picture list, reference picture, motion vector predictor index, motion vector predictor candidate, motion vector candidate list, whether or not to use merge mode, merge index, merge candidate, merge candidate list, whether or not to use skip mode, interpolation filter type, interpolation filter tap, interpolation filter coefficient, motion vector size, representation precision of motion vector, transform type, transform size, information whether or not first (primary) transform is used, information whether or not secondary transform is used, first transform index, secondary transform index, information whether or not residual signal exists, coding block pattern, coding block flag (CBF), quantization parameter, quantization parameter residual, quantization matrix, whether or not to apply intra loop filter, intra loop filter coefficient, intra loop filter tap, intra loop filter shape / form, whether or not to apply deblocking filter, deblocking filter coefficient, deblocking filter tap, deblocking filter strength, deblocking filter shape / form, whether or not to apply adaptive sample offset, adaptive sample offset value, adaptive sample offset class, adaptive sample offset type, whether or not to apply adaptive loop filter, adaptive loop filter coefficient, adaptive loop filter tap, adaptive loop filter shape / form,binarization / de-binarization method, context model determination method, context model update method, whether to perform normal mode, whether to perform bypass mode, context bin, bypass bin, significant coefficient flag, last significant coefficient flag, coding flag for a unit of a coefficient group, position of last significant coefficient, flag as to whether a value of a coefficient is greater than 1, flag as to whether a value of a coefficient is greater than 2, flag as to whether a value of a coefficient is greater than 3, information on a residual coefficient value, sign information, reconstructed luma sample, reconstructed chroma sample, residual luma sample, residual chroma sample, luma transform coefficient, chroma transform coefficient, quantized luma level, quantized chroma level, transform coefficient level scanning method, size of a motion vector search region at a decoder side, shape of a motion vector search region at a decoder side, number of times of motion vector search at a decoder side, information on a CTU size, information on a minimum block size, information on a maximum block size, information on a maximum block depth, information on a minimum block depth, image display / output order, slice identification information, slice type, slice partition information, parallel block identification information, parallel block type, parallel block partition information, parallel block group indication information, parallel block group type, parallel block group partition information, picture type, bit depth of an input sample, bit depth of a reconstructed sample, bit depth of a residual sample, bit depth of a transform coefficient, bit depth of a quantized level, and information on a luma signal or information on a chroma signal.
[0138] Here, signaling a flag or an index can mean that the corresponding flag or index is entropy-encoded by an encoder and included in a bitstream, and can mean that the corresponding flag or index is entropy-decoded by a decoder from the bitstream.
[0139] When the encoding apparatus 100 performs encoding through inter prediction, the encoded current picture can be used as a reference picture for another picture which is subsequently processed. Accordingly, the encoding apparatus 100 can reconstruct or decode the encoded current picture, or store the reconstructed or decoded picture in the reference picture buffer 190 as a reference picture.
[0140] The quantized level can be dequantized in the dequantization unit 160, or can be inverse-transformed in the inverse transform unit 170. The coefficient which is dequantized or inverse-transformed or both can be added to the prediction block by the adder 175. By adding the coefficient which is dequantized or inverse-transformed or both to the prediction block, a reconstructed block can be generated. Here, the coefficient which is dequantized or inverse-transformed or both can mean a coefficient for which at least one of dequantization and inverse transformation is performed, and can mean a reconstructed residual block.
[0141] The reconstructed block can pass through a filter unit 180. The filter unit 180 can apply at least one of a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the reconstructed sample, the reconstructed block, or the reconstructed picture. The filter unit 180 can be referred to as an in-loop filter.
[0142] The deblocking filter can remove blocking distortion generated in a boundary between blocks. In order to determine whether to apply the deblocking filter, it can be determined whether to apply the deblocking filter to the current block based on samples included in a number of rows or columns included in the block. When the deblocking filter is applied to the block, another filter can be applied according to a required deblocking filter strength.
[0143] In order to compensate for encoding errors, a suitable offset value can be added to a sample value by using a sample adaptive offset. The sample adaptive offset can correct an offset of a deblocked picture from an original picture in units of samples. A method of applying an offset considering edge information about each sample can be used, or a method of partitioning samples of a picture into a predetermined number of regions, determining a region to which an offset is applied, and applying the offset to the determined region can be used.
[0144] The adaptive loop filter can perform filtering based on a comparison result of a filtered reconstructed picture and an original picture. Samples included in a picture can be partitioned into a predetermined group, a filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information on whether to apply the ALF can be signaled through a coding unit (CU), and a form and coefficients of the ALF to be applied to each block can vary.
[0145] The reconstructed block or the reconstructed picture that has passed through the filter unit 180 can be stored in a reference picture buffer 190. The reconstructed block processed by the filter unit 180 can be a part of a reference picture. That is, the reference picture is a reconstructed picture composed of the reconstructed blocks processed by the filter unit 180. The stored reference picture can be used later in inter prediction or motion compensation.
[0146] Figure 2 FIG. 1 is a block diagram illustrating a configuration of a decoding apparatus according to an embodiment of the present application.
[0147] The decoding apparatus 200 can be a decoder, a video decoding apparatus, or a picture decoding apparatus.
[0148] Referring to Figure 2 , the decoding apparatus 200 can include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra prediction unit 240, a motion compensation unit 250, a summer 225, a filter unit 260, and a reference picture buffer 270.
[0149] The decoding apparatus 200 can receive a bitstream output from the encoding apparatus 100. The decoding apparatus 200 can receive a bitstream stored in a computer-readable recording medium, or can receive a bitstream being streamed through a wired / wireless transmission medium. The decoding apparatus 200 can decode the bitstream by using an intra mode or an inter mode. Furthermore, the decoding apparatus 200 can generate a reconstructed image or a decoded image produced by decoding, and output the reconstructed image or the decoded image.
[0150] When the prediction mode used at the time of decoding is the intra mode, the switch can be switched to the intra. Alternatively, when the prediction mode used at the time of decoding is the inter mode, the switch can be switched to the inter mode.
[0151] The decoding apparatus 200 can obtain a reconstructed residual block by decoding an input bitstream, and generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding apparatus 200 can generate a reconstructed block that is a decoding target by adding the reconstructed residual block to the prediction block. The decoding target block can be referred to as a current block.
[0152] The entropy decoding unit 210 can generate a symbol by entropy-decoding a bitstream according to a probability distribution. The generated symbol can include a quantized level form of a symbol. Here, the entropy-decoding method can be an inverse process of the above-described entropy-encoding method.
[0153] In order to decode a transform coefficient level (quantized level), the entropy decoding unit 210 can change a coefficient in a one-dimensional vector form to a two-dimensional block form by using a transform coefficient scanning method.
[0154] The quantized level can be inverse-quantized in the inverse quantization unit 220, or can be inverse-transformed in the inverse transform unit 230. The quantized level can be a result of inverse-quantization or inverse-transformation, or both inverse-quantization and inverse-transformation, and can be generated as a reconstructed residual block. Here, the inverse quantization unit 220 can apply a quantization matrix to the quantized level.
[0155] When the intra mode is used, the intra prediction unit 240 can generate a prediction block by performing spatial prediction on the current block, in which the spatial prediction uses sample values of blocks adjacent to the decoding target block and already decoded.
[0156] When the inter mode is used, the motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block, in which the motion compensation uses a motion vector and a reference image stored in the reference picture buffer 270.
[0157] The adder 225 can generate a reconstructed block by adding the reconstructed residual block to the prediction block. The filter unit 260 can apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the reconstructed block or the reconstructed picture. The filter unit 260 can output the reconstructed picture. The reconstructed block or the reconstructed picture can be stored in the reference picture buffer 270 and used when performing inter prediction. The reconstructed block processed by the filter unit 260 can be a part of a reference picture. That is, the reference picture is a reconstructed picture composed of the reconstructed blocks processed by the filter unit 260. The stored reference picture can be used later when performing inter prediction or motion compensation.
[0158] Figure 3 FIG. 1 is a diagram schematically illustrating a partition structure of a picture when encoding and decoding the picture. Figure 3 FIG. 2 schematically illustrates an example of partitioning a single unit into a plurality of lower-level units.
[0159] To effectively partition a picture, a coding unit (CU) can be used when encoding and decoding. The coding unit can be used as a basic unit when encoding / decoding a picture. Also, the coding unit can be used as a unit for distinguishing an intra prediction mode from an inter prediction mode when encoding / decoding a picture. The coding unit can be a basic unit for prediction, transform, quantization, inverse transform, dequantization, or encoding / decoding processing of transform coefficients.
[0160] Referring to Figure 3 , the picture 300 is sequentially partitioned in a maximum coding unit (LCU) and the LCU unit is determined as a partition structure. Here, the LCU can be used in the same meaning as a coding tree unit (CTU). The unit partitioning can denote partitioning of a block associated with the unit. In the block partitioning information, information of a unit depth can be included. The depth information can denote either or both of a number or degree of partitioning of a unit or a number and degree of partitioning of a unit. A single unit can be partitioned into a plurality of lower-level units hierarchically associated with the depth information based on a tree structure. In other words, the unit and the lower-level units generated by partitioning the unit can correspond to a node and child nodes of the node, respectively. Each of the partitioned lower-level units can have the depth information. The depth information can be information denoting a size of a CU and can be stored in each CU. The unit depth denotes a number and / or degree related to partitioning of a unit. Accordingly, the partitioning information of the lower-level units can include information on sizes of the lower-level units.
[0161] The partition structure can represent a distribution of coding units (CUs) within the LCU 310. The distribution can be determined according to whether a single CU is partitioned into a plurality of (a positive integer equal to or greater than 2, including 2, 4, 8, 16, etc.) CUs. The horizontal and vertical sizes of the CUs generated by the partitioning can be half of the horizontal and vertical sizes of the CU before the partitioning, respectively, or can have sizes smaller than the horizontal and vertical sizes before the partitioning according to the number of times of the partitioning, respectively. The CU can be recursively partitioned into a plurality of CUs. At least one of the height and width of the CU after the partitioning can be reduced compared to at least one of the height and width of the CU before the partitioning by the recursive partitioning. The partitioning of the CU can be performed recursively until a predefined depth or a predefined size. For example, the depth of the LCU can be 0, and the depth of a smallest coding unit (SCU) can be a predefined maximum depth. Here, as described above, the LCU can be a coding unit having a maximum coding unit size, and the SCU can be a coding unit having a minimum coding unit size. The partitioning starts from the LCU 310, and the depth of the CU increases by 1 when the horizontal size or the vertical size, or both, of the CU is reduced by the partitioning. For example, the size of the CU that is not partitioned can be 2Nx2N for each depth. Also, in the case of the CU that is partitioned, the CU having a size of 2Nx2N can be partitioned into four CUs having a size of NxN. As the depth increases by 1, the size of N can be halved.
[0162] Also, information on whether the CU is partitioned can be represented by using partitioning information of the CU. The partitioning information can be 1-bit information. All CUs except for the SCU can include the partitioning information. For example, when the value of the partitioning information is 1, the CU can not be partitioned, and when the value of the partitioning information is 2, the CU can be partitioned.
[0163] Referring to Figure 3 The LCU having a depth of 0 can be a 64x64 block. 0 can be a minimum depth. The SCU having a depth of 3 can be an 8x8 block. 3 can be a maximum depth. The CUs of the 32x32 block and the 16x16 block can be represented as depths 1 and 2, respectively.
[0164] For example, when a single coding unit is partitioned into four coding units, the horizontal and vertical sizes of the partitioned four coding units can be half the size of the horizontal and vertical sizes of the CU before being partitioned. In one embodiment, when a coding unit having a size of 32x32 is partitioned into four coding units, each of the partitioned four coding units can have a size of 16x16. When a single coding unit is partitioned into four coding units, the coding unit can be said to be partitioned in a quad-tree form.
[0165] For example, when one coding unit is partitioned into two sub-coding units, each of the two sub-coding units can have a horizontal size or a vertical size (width or height) that is half of a horizontal size or a vertical size of the original coding unit. For example, when a coding unit having a size of 32x32 is vertically partitioned into two sub-coding units, each of the two sub-coding units can have a size of 16x32. For example, when a coding unit having a size of 8x32 is horizontally partitioned into two sub-coding units, each of the two sub-coding units can have a size of 8x16. When one coding unit is partitioned into two sub-coding units, the coding unit can be referred to as being bipartitioned, or partitioned according to a binary tree partitioning structure.
[0166] For example, when one coding unit is partitioned into three sub-coding units, a horizontal size or a vertical size of the coding unit can be partitioned in a ratio of 1:2:1, thereby resulting in three sub-coding units having a ratio of 1:2:1 in horizontal size or vertical size. For example, when a coding unit having a size of 16x32 is horizontally partitioned into three sub-coding units, the three sub-coding units can have sizes of 16x8, 16x16, and 16x8, in order from a topmost sub-coding unit to a bottommost sub-coding unit. For example, when a coding unit having a size of 32x32 is vertically partitioned into three sub-coding units, the three sub-coding units can have sizes of 8x32, 16x32, and 8x32, in order from a leftmost sub-coding unit to a rightmost sub-coding unit. When one coding unit is partitioned into three sub-coding units, the coding unit can be referred to as being tripartitioned, or partitioned according to a ternary tree partitioning structure.
[0167] In Figure 3 In the above-described example, a coding tree unit (CTU) 320 is an example of a CTU to which all of the quad-tree partitioning structure, the binary tree partitioning structure, and the ternary tree partitioning structure are applied.
[0168] As described above, in order to partition a CTU, at least one of the quad-tree partitioning structure, the binary tree partitioning structure, and the ternary tree partitioning structure can be applied. The various tree partitioning structures can be applied to the CTU sequentially according to a predetermined priority order. For example, the quad-tree partitioning structure can be preferentially applied to the CTU. A coding unit that cannot be further partitioned using the quad-tree partitioning structure can correspond to a leaf node of a quad-tree. The coding unit corresponding to the leaf node of the quad-tree can be used as a root node of a binary tree and / or a ternary tree partitioning structure. That is, the coding unit corresponding to the leaf node of the quad-tree can be further partitioned according to the binary tree partitioning structure or the ternary tree partitioning structure, or can not be further partitioned. Accordingly, by preventing a coding block resulting from binary tree partitioning or ternary tree partitioning of a coding unit corresponding to a leaf node of a quad-tree from being further quad-tree partitioned, a block partitioning operation and / or an operation of signaling partitioning information can be efficiently performed.
[0169] The fact that a coding unit corresponding to a node of a quadtree is partitioned can be signaled using quad-partition information. The quad-partition information having a first value (e.g., "1") can indicate that the current coding unit is partitioned according to a quadtree partition structure. The quad-partition information having a second value (e.g., "0") can indicate that the current coding unit is not partitioned according to a quadtree partition structure. The quad-partition information can be a flag having a predetermined length (e.g., one bit).
[0170] There can be no priority between binary tree partitioning and ternary tree partitioning. That is, a coding unit corresponding to a leaf node of a quadtree can be further partitioned by any of binary tree partitioning and ternary tree partitioning. Further, a coding unit resulting from binary tree partitioning or ternary tree partitioning can be further partitioned by binary tree partitioning or further partitioned by ternary tree partitioning, or can not be further partitioned.
[0171] A tree structure in which there is no priority between binary tree partitioning and ternary tree partitioning is referred to as a multi-type tree structure. A coding unit corresponding to a leaf node of a quadtree can serve as a root node of a multi-type tree. Whether a coding unit corresponding to a node of a multi-type tree is partitioned can be signaled using at least one of multi-type tree partitioning indication information, partition direction information, and partition tree information. In order to partition a coding unit corresponding to a node of a multi-type tree, the multi-type tree partitioning indication information, the partition direction information, and the partition tree information can be sequentially signaled.
[0172] The multi-type tree partitioning indication information having a first value (e.g., "1") can indicate that the current coding unit is to be partitioned by a multi-type tree partition. The multi-type tree partitioning indication information having a second value (e.g., "0") can indicate that the current coding unit is not to be partitioned by a multi-type tree partition.
[0173] When a coding unit corresponding to a node of a multi-type tree is further partitioned according to a multi-type tree partition structure, the coding unit can include partition direction information. The partition direction information can indicate in which direction the current coding unit is to be partitioned according to the multi-type tree partition. The partition direction information having a first value (e.g., "1") can indicate that the current coding unit is to be vertically partitioned. The partition direction information having a second value (e.g., "0") can indicate that the current coding unit is to be horizontally partitioned.
[0174] When a coding unit corresponding to a node of a multi-type tree is further partitioned according to a multi-type tree partition structure, the current coding unit can include partition tree information. The partition tree information can indicate a tree partition structure to be used for partitioning a node of a multi-type tree. The partition tree information having a first value (e.g., "1") can indicate that the current coding unit is to be partitioned according to a binary tree partition structure. The partition tree information having a second value (e.g., "0") can indicate that the current coding unit is to be partitioned according to a ternary tree partition structure.
[0175] The partition indication information, the partition tree information, and the partition direction information can each be a flag having a predetermined length (e.g., one bit).
[0176] At least any one of the quad-tree partition indication information, the multi-type tree partition indication information, the partition direction information, and the partition tree information can be entropy coded / entropy decoded. In order to entropy code / entropy decode those types of information, information about neighboring coding units adjacent to the current coding unit can be used. For example, it is highly likely that the partition type (partitioned or not partitioned, partition tree, and / or partition direction) of a left neighboring coding unit and / or an above neighboring coding unit of the current coding unit is similar to the partition type of the current coding unit. Thus, context information for entropy coding / entropy decoding information about the current coding unit can be derived from information about the neighboring coding units. The information about the neighboring coding units can include at least any one of the quad-tree partition information, the multi-type tree partition indication information, the partition direction information, and the partition tree information.
[0177] As another example, in the binary tree partition and the ternary tree partition, the binary tree partition can be preferentially performed. That is, the current coding unit can first undergo the binary tree partition, and then coding units corresponding to leaf nodes of the binary tree can be set as root nodes for the ternary tree partition. In this case, for coding units corresponding to nodes of the ternary tree, neither the quad-tree partition nor the binary tree partition can be performed.
[0178] A coding unit that cannot be partitioned according to the quad-tree partition structure, the binary tree partition structure, and / or the ternary tree partition structure becomes a basic unit for encoding, prediction, and / or transformation. That is, the coding unit cannot be further partitioned for prediction and / or transformation. Thus, in a bitstream, there can be no partition structure information and partition information for partitioning a coding unit into prediction units and / or transformation units.
[0179] However, when the size of a coding unit (i.e., a basic unit for partitioning) is larger than the size of the maximum transform block, the coding unit can be recursively partitioned until the size of the coding unit is reduced to be equal to or smaller than the size of the maximum transform block. For example, when the size of the coding unit is 64x64 and when the size of the maximum transform block is 32x32, the coding unit can be partitioned into four 32x32 blocks for transform. For example, when the size of the coding unit is 32x64 and the size of the maximum transform block is 32x32, the coding unit can be partitioned into two 32x32 blocks for transform. In this case, the partitioning of the coding unit for transform is not signaled separately and can be determined by a comparison between the horizontal size or the vertical size of the coding unit and the horizontal size or the vertical size of the maximum transform block. For example, when the horizontal size (width) of the coding unit is larger than the horizontal size (width) of the maximum transform block, the coding unit can be vertically bisected. For example, when the vertical size (length) of the coding unit is larger than the vertical size (length) of the maximum transform block, the coding unit can be horizontally bisected.
[0180] Information of the maximum size and / or the minimum size of the coding unit and information of the maximum size and / or the minimum size of the transform block can be signaled or determined at a higher level of the coding unit. The higher level can be, for example, a sequence level, a picture level, a slice level, a tile group level, a tile level, etc. For example, the minimum size of the coding unit can be determined to be 4x4. For example, the maximum size of the transform block can be determined to be 64x64. For example, the minimum size of the transform block can be determined to be 4x4.
[0181] Information of the minimum size of the coding unit corresponding to a leaf node of a quad tree (quad tree minimum size) and / or information of the maximum depth of a multi-type tree from a root node to a leaf node (maximum tree depth of the multi-type tree) can be signaled or determined at a higher level of the coding unit. For example, the higher level can be a sequence level, a picture level, a slice level, a tile group level, a tile level, etc. The information of the minimum size of the quad tree and / or the information of the maximum depth of the multi-type tree can be signaled or determined for each of an intra slice and an inter slice.
[0182] The difference information between the size of the CTU and the maximum size of the transform block can be signaled or determined at a higher level of the coding unit. For example, the higher level can be a sequence level, a picture level, a slice level, a parallel block group level, a parallel block level, or the like. Information of the maximum size of the coding unit corresponding to each node of the binary tree (hereinafter, referred to as the maximum size of the binary tree) can be determined based on the size of the coding tree unit and the difference information. The maximum size of the coding unit corresponding to each node of the ternary tree (hereinafter, referred to as the maximum size of the ternary tree) can vary depending on the type of the slice. For example, for an intra slice, the maximum size of the ternary tree can be 32x32. For example, for an inter slice, the maximum size of the ternary tree can be 128x128. For example, the minimum size of the coding unit corresponding to each node of the binary tree (hereinafter, referred to as the minimum size of the binary tree) and / or the minimum size of the coding unit corresponding to each node of the ternary tree (hereinafter, referred to as the minimum size of the ternary tree) can be set to the minimum size of the coding block.
[0183] As another example, the maximum size of the binary tree and / or the maximum size of the ternary tree can be signaled or determined at the slice level. Alternatively, the minimum size of the binary tree and / or the minimum size of the ternary tree can be signaled or determined at the slice level.
[0184] Depending on the size information and the depth information of the various blocks described above, the quad partition information, the multi-type tree partitioning indication information, the partition tree information, and / or the partition direction information can or can not be included in the bitstream.
[0185] For example, when the size of the coding unit is not greater than the minimum size of the quad tree, the coding unit does not include the quad partition information. Thus, the quad partition information can be derived from the second value.
[0186] For example, when the size (horizontal size and vertical size) of the coding unit corresponding to the node of the multi-type tree is greater than the maximum size (horizontal size and vertical size) of the binary tree and / or the maximum size (horizontal size and vertical size) of the ternary tree, the coding unit can not be bi-partitioned or tri-partitioned. Thus, the multi-type tree partitioning indication information can not be signaled, but can be derived from the second value.
[0187] Optionally, the multi-type tree partitioning indication information can not be signaled, but can be derived from the second value, when the size (horizontal size and vertical size) of the coding unit corresponding to the node of the multi-type tree is the same as the maximum size (horizontal size and vertical size) of the binary tree and / or twice as large as the maximum size (horizontal size and vertical size) of the ternary tree. This is because, when partitioning the coding unit according to the binary tree partitioning structure and / or the ternary tree partitioning structure, a coding unit smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree is generated.
[0188] Optionally, the multi-type tree partitioning indication information can not be signaled, but can be derived from the second value, when the depth of the coding unit corresponding to the node of the multi-type tree is equal to the maximum depth of the multi-type tree. This is because, when partitioning the coding unit according to the binary tree partitioning structure and / or the ternary tree partitioning structure, a coding unit smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree is generated.
[0189] Optionally, the multi-type tree partitioning indication information can not be signaled, but can be derived from the second value, when the size (horizontal size and vertical size) of the coding unit corresponding to the node of the multi-type tree is the same as the maximum size (horizontal size and vertical size) of the binary tree and / or twice as large as the maximum size (horizontal size and vertical size) of the ternary tree. This is because, when partitioning the coding unit according to the binary tree partitioning structure and / or the ternary tree partitioning structure, a coding unit smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree is generated.
[0190] Optionally, the partitioning direction information can not be signaled, but can be derived from a value indicating the possible partitioning directions, when both the vertical binary tree partitioning and the horizontal binary tree partitioning or both the vertical ternary tree partitioning and the horizontal ternary tree partitioning are feasible for the coding tree corresponding to the node of the multi-type tree. This is because, when partitioning the coding tree according to the binary tree partitioning structure and / or the ternary tree partitioning structure, a coding tree smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree is generated.
[0191] Optionally, the partitioning tree information can not be signaled, but can be derived from a value indicating the possible partitioning tree structures, when both the vertical binary tree partitioning and the vertical ternary tree partitioning or both the horizontal binary tree partitioning and the horizontal ternary tree partitioning are feasible for the coding tree corresponding to the node of the multi-type tree. This is because, when partitioning the coding tree according to the binary tree partitioning structure and / or the ternary tree partitioning structure, a coding tree smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree is generated.
[0192] Figure 4 is a diagram illustrating an intra prediction process.
[0193] Figure 4 The arrows from the center to the outside in the diagram of
[0194] Intra coding and / or decoding can be performed by using reference samples of neighboring blocks of a current block. The neighboring blocks can be reconstructed neighboring blocks. For example, the intra coding and / or decoding can be performed by using coding parameters or values of the reference samples included in the reconstructed neighboring blocks.
[0195] A prediction block can represent a block generated by performing intra prediction. The prediction block can correspond to at least one of a CU, a PU, and a TU. A unit of the prediction block can have a size of one of the CU, the PU, and the TU. The prediction block can be a square block having a size of 2x2, 4x4, 16x16, 32x32, or 64x64, etc., or can be a rectangular block having a size of 2x8, 4x8, 2x16, 4x16, and 8x16, etc.
[0196] Intra prediction can be performed according to an intra prediction mode for a current block. The number of the intra prediction modes that the current block can have can be a fixed value, and can be a value determined differently according to properties of the prediction block. For example, the properties of the prediction block can include a size of the prediction block and a shape of the prediction block, etc.
[0197] Regardless of the block size, the number of the intra prediction modes can be fixed to N. Alternatively, the number of the intra prediction modes can be 3, 5, 9, 17, 34, 35, 36, 65, or 67, etc. Alternatively, the number of the intra prediction modes can vary according to the block size or the color component type or both the block size and the color component type. For example, the number of the intra prediction modes can vary according to whether the color component is a luma signal or a chroma signal. For example, as the block size becomes larger, the number of the intra prediction modes can increase. Alternatively, the number of the intra prediction modes for a luma component block can be greater than the number of the intra prediction modes for a chroma component block.
[0198] The intra prediction mode can be a non-angular mode or an angular mode. The non-angular mode can be a DC mode or a planar mode, and the angular mode can be a prediction mode having a specific direction or angle. The intra prediction mode can be represented by at least one of a mode number, a mode value, a mode number, a mode angle, and a mode direction. The number of the intra prediction modes can be M which is greater than or equal to 1, including the non-angular mode and the angular mode.
[0199] In order to perform intra prediction on a current block, a step of determining whether a sample included in a reconstructed neighboring block can be used as a reference sample of the current block can be performed. When there is a sample that cannot be used as the reference sample of the current block, a value obtained by copying or performing interpolation on at least one of the sample values of the samples included in the reconstructed neighboring block, or both copying and interpolation, can be used to replace the unavailable sample value of the sample, and thus the replaced sample value is used as the reference sample of the current block.
[0200] When intra prediction is performed, a filter can be applied to at least one of reference samples and prediction samples based on an intra prediction mode and a size of the current block.
[0201] In the case of the planar mode, when a prediction block of the current block is generated, depending on a position of a prediction target sample within the prediction block, a sample value of the prediction target sample can be generated by using a weighted sum of an upper reference sample and a left reference sample of the current sample and an upper right reference sample and a lower left reference sample of the current block. Also, in the case of the DC mode, when the prediction block of the current block is generated, an average of the upper reference sample and the left reference sample of the current block can be used. Also, in the case of the angular mode, the prediction block can be generated by using the upper reference sample, the left reference sample, the upper right reference sample, and / or the lower left reference sample of the current block. To generate the prediction sample value, interpolation can be performed on real number units.
[0202] An intra prediction mode of a current block can be entropy encoded / decoded by predicting an intra prediction mode of a block that exists adjacent to the current block. When the intra prediction modes of the current block and the adjacent block are the same, information that the intra prediction modes of the current block and the adjacent block are the same can be signaled by using predetermined flag information. Also, indicator information of an intra prediction mode that is the same as the intra prediction mode of the current block among intra prediction modes of a plurality of adjacent blocks can be signaled. When the intra prediction modes of the current block and the adjacent block are not the same, the intra prediction mode information of the current block can be entropy encoded / decoded by performing entropy encoding / decoding based on the intra prediction modes of the adjacent blocks.
[0203] Figure 5 is a diagram illustrating an embodiment of an inter prediction process.
[0204] In Figure 5 , a rectangle can represent a picture. In Figure 5 , an arrow indicates a prediction direction. Depending on an encoding type of a picture, the picture can be classified into an intra picture (I picture), a predicted picture (P picture), and a bi-predicted picture (B picture).
[0205] An I picture can be encoded by intra prediction without inter prediction. A P picture can be encoded by inter prediction by using a reference picture that exists in one direction (i.e., a forward direction or a backward direction) with respect to a current block. A B picture can be encoded by inter prediction by using reference pictures that exist in two directions (i.e., a forward direction and a backward direction) with respect to a current block. When inter prediction is used, an encoder can perform inter prediction or motion compensation, and a decoder can perform corresponding motion compensation.
[0206] Hereinafter, an embodiment of inter prediction will be described in detail.
[0207] Inter-prediction or motion compensation can be performed using the reference picture and the motion information.
[0208] The motion information of the current block can be derived during inter-prediction by each of the encoding apparatus 100 and the decoding apparatus 200. The motion information of the current block can be derived by using the motion information of the reconstructed neighboring block, the motion information of the collocated block (also referred to as a col block or a collocated block), and / or the motion information of a block neighboring the collocated block. The collocated block can denote a block spatially collocated with the current block within a previously reconstructed collocated picture (also referred to as a col picture or a collocated picture). The collocated picture can be one of the one or more reference pictures included in the reference picture list.
[0209] The method of deriving the motion information of the current block can vary according to a prediction mode of the current block. For example, as the prediction mode for inter-prediction, there can be an AMVP mode, a merge mode, a skip mode, a current picture reference mode, etc. The merge mode can be referred to as a motion merge mode.
[0210] For example, when the AMVP is used as the prediction mode, at least one of the motion vector of the reconstructed neighboring block, the motion vector of the collocated block, the motion vector of the block neighboring the collocated block, and a (0, 0) motion vector can be determined as a motion vector candidate for the current block, and a motion vector candidate list can be generated by using the motion vector candidate. The motion vector candidate of the current block can be derived by using the generated motion vector candidate list. The motion information of the current block can be determined based on the derived motion vector candidate. The motion vector of the collocated block or the motion vector of the block neighboring the collocated block can be referred to as a temporal motion vector candidate, and the motion vector of the reconstructed neighboring block can be referred to as a spatial motion vector candidate.
[0211] The encoding apparatus 100 can calculate a motion vector difference (MVD) between the motion vector of the current block and the motion vector candidate, and can perform entropy encoding on the motion vector difference (MVD). In addition, the encoding apparatus 100 can perform entropy encoding on a motion vector candidate index and generate a bitstream. The motion vector candidate index can indicate a best motion vector candidate among the motion vector candidates included in the motion vector candidate list. The decoding apparatus can perform entropy decoding on the motion vector candidate index included in the bitstream, and can select a motion vector candidate of a decoding target block from among the motion vector candidates included in the motion vector candidate list by using the entropy-decoded motion vector candidate index. In addition, the decoding apparatus 200 can add the entropy-decoded MVD to the motion vector candidate extracted through the entropy decoding, thereby deriving the motion vector of the decoding target block.
[0212] The bitstream can include a reference picture index indicating a reference picture. The reference picture index can be entropy encoded by the encoding apparatus 100 and then signaled to the decoding apparatus 200 as the bitstream. The decoding apparatus 200 can generate a prediction block of a decoded target block based on the derived motion vector and the reference picture index information.
[0213] Another example of a method of deriving motion information of a current block can be a merge mode. The merge mode can denote a method of merging motion of a plurality of blocks. The merge mode can denote a mode of deriving motion information of a current block from motion information of neighboring blocks. When the merge mode is applied, a merge candidate list can be generated using motion information of reconstructed neighboring blocks and / or motion information of collocated blocks. The motion information can include at least one of a motion vector, a reference picture index, and an inter prediction indicator. The prediction indicator can indicate a uni-prediction (L0 prediction or L1 prediction) or bi-prediction (L0 prediction and L1 prediction).
[0214] The merge candidate list can be a list of stored motion information. The motion information included in the merge candidate list can be at least one of a zero merge candidate and new motion information, wherein the new motion information is motion information of one neighboring block adjacent to the current block (spatial merge candidate), motion information of a collocated block of the current block included in a reference picture (temporal merge candidate), and a combination of motion information existing in the merge candidate list.
[0215] The encoding apparatus 100 can generate a bitstream by performing entropy encoding on at least one of a merge flag and a merge index, and can signal the bitstream to the decoding apparatus 200. The merge flag can be information indicating whether the merge mode is performed for each block, and the merge index can be information indicating which of the neighboring blocks of the current block is a merge target block. For example, the neighboring blocks of the current block can include a left neighboring block disposed to the left of the current block, an above neighboring block disposed above the current block, and a temporal neighboring block adjacent in time to the current block.
[0216] The skip mode can be a mode of applying motion information of a neighboring block as it is to a current block. When the skip mode is applied, the encoding apparatus 100 can perform entropy encoding on information of a fact that which block's motion information is to be used as motion information of the current block to generate a bitstream, and can signal the bitstream to the decoding apparatus 200. The encoding apparatus 100 can not signal syntax elements regarding at least any one of motion vector difference information, a coded block flag, and transform coefficient levels to the decoding apparatus 200.
[0217] The current picture reference mode can represent a prediction mode in which a previously reconstructed region within the current picture to which the current block belongs is used for prediction. Here, a vector can be used to specify the previously reconstructed region. Information indicating whether the current block is to be encoded in the current picture reference mode can be encoded by using a reference picture index of the current block. A flag or index indicating whether the current block is a block encoded in the current picture reference mode can be signaled, and the flag or index can be derived based on the reference picture index of the current block. In case that the current block is encoded in the current picture reference mode, the current picture can be added to a reference picture list for the current block so as to be located at a fixed position or an arbitrary position in the reference picture list. The fixed position can be, for example, a position indicated by a reference picture index 0, or a last position in the list. When the current picture is added to the reference picture list so as to be located at the arbitrary position, a reference picture index indicating the arbitrary position can be signaled.
[0218] Figure 6 is a diagram illustrating transform and quantization processes.
[0219] As Figure 6 illustrated in FIG. 1, transform and / or quantization processes are performed on a residual signal to generate a quantized level signal. The residual signal is a difference between an original block and a prediction block (i.e., an intra-predicted block or an inter-predicted block). The prediction block is a block generated by intra-prediction or inter-prediction. The transform can be a primary transform, a secondary transform, or both the primary transform and the secondary transform. The primary transform of the residual signal generates transform coefficients, and the secondary transform of the transform coefficients generates secondary transform coefficients.
[0220] At least one scheme selected from various pre-defined transform schemes is used to perform the primary transform. For example, examples of the pre-defined transform schemes include a discrete cosine transform (DCT), a discrete sine transform (DST), and a Karhunen-Loève transform (KLT). The transform coefficients generated by the primary transform can be subjected to the secondary transform. The transform scheme used for the primary transform and / or the secondary transform can be determined according to an encoding parameter of the current block and / or a neighboring block of the current block. Alternatively, the transform scheme can be determined by signaling of transform information.
[0221] Since the residual signal is quantized by the first transform and the second transform, a quantized level signal (quantization coefficient) is generated. Depending on the intra prediction mode or the block size / shape of the block, the quantized level signal can be scanned according to at least one of diagonal up-right scanning, vertical scanning, and horizontal scanning. For example, when the coefficients are scanned according to diagonal up-right scanning, the coefficients in the form of a block become in the form of a one-dimensional vector. In addition to diagonal up-right scanning, horizontal scanning or vertical scanning can be used to scan the coefficients in the form of a two-dimensional block horizontally or vertically, depending on the intra prediction mode and / or size of the transform block. The scanned quantized level coefficients can be entropy coded to be inserted into a bitstream.
[0222] The decoder entropy decodes the bitstream to obtain the quantized level coefficients. The quantized level coefficients can be arranged in the form of a two-dimensional block by inverse scanning. For the inverse scanning, at least one of diagonal up-right scanning, vertical scanning, and horizontal scanning can be used.
[0223] Then, the quantized level coefficients can be dequantized, and then inverse-transformed secondarily as necessary, and finally inverse-transformed first as necessary, to generate a reconstructed residual signal.
[0224] Hereinafter, a method of in-loop filtering using sub-sampling-based block classification according to an embodiment of the present application will be described with reference to Figures 7 to 5 5.
[0225] In the present application, the in-loop filtering method includes deblocking filtering, sample adaptive offset (SAO), bilateral filtering, and adaptive in-loop filtering, etc.
[0226] By applying at least one of deblocking filtering and SAO to a reconstructed picture (i.e., a video frame) generated by summing a reconstructed intra / inter prediction block and a reconstructed residual block, it is possible to effectively reduce blocking artifacts and ringing artifacts within the reconstructed picture. Deblocking filtering is intended to reduce blocking artifacts around a block boundary by performing vertical filtering and horizontal filtering on the block boundary. However, deblocking filtering has a problem in that it cannot minimize distortion between an original picture and a reconstructed picture when the block boundary is filtered. Sample adaptive offset (SAO) is a filtering technique to reduce ringing artifacts by adding an offset to a certain sample after comparing a pixel value of the sample with pixel values of neighboring samples on a sample-by-sample basis, or by adding an offset to samples whose pixel values are within a certain pixel value range. SAO has an effect of reducing distortion between an original picture and a reconstructed picture to some extent by using rate-distortion optimization. However, there is a limit in minimizing distortion when the difference between the original picture and the reconstructed picture is large.
[0227] Bi-directional filtering refers to a filtering technique that determines filter coefficients based on a distance from a center sample in a filtering target region to each of other samples in the filtering target region and based on a difference between a pixel value of the center sample and a pixel value of each of the other samples.
[0228] Adaptive in-loop filtering refers to a filtering technique that minimizes a distortion between an original picture and a reconstructed picture by using a filter that can minimize the distortion.
[0229] Unless specifically stated otherwise in the description of the present invention, in-loop filtering refers to adaptive in-loop filtering.
[0230] In the present invention, filtering refers to a process of applying a filter to at least one basic unit selected from a sample, a block, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree unit (CTU), a slice, a parallel block, a group of parallel blocks (a parallel block group), a picture, and a sequence. The filtering includes at least one of a block classification process, a filter execution process, and a filter information encoding / decoding process.
[0231] In the present invention, a coding unit (CU), a prediction unit (PU), a transform unit (TU), and a coding tree unit (CTU) respectively have the same meaning as a coding block (CB), a prediction block (PB), a transform block (TB), and a coding tree block (CTB).
[0232] In the present invention, a block refers to at least one of a CU, a PU, a TU, a CB, a PB, and a TB used as a basic unit at the time of encoding / decoding processing.
[0233] The in-loop filtering is performed so that bi-directional filtering, deblocking filtering, sample adaptive offset, and adaptive in-loop filtering are sequentially applied to a reconstructed picture to generate a decoded picture. However, the order in which filtering schemes to be classified as in-loop filtering are applied to the reconstructed picture is varied.
[0234] For example, the in-loop filtering can be performed so that deblocking filtering, sample adaptive offset, and adaptive in-loop filtering are sequentially applied to the reconstructed picture in this order.
[0235] Alternatively, the in-loop filtering can be performed so that bi-directional filtering, adaptive in-loop filtering, deblocking filtering, and sample adaptive offset are sequentially applied to the reconstructed picture in this order.
[0236] Further alternatively, the in-loop filtering can be performed so that adaptive in-loop filtering, deblocking filtering, and sample adaptive offset are sequentially applied to the reconstructed picture in this order.
[0237] Further alternatively, the in-loop filtering can be performed such that the adaptive in-loop filtering, the sample adaptive offset, and the deblocking filtering are sequentially applied to the reconstructed picture in this order.
[0238] In the present disclosure, a decoded picture refers to an output from performing in-loop filtering or post-processing filtering on a reconstructed picture composed of reconstructed blocks, wherein each reconstructed block is generated by summing a reconstructed residual block and a corresponding intra prediction block or summing a reconstructed block and a corresponding inter prediction block. In the present disclosure, the meaning of a decoded sample, a decoded block, a decoded CTU, or a decoded picture is the same as that of a reconstructed sample, a reconstructed block, a reconstructed CTU, or a reconstructed picture, respectively.
[0239] The adaptive in-loop filtering is performed on the reconstructed picture to generate a decoded picture. The adaptive in-loop filtering can be performed on the decoded picture that has already undergone at least one of the deblocking filtering, the sample adaptive offset, and the bi-directional filtering. In addition, the adaptive in-loop filtering can be performed on the reconstructed picture that has already undergone the adaptive in-loop filtering. In this case, the adaptive in-loop filtering can be repeatedly performed N times on the reconstructed picture or the decoded picture. In this case, N is a positive integer.
[0240] The in-loop filtering can be performed on the decoded picture that has already undergone at least one of the in-loop filtering methods. For example, when at least one of the in-loop filtering methods is performed on the decoded picture that has already undergone at least one of the other in-loop filtering methods, a parameter for a latter filtering method can be changed, and then a former filtering can be performed on the decoded picture using the changed parameter. In this case, the parameter includes an encoding parameter, a filter coefficient, a number of filter taps (filter length), a filter shape, a filter type, a number of filtering executions, a filter strength, a threshold, and / or a combination of these parameters.
[0241] The filter coefficient indicates a coefficient constituting a filter. Alternatively, the filter coefficient indicates a coefficient value corresponding to a specific mask position in a mask form, and a reconstructed sample is multiplied by the coefficient value.
[0242] The number of filter taps refers to a length of a filter. When a filter is symmetrical with respect to one specific direction, the filter coefficients to be encoded / decoded can be reduced by half. In addition, the filter tap refers to a width (horizontal dimension) or a height (vertical dimension) of a filter. Alternatively, the filter tap refers to both a width (dimension in a transverse direction) and a height (dimension in a longitudinal direction) of a two-dimensional filter. In addition, the filter can be symmetrical with respect to two or more specific directions.
[0243] When the filter has a mask form, the filter can be a two-dimensional geometric figure having a square / diamond shape, a non-square rectangular shape, a square shape, a trapezoidal shape, a diagonal line shape, a snowflake shape, a numeral symbol shape, a shamrock shape, a cross shape, a triangular shape, a pentagonal shape, a hexagonal shape, an octagonal shape, a decagonal shape, a dodecagonal shape, or any combination of these shapes. Alternatively, the filter shape can be a shape obtained by projecting a three-dimensional figure onto a two-dimensional plane.
[0244] The filter type indicates a filter selected from among a Wiener filter, a low-pass filter, a high-pass filter, a linear filter, a non-linear filter, and a bidirectional filter.
[0245] In the present disclosure, a Wiener filter will be described in detail among various filters. However, the present disclosure is not limited thereto, and a combination of the above-described filters can be used in an embodiment of the present disclosure.
[0246] As a filter type for adaptive in-loop filtering, a Wiener filter can be used. The Wiener filter is an optimal linear filter for effectively removing noise, blur, and distortion within a picture, thereby improving coding efficiency. The Wiener filter is designed to minimize distortion between an original picture and a reconstructed / decoded picture.
[0247] At least one of the filtering methods can be performed at the time of encoding processing or decoding processing. The encoding processing or the decoding processing refers to encoding or decoding performed in units of at least one of a slice, a parallel block, a parallel block group, a picture, a sequence, a CTU, a block, a CU, a PU, and a TU. At least one of the filtering methods is performed during encoding or decoding performed in units of a slice, a parallel block, a parallel block group, a picture, and the like. For example, the Wiener filter is used for adaptive in-loop filtering during encoding or decoding. That is, in the phrase "adaptive in-loop filtering", the term "in-loop" indicates that filtering is performed during encoding or decoding processing. When adaptive in-loop filtering is performed, a decoded picture that has undergone adaptive in-loop filtering can be used as a reference picture when a subsequent picture is encoded or decoded. In this case, since intra prediction or motion compensation is performed on a subsequent picture to be encoded / decoded by referring to a reconstructed picture that has undergone adaptive in-loop filtering, coding efficiency of the subsequent picture and coding efficiency of a current picture that has undergone in-loop filtering can be improved.
[0248] Also, at least one of the above filtering methods is performed when a CTU-based or block-based encoding or decoding process is performed. For example, a Wiener filter is used for adaptive in-loop filtering when a CTU-based or block-based encoding or decoding process is performed. That is, in the phrase "adaptive in-loop filtering", the term "in-loop" indicates that filtering is performed during a CTU-based or block-based encoding or decoding process. When adaptive in-loop filtering is performed on a per-CTU or per-block basis, a decoded CTU or block that has undergone adaptive in-loop filtering is used as a reference CTU or block for a subsequent CTU or block to be encoded / decoded. In this case, since intra prediction or motion compensation is performed on the subsequent CTU or block by referring to the current CTU or block to which adaptive in-loop filtering is applied, the coding efficiency of the current CTU or block to which in-loop filtering is applied is improved, and the coding efficiency of the subsequent CTU or block to be encoded / decoded is improved.
[0249] Also, at least one of the filtering methods can be performed as post-processing filtering after a decoding process is performed. For example, a Wiener filter can be used as a post-processing filter after a decoding process is performed. When a Wiener filter is used after a decoding process, the Wiener filter is applied to a reconstructed / decoded picture before the reconstructed / decoded picture is output (i.e., displayed). When post-processing filtering is performed, a decoded picture that has undergone post-processing filtering can not be used as a reference picture for a subsequent picture to be encoded / decoded.
[0250] Adaptive in-loop filtering cannot be performed on a per-block basis. That is, block-based filter adaptation cannot be performed. Here, block-based filter adaptation indicates that different filters are selected for different blocks, respectively. Block-based filter adaptation also indicates block classification.
[0251] Figure 7 is a flowchart illustrating a video decoding method according to an embodiment of the present application.
[0252] Referring to Figure 7 , the decoder decodes filter information for each coding unit (S701).
[0253] The filter information is not limited to filter information on a per-coding unit basis. The filter information also indicates filter information on a per-slice, parallel block, parallel block group, picture, sequence, CTU, block, CU, PU, or TU basis.
[0254] The filter information includes information on whether filtering is performed, filter coefficient values, the number of filters, the number of filter taps (filter length), filter shape information, filter type information, information on whether a fixed filter is used for a block classification index, and / or filter symmetry type information.
[0255] The filter shape information includes at least one shape selected from a diamond (square) shape, a rectangular shape, a square shape, a trapezoidal shape, a diagonal line shape, a snowflake shape, a numeral symbol shape, a shamrock shape, a cross shape, a triangular shape, a pentagonal shape, a hexagonal shape, an octagonal shape, a decagonal shape, and a dodecagonal shape.
[0256] The filter coefficient value includes a filter coefficient value of a geometric transform for each block classification unit.
[0257] On the other hand, examples of the filter symmetry type include at least one of point symmetry, horizontal symmetry, vertical symmetry, and diagonal symmetry.
[0258] In addition, the decoder performs block classification on the sample of the coding unit based on each block classification unit (step S702). Furthermore, the decoder assigns a block classification index to the block classification unit in the coding unit.
[0259] The block classification is not limited to the classification based on each coding unit. That is, the block classification can be performed in units of a slice, a parallel block, a parallel block group, a picture, a sequence, a CTU, a block, a CU, a PU, or a TU.
[0260] The gradient value is determined based on the directionality information and the activity information.
[0261] At least one of the directionality information and the activity information is determined according to a gradient value with respect to at least one of a vertical, a horizontal, a first diagonal line, and a second diagonal line direction.
[0262] On the other hand, the gradient value is obtained based on each block classification unit using a one-dimensional Laplacian operation.
[0263] The one-dimensional Laplacian operation is preferably a one-dimensional Laplacian operation in which an operation position is a sub-sampled position.
[0264] Optionally, the gradient value can be determined according to a temporal layer identifier.
[0265] In addition, the decoder filters the coding unit on which the block classification has been performed based on each block classification unit by using the filter information (S703).
[0266] The filtering target unit is not limited to the coding unit. That is, the filtering can be performed in units of a slice, a parallel block, a parallel block group, a picture, a sequence, a CTU, a block, a CU, a PU, or a TU.
[0267] Figure 8 is a flowchart illustrating a video encoding method according to an embodiment of the present application;
[0268] Referring to Figure 8The encoder classifies the samples in the coding unit into classes based on each block classification unit (step S801). In addition, the encoder assigns a block classification index to the block classification unit in each coding unit.
[0269] The basic unit for block classification is not limited to a coding unit. That is, block classification can be performed in units of a slice, a parallel block, a parallel block group, a picture, a sequence, a CTU, a block, a CU, a PU, or a TU.
[0270] The block classification index is determined based on directionality information and activity information.
[0271] At least one of the directionality information and the activity information is determined based on gradient values with respect to at least one of a vertical, a horizontal, a first diagonal, and a second diagonal direction.
[0272] The gradient values are obtained based on each block classification unit using a one-dimensional Laplacian operation.
[0273] The one-dimensional Laplacian operation is preferably a one-dimensional Laplacian operation in which the operation position is a sub-sampled position.
[0274] Optionally, the gradient values are determined according to a temporal layer identifier.
[0275] In addition, the encoder filters the coding unit samples classified based on each block classification unit by using filter information of the coding unit (S802).
[0276] The basic unit for filtering is not limited to a coding unit. That is, filtering can be performed in units of a slice, a parallel block, a parallel block group, a picture, a sequence, a CTU, a block, a CU, a PU, or a TU.
[0277] The filter information includes information on whether to perform filtering, filter coefficient values, the number of filters, the number of filter taps (filter length), filter shape information, filter type information, information on whether to use a fixed filter for a block classification index, and / or filter symmetry type information.
[0278] Examples of the filter shape include at least one of a diamond (square) shape, a rectangular shape, a square shape, a trapezoidal shape, a diagonal shape, a snowflake shape, a numeral symbol shape, a shamrock shape, a cross shape, a triangular shape, a pentagonal shape, a hexagonal shape, an octagonal shape, a decagonal shape, and a dodecagonal shape.
[0279] The filter coefficient values include filter coefficient values that are geometrically transformed based on each block classification unit.
[0280] Next, the encoder encodes the filter information (S803).
[0281] The filter information is not limited to the filter information on a per coding unit basis. The filter information can be filter information on a per slice, parallel block, parallel block group, picture, sequence, CTU, block, CU, PU, or TU basis.
[0282] At the encoder side, the adaptive in-loop filtering process can be divided into several sub-steps, such as block classification, filtering, and filter information encoding.
[0283] More specifically, at the encoder side, the adaptive in-loop filtering can be divided into several sub-steps, such as block classification, filter coefficient derivation, filtering execution determination, filter shape determination, filtering execution, and filter information encoding. The filter coefficient derivation, filtering execution determination, and filter shape determination do not fall within the scope of the present application. Therefore, these sub-steps are not described in depth, but are merely briefly described. Thus, at the encoder side, the in-loop filtering process is divided into block classification, filtering, filter information encoding, etc.
[0284] At the filter coefficient derivation step, the Wiener filter coefficients that minimize the distortion between the original picture and the filtered picture are derived. In this case, the Wiener filter coefficients are derived on a per block classification basis. In addition, the Wiener filter coefficients are derived according to at least one of the number of filter taps and the filter shape. When the Wiener filter coefficients are derived, the autocorrelation function for the reconstructed samples, the cross-correlation function for the original samples and the reconstructed samples, the autocorrelation matrix, and the cross-correlation matrix can be derived. The filter coefficients are calculated by deriving the Wiener-Hopf equation based on the autocorrelation matrix and the cross-correlation matrix. In this case, the filter coefficients are obtained by calculating the Wiener-Hopf equation based on Gaussian elimination or Cholesky decomposition.
[0285] At the filtering execution determination step, whether to perform the adaptive in-loop filtering on a per slice, picture, parallel block, or parallel block group basis, whether to perform the adaptive in-loop filtering on a per block basis, or whether not to perform the adaptive in-loop filtering is determined according to rate-distortion optimization. Here, the rate includes the filter information to be encoded. The distortion is the difference between the original picture and the reconstructed picture or the difference between the original picture and the filtered reconstructed picture. The distortion is represented by mean square error (MSE), sum of mean square error (SSE), sum of absolute difference, etc. At the filtering execution determination step, whether to perform filtering on a chroma component and whether to perform filtering on a luma component are determined.
[0286] At the filter shape determination step, when the in-loop adaptive filtering is applied, which filter shape to use, what tap number filter to use, etc. can be determined according to rate-distortion optimization.
[0287] In addition, at the decoder side, the adaptive in-loop filtering process is divided into the filter information decoding, block classification, and filtering steps.
[0288] Hereinafter, for the sake of avoiding redundant explanation, the filter information encoding step and the filter information decoding step will be collectively referred to as a filter information encoding / decoding step.
[0289] Hereinafter, the block classification step will be first described.
[0290] The block classification index is assigned within the reconstructed picture on a block basis of MxN size (or on a per block classification unit basis) so that the blocks within the reconstructed picture can be classified into L classes. Here, the block classification index can be assigned not only to the reconstructed / decoded picture but also to at least one of a restored / decoded slice, a restored / decoded parallel block group, a restored / decoded parallel block, a restored / decoded CTU, and a restored / decoded block.
[0291] Here, N, M, and L are each a positive integer. For example, N and M are each a positive integer selected from 2, 4, 8, 16, and 32, and L is a positive integer selected from 4, 8, 16, 20, 25, and 32. When N and M are the same integer 1, the block classification is performed on a sample basis rather than on a block basis. On the other hand, when N and M are different positive integers, the block of N x M size is a non-square shape. Alternatively, N and M can be the same positive integer.
[0292] For example, a total of 25 block classification indexes can be assigned to the reconstructed picture on a per 2 x 2 size block basis. For example, a total of 25 block classification indexes can be assigned to the reconstructed picture on a per 4 x 4 size block basis.
[0293] The block classification index is a value in the range from 0 to L-1, or can be a value in the range from 1 to L.
[0294] The block classification index C is determined based on at least one of a quantized activity value A of a directionality value D and an activity value A, and is represented by Equation 1. q
[0295] [Equation 1]
[0296] C = 5D + A q
[0297] In Equation 1, 5 is an exemplary constant value. The constant value can be represented by J. In this case, J is a positive integer having a value smaller than L.
[0298] For example, in one embodiment in which the block classification is performed on a per 2 x 2 size block basis, the sum of the one-dimensional Laplacian gradient values for the vertical direction is represented by g v , and the sums of the one-dimensional Laplacian gradient values for the horizontal direction, the first diagonal direction (angle 135°), and the second diagonal direction (angle 45°) are represented by g h , g d1 , and gd2 The Laplacian operations in the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction are represented by Expression 2, Expression 3, Expression 4, and Expression 5, respectively. The directionality value D and the activity value A are derived by using the sum of the gradient values. In one embodiment, the sum of the gradient values is used. Alternatively, any statistical value of the gradient values can be used instead of the sum of the gradient values.
[0299] [Equation 2]
[0300]
[0301] [Equation 3]
[0302]
[0303] [Equation 4]
[0304]
[0305] D1 k,l = |2R(k, l) - R(k - 1, l - 1) - R(k + 1, l + 1)|
[0306] [Equation 5]
[0307]
[0308] D2 k,l = |2R(k, l) - R(k - 1, l + 1) - R(k + 1, l - 1)|
[0309] In Equations 2 to 5, i and j represent the coordinates of the upper left position in the horizontal direction and the vertical direction, respectively, and R(i, j) represents the reconstructed sample value at position (i, j).
[0310] In Equations 2 to 5, k and l represent the horizontal operation range and the vertical operation range of the sum of the results V k,l , H k,l , D1 k,l , D2 k,l of the sample-based one-dimensional Laplacian operation generated for each direction. The result of the sample-based one-dimensional Laplacian operation for one direction represents the sample-based gradient value for the corresponding direction. That is, the result of the one-dimensional Laplacian operation represents the gradient value. The one-dimensional Laplacian operation is performed in each of the vertical, horizontal, first diagonal, and second diagonal directions, and the one-dimensional Laplacian operation indicates the gradient value for the corresponding direction. In addition, the results of the one-dimensional Laplacian operations for the vertical, horizontal, first diagonal, and second diagonal directions are represented by V k,l , H k,l , D1k,l D2 k,l D2
[0311] For example, k and l can be the same range. That is, the horizontal length and the vertical length of the operation range for computing the one-dimensional Laplacian sum can be the same.
[0312] Alternatively, k and l can be different ranges. That is, the horizontal length and the vertical length of the operation range for computing the one-dimensional Laplacian sum can be different.
[0313] As an example, k is a range from i-2 to i+3, and l is a range from j-2 to j+3. In this case, the operation range for computing the one-dimensional Laplacian sum is 6x6 size. In this case, the operation range for computing the one-dimensional Laplacian sum is greater than the size of the block classification unit.
[0314] As another example, k is a range from i-1 to i+2, and l is a range from j-1 to j+2. In this case, the operation range for computing the one-dimensional Laplacian sum is 4x4 size. In this case, the operation range for computing the one-dimensional Laplacian sum is greater than the size of the block classification unit.
[0315] As another example, k is a range from i to i+1, and l is a range from j to j+1. In this case, the operation range for computing the one-dimensional Laplacian sum is 2x2 size. In this case, the operation range for computing the one-dimensional Laplacian sum is equal to the size of the block classification unit.
[0316] For example, the operation range for computing the sum of the results of the one-dimensional Laplacian operation has a two-dimensional geometric shape selected from a rhombus, a rectangle, a square, a trapezoid, a diagonal line, a snowflake, a numeral symbol, a shamrock, a cross, a triangle, a pentagon, a hexagon, a decagon, and a dodecagon.
[0317] For example, the block classification unit has a two-dimensional geometric shape selected from a rhombus / square, a rectangle, a square, a trapezoid, a diagonal line, a snowflake, a numeral symbol, a shamrock, a cross, a triangle, a pentagon, a hexagon, a decagon, and a dodecagon.
[0318] For example, the range for computing the sum of the one-dimensional Laplacian operation is SxT size. In this case, S and T are each zero or a positive integer.
[0319] In addition, D1 representing the first diagonal line and D2 representing the second diagonal line can respectively refer to D0 representing the first diagonal line and D1 representing the second diagonal line.
[0320] For example, in one embodiment of performing block classification for each 4x4 size block, the sum of gradient values for vertical, horizontal, first diagonal and second diagonal directions are calculated by Equation 6, Equation 7, Equation 8, Equation 9 based on one-dimensional Laplacian operation v h d1 d2 The directionality value D and the activity value A are derived by using the sum of gradient values. In one embodiment, the sum of gradient values is used. Alternatively, any statistical value of gradient values can be used instead of the sum of gradient values.
[0321] [Equation 6]
[0322]
[0323] [Equation 7]
[0324]
[0325] [Equation 8]
[0326]
[0327] D1 k,l = |2R(k,l) - R(k-1,l-1) - R(k+1,l+1)|
[0328] [Equation 9]
[0329]
[0330] D2 k,l = |2R(k,l) - R(k-1,l+1) - R(k+1,l-1)|
[0331] In Equation 6 to Equation 9, i and j represent the coordinates of the top-left position in horizontal and vertical directions, respectively, and R(i,j) represents the reconstructed sample value at position (i,j).
[0332] In Equation 6 to 9, k and l represent the results of one-dimensional Laplacian operation based on samples for calculating for each direction V k,l k,l k,l k,l horizontal and vertical operation ranges of the sum of the one-dimensional Laplacian operations. The result of the sample-based one-dimensional Laplacian operation for one direction represents a sample-based gradient value for the corresponding direction. That is, the result of the one-dimensional Laplacian operation represents a gradient value. The one-dimensional Laplacian operation is performed for each of the vertical, horizontal, first diagonal, and second diagonal directions, and the one-dimensional Laplacian operation indicates a gradient value for the corresponding direction. In addition, the results of the one-dimensional Laplacian operations for the vertical, horizontal, first diagonal, and second diagonal directions are respectively represented as V kl kl kl kl .
[0333] For example, k and l can be the same range. That is, the horizontal length and the vertical length of the operation range for which the sum of the one-dimensional Laplacian operations is calculated can be the same.
[0334] Alternatively, k and l can be different ranges. That is, the horizontal length and the vertical length of the operation range for which the sum of the one-dimensional Laplacian operations is calculated can be different.
[0335] As an example, k is a range from i-2 to i+5, and l is a range from j-2 to j+5. In this case, the operation range for which the sum of the one-dimensional Laplacian operations is calculated is 8x8 size. In this case, the operation range for which the sum of the one-dimensional Laplacian operations is calculated is greater than the size of the block classification unit.
[0336] As another example, k is a range from i to i+3, and l is a range from j to j+3. In this case, the operation range for which the sum of the one-dimensional Laplacian operations is calculated is 4x4 size. In this case, the operation range for which the sum of the one-dimensional Laplacian operations is calculated is equal to the size of the block classification unit.
[0337] For example, the operation range of the sum of the results of the one-dimensional Laplacian operations has a two-dimensional geometric shape selected from a diamond, a rectangle, a square, a trapezoid, a diagonal line, a snowflake, a numeral symbol, a shamrock, a cross, a triangle, a pentagon, a hexagon, a decagon, and a dodecagon.
[0338] For example, the operation range of the sum of the one-dimensional Laplacian operations is SxT size. In this case, S and T are each zero or a positive integer.
[0339] For example, the block classification unit has a two-dimensional geometric shape selected from a diamond / square, a rectangle, a square, a trapezoid, a diagonal line, a snowflake, a numeral symbol, a shamrock, a cross, a triangle, a pentagon, a hexagon, an octagon, a decagon, and a dodecagon.
[0340] Figure 9 is a diagram showing an exemplary method of determining gradient values for horizontal, vertical, first diagonal, and second diagonal directions, respectively.
[0341] As shown in Figure 9 , when block classification is performed based on each 4x4 size block, the sum g v , g h , g d1 , g d2 of gradient values for vertical, horizontal, first diagonal, and second diagonal directions can be calculated. Here, V, H, D1, and D2 represent the results of one-dimensional Laplacian operation based on samples for vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operation is performed for vertical, horizontal, first diagonal, and second diagonal directions at positions V, H, D1, and D2, respectively. In Figure 9 , the block classification index C is assigned to the 4x4 size block which is shaded. In this case, the operation range for calculating one-dimensional Laplacian sum is larger than the size of the block classification unit. Here, the thin solid rectangle represents the reconstructed sample position, and the thick solid rectangle represents the operation range for calculating one-dimensional Laplacian sum.
[0342] For example, in one embodiment where block classification is performed based on each 4x4 size block, one-dimensional Laplacian sum g v , g h , g d1 , g d2 of gradient values for vertical, horizontal, first diagonal, and second diagonal directions is calculated by Equations 10 to 13, respectively. The gradient values are represented based on sub-samples to reduce the calculation complexity of block classification. The directionality value D and the activity value A are derived by using the sum of gradient values. In one embodiment, the sum of gradient values is used. Alternatively, any statistical value of gradient values can be used instead of the sum of gradient values.
[0343] [Equation 10]
[0344]
[0345] k = i-2, i, i+2, i+4, l = j-2,..., j+5
[0346] [Equation 11]
[0347]
[0348] k = i-2,..., i+5, l = j-2, j, j+2, j+4
[0349] [Equation 12]
[0350]
[0351] D1 k,l = |2R(k,l) - R(k-1,l-1) - R(k+1,l+1)|,
[0352] k = i-2,..., i+5, l = j-2,..., j+5
[0353]
[0354] [Equation 13]
[0355]
[0356] D2 k,l = |2R(k,l) - R(k-1,l+1) - R(k+1,l-1)|,
[0357] k = i-2,..., i+5, l = j-2,..., j+5
[0358]
[0359] In Equations 10 to 13, i and j respectively denote coordinates of a top-left position in a horizontal direction and a vertical direction, and R(i,j) denotes a reconstructed sample value at a position (i,j).
[0360] In Equations 10 to 13, k and l respectively denote a horizontal operation range and a vertical operation range in which a sum of results V k,l , H k,l , D1 k,l , D2 k,l of sample-based one-dimensional Laplacian operations are calculated. The result of the sample-based one-dimensional Laplacian operation for one direction denotes a sample-based gradient value for the corresponding direction. That is, the result of the one-dimensional Laplacian operation denotes a gradient value. The one-dimensional Laplacian operation is performed for each of the vertical, horizontal, first diagonal, and second diagonal directions, and the one-dimensional Laplacian operation indicates a gradient value for the corresponding direction. In addition, the results of the one-dimensional Laplacian operations for the vertical, horizontal, first diagonal, and second diagonal directions are respectively denoted as V k,l , H k,l , D1 k,l , D2 k,l .
[0361] For example, k and l can be the same range. That is, a horizontal length and a vertical length of an operation range in which the sum of the one-dimensional Laplacian operations is calculated are the same.
[0362] Alternatively, k and l can be different ranges. That is, the horizontal length and the vertical length of the operation range in which the sum of the one-dimensional Laplacian operation is calculated can be different.
[0363] As an example, k is a range from i-2 to i+5, and l is a range from j-2 to j+5. In this case, the operation range in which the sum of the one-dimensional Laplacian operation is calculated is 8x8 size. In this case, the operation range in which the one-dimensional Laplacian sum is calculated is greater than the size of the block classification unit.
[0364] As another example, k is a range from i to i+3, and l is a range from j to j+3. In this case, the operation range in which the sum of the one-dimensional Laplacian operation is calculated is 4x4 size. In this case, the operation range in which the sum of the one-dimensional Laplacian operation is calculated is equal to the size of the block classification unit.
[0365] For example, the operation range in which the sum of the result of the one-dimensional Laplacian operation is calculated has a two-dimensional geometric shape selected from a diamond, a rectangle, a square, a trapezoid, a diagonal line, a snowflake, a numeral symbol, a shamrock, a cross, a triangle, a pentagon, a hexagon, a decagon, and a dodecagon.
[0366] For example, the operation range in which the sum of the one-dimensional Laplacian operation is calculated has SxT size. In this case, S and T are zero or a positive integer.
[0367] For example, the block classification unit has a two-dimensional geometric shape selected from a diamond / square, a rectangle, a square, a trapezoid, a diagonal line, a snowflake, a numeral symbol, a shamrock, a cross, a triangle, a pentagon, a hexagon, an octagon, a decagon, and a dodecagon.
[0368] According to an embodiment of the present invention, the sample-based gradient value calculation method can calculate a gradient value by performing a one-dimensional Laplacian operation on samples within an operation range in a corresponding direction. Here, a statistical value of the gradient value can be calculated by calculating a statistical value of a result of the one-dimensional Laplacian operation performed on at least one of the samples within the operation range in which the sum of the one-dimensional Laplacian operation is calculated. In this case, the statistical value is any one of a sum, a weighted sum, and an average.
[0369] For example, in order to calculate a gradient value for a horizontal direction, a one-dimensional Laplacian operation is performed at each sample position within the operation range in which the sum of the one-dimensional Laplacian operation is calculated. In this case, the gradient value for the horizontal direction can be calculated at intervals of P rows within the operation range in which the sum of the one-dimensional Laplacian operation is calculated. Here, P is a positive integer.
[0370] Alternatively, to calculate the gradient value in the vertical direction, a one-dimensional Laplacian operation is performed at each sample position on the columns within the operation range for calculating the sum of the one-dimensional Laplacian operation. In this case, the gradient value in the vertical direction can be calculated at intervals of P columns within the operation range for calculating the sum of the one-dimensional Laplacian operation. Here, P is a positive integer.
[0371] Further alternatively, to calculate the gradient value in the first diagonal line direction, a one-dimensional Laplacian operation is performed on the sample positions at intervals of P columns or Q rows in at least one of the horizontal direction and the vertical direction within the operation range for calculating the sum of the one-dimensional Laplacian operation, thereby obtaining the gradient value in the first diagonal line direction. Here, P and Q are zero or positive integers.
[0372] Further alternatively, to calculate the gradient value in the second diagonal line direction, a one-dimensional Laplacian operation is performed on the sample positions at intervals of P columns or Q rows in at least one of the horizontal direction and the vertical direction within the operation range for calculating the sum of the one-dimensional Laplacian operation, thereby obtaining the gradient value in the second diagonal line direction. Here, P and Q are zero or positive integers.
[0373] According to an embodiment of the present application, a sample-based gradient value calculation method can calculate a gradient value by performing a one-dimensional Laplacian operation on at least one sample within an operation range for calculating the sum of the one-dimensional Laplacian operation. Here, a statistical value of the gradient value can be calculated by calculating a statistical value of the result of the one-dimensional Laplacian operation performed on at least one of the samples within the operation range for calculating the sum of the one-dimensional Laplacian operation. In this case, the statistical value is any one of a sum, a weighted sum, and an average.
[0374] For example, to calculate the gradient value, a one-dimensional Laplacian operation is performed at each sample position within the operation range for calculating the sum of the one-dimensional Laplacian operation. In this case, the gradient value can be calculated at intervals of P rows within the operation range for calculating the sum of the one-dimensional Laplacian operation. Here, P is a positive integer.
[0375] Alternatively, to calculate the gradient value, a one-dimensional Laplacian operation is performed at each sample position on the columns within the operation range for calculating the sum of the one-dimensional Laplacian operation. In this case, the gradient value can be calculated at intervals of P rows within the operation range for calculating the sum of the one-dimensional Laplacian operation. Here, P is a positive integer.
[0376] Further alternatively, to calculate the gradient value, a one-dimensional Laplacian operation is performed on the sample positions at intervals of P columns or Q rows in at least one of the horizontal direction and the vertical direction within the operation range for calculating the sum of the one-dimensional Laplacian operation, thereby obtaining the gradient value. Here, P and Q are zero or positive integers.
[0377] Further optionally, to calculate the gradient value, within the operation range for calculating the sum of the one-dimensional Laplacian operation, the sample positions are subjected to the one-dimensional Laplacian operation at intervals of P columns and Q rows in the horizontal direction and the vertical direction, thereby obtaining the gradient value. Here, P and Q are zero or a positive integer.
[0378] On the other hand, the gradient refers to at least one of a gradient with respect to the horizontal direction, a gradient with respect to the vertical direction, a gradient with respect to the first diagonal line direction, and a gradient with respect to the second diagonal line direction.
[0379] Figures 10 to 12 is a diagram illustrating a sub-sampling-based method of determining gradient values with respect to the horizontal, vertical, first diagonal, and second diagonal directions.
[0380] As shown in Figure 10 , when block classification is performed based on each 2x2 stored block, at least one of the sums g v , g h , g d1 , g d2 of the gradient values with respect to the vertical, horizontal, first diagonal, and second diagonal directions can be calculated based on sub-sampling. Here, V, H, D1, and D2 respectively denote the results of the sample-based one-dimensional Laplacian operation with respect to the vertical direction, the horizontal direction, the first diagonal line direction, and the second diagonal line direction. That is, the one-dimensional Laplacian operation is performed with respect to the vertical, horizontal, first diagonal, and second diagonal directions at positions V, H, D1, and D2, respectively. In addition, the positions at which the one-dimensional Laplacian operation is performed are the positions of sub-sampling. In Figure 10 , the block classification index C is assigned to the 2x2-sized block that is shaded. In this case, the operation range for calculating the one-dimensional Laplacian sum is greater than the size of the block classification unit. Here, the thin solid line rectangle denotes the reconstructed sample position, and the thick solid line rectangle denotes the operation range for calculating the one-dimensional Laplacian sum.
[0381] In the drawings of the present application, the positions not indicated by V, H, D1, or D2 are sample positions at which the one-dimensional Laplacian operation is not performed in the direction. That is, the one-dimensional Laplacian operation is performed in each direction only at the sample positions indicated by V, H, D1, or D2. When the one-dimensional Laplacian operation is not performed, the result of the one-dimensional Laplacian operation at the corresponding sample position is determined to be a certain value, for example, H. Here, H can be at least one of a negative integer, 0, and a positive integer.
[0382] As shown in Figure 11 , when block classification is performed based on a 4x4-sized block, at least one of the sums g v , g h , gd1 g d2 At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations for the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Furthermore, the positions where the one-dimensional Laplacian operations are performed are the sub-sampling positions. Figure 11 In this context, the block classification index C is assigned to the shaded 4×4 block. In this case, the computational range for calculating the one-dimensional Laplacian sum is larger than the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplacian sum.
[0383] like Figure 12 As shown, when performing block classification based on 4×4 size blocks, the sum of gradient values g for the vertical, horizontal, first diagonal, and second diagonal directions can be calculated based on subsampling. v g h g d1 g d2 At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations for the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Furthermore, the positions where the one-dimensional Laplacian operations are performed are the sub-sampling positions. Figure 12 In this context, the block classification index C is assigned to the shaded 4×4 block. In this case, the operational range for calculating the one-dimensional Laplacian sum is equal to the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample locations, and thick solid rectangles represent the operational range for calculating the one-dimensional Laplacian sum.
[0384] According to an embodiment of the present invention, the gradient value can be calculated by performing a one-dimensional Laplace operation on sample points located at specific positions within a block of size N×M based on subsampling. In this case, the specific position can be at least one of an absolute position and a relative position within the block. Here, the statistical value of the gradient value can be calculated by calculating the statistical value of the result of performing a one-dimensional Laplace operation on at least one sample point within the range of the calculation of the one-dimensional Laplace sum. In this case, the statistical value is any one of the sum, weighted sum, and average.
[0385] For example, an absolute position represents the top-left position within an N×M block.
[0386] Optionally, the absolute position represents the bottom right position within an N×M block.
[0387] Further optionally, the relative position indicates a center position within the NxM block.
[0388] According to embodiments of the present invention, the gradient values can be calculated by performing one-dimensional Laplacian operations on R samples within an NxM sized block based on sub-sampling. In this case, P and Q are zero or positive integers. Further, R is equal to or smaller than the product of N and M. Here, the statistical value of the gradient values can be calculated by calculating a statistical value of results of one-dimensional Laplacian operations performed on at least one of the samples within the operation range for calculating one-dimensional Laplacian sum. In this case, the statistical value is any one of sum, weighted sum and average.
[0389] For example, when R is 1, one-dimensional Laplacian operations are performed only on one sample within the NxM block.
[0390] Optionally, when R is 2, one-dimensional Laplacian operations are performed only on two samples within the NxM block.
[0391] Further optionally, when R is 4, one-dimensional Laplacian operations are performed only on 4 samples within each NxM sized block.
[0392] According to embodiments of the present invention, the gradient values can be calculated by performing one-dimensional Laplacian operations on R samples within an NxM sized block based on sub-sampling. In this case, P and Q are zero or positive integers. Further, R is equal to or smaller than the product of N and M. Here, the statistical value of the gradient values can be calculated by calculating a statistical value of results of one-dimensional Laplacian operations performed on at least one of the samples within the operation range for calculating one-dimensional Laplacian sum. In this case, the statistical value is any one of sum, weighted sum and average.
[0393] For example, when R is 1, one-dimensional Laplacian operations are performed only on one sample within each NxM sized block for calculating one-dimensional Laplacian sum.
[0394] Optionally, when R is 2, one-dimensional Laplacian operations are performed only on two samples within each NxM sized block for calculating one-dimensional Laplacian sum.
[0395] Further optionally, when R is 4, one-dimensional Laplacian operations are performed only on 4 samples within each NxM sized block for calculating one-dimensional Laplacian sum.
[0396] Figures 13 to 18 is a diagram illustrating an exemplary method based on sub-sampling for determining gradient values along horizontal, vertical, first diagonal and second diagonal directions.
[0397] As Figure 13At least one of g v , g h , g d1 , g d2 Here, V, H, D1 and D2 respectively denote results of the sample-based one-dimensional Laplacian operation for the vertical direction, the horizontal direction, the first diagonal direction and the second diagonal direction. That is, the one-dimensional Laplacian operation is performed along the vertical direction, the horizontal direction, the first diagonal direction and the second diagonal direction at positions V, H, D1 and D2, respectively. In addition, the positions at which the one-dimensional Laplacian operation is performed can be the positions of the sub-sampling. In Figure 13 , the block classification index C is assigned to the 4x4-sized block which is hatched. In this case, the operation range for calculating the sum of the one-dimensional Laplacian operation is equal to the size of the block classification unit. Here, the thin solid rectangle denotes the reconstructed sample position, and the thick solid rectangle denotes the operation range for calculating the sum of the one-dimensional Laplacian operation.
[0398] As shown in Figure 14 , when the block classification is performed on a per 4x4-sized block basis, the sum g v , g h , g d1 , g d2 of the gradient values for the vertical, horizontal, first diagonal and second diagonal directions can be calculated by using the samples at certain positions within each NxM-sized block on a sub-sampling basis. Here, V, H, D1 and D2 respectively denote results of the sample-based one-dimensional Laplacian operation for the vertical direction, the horizontal direction, the first diagonal direction and the second diagonal direction. That is, the one-dimensional Laplacian operation is performed along the vertical direction, the horizontal direction, the first diagonal direction and the second diagonal direction at positions V, H, D1 and D2, respectively. In addition, the positions at which the one-dimensional Laplacian operation is performed are the positions of the sub-sampling. In Figure 14 , the block classification index C is assigned to the 4x4-sized block which is hatched. In this case, the operation range for calculating the sum of the one-dimensional Laplacian operation is smaller than the size of the block classification unit. Here, the thin solid rectangle denotes the reconstructed sample position, and the thick solid rectangle denotes the operation range for calculating the sum of the one-dimensional Laplacian operation.
[0399] As shown in Figure 15 , when the block classification is performed on a per 4x4-sized block basis, the sum g v , gh d1 d2 at least one of gV, gH, gD1, and gD2. Here, V, H, D1, and D2 respectively denote results of a sample-based one-dimensional Laplacian operation for a vertical direction, a horizontal direction, a first diagonal direction, and a second diagonal direction. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 respectively along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. In addition, the positions at which the one-dimensional Laplacian operations are performed can be sub-sampled positions. In Figure 15 In this case, the operation range for calculating the sum of the one-dimensional Laplacian operations is smaller than the size of the block classification unit. Here, the thin solid rectangle denotes the reconstructed sample positions, and the thick solid rectangle denotes the operation range for calculating the sum of the one-dimensional Laplacian operations.
[0400] As shown in Figure 16 , when the block classification is performed on a per 4x4 size block basis, the sum gV, gH, gD1, and gD2 of the gradient values for the vertical, horizontal, first diagonal, and second diagonal directions can be calculated using samples at specific positions within each NxM size block on a sub-sampling basis. v h d1 d2 at least one of gV, gH, gD1, and gD2. Here, V, H, D1, and D2 respectively denote results of a sample-based one-dimensional Laplacian operation for a vertical direction, a horizontal direction, a first diagonal direction, and a second diagonal direction. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 respectively along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. In addition, the positions at which the one-dimensional Laplacian operations are performed are sub-sampled positions. In Figure 16 In this case, the operation range for calculating the sum of the one-dimensional Laplacian operations is smaller than the size of the block classification unit. Here, the thin solid rectangle denotes the reconstructed sample positions, and the thick solid rectangle denotes the operation range for calculating the sum of the one-dimensional Laplacian operations.
[0401] As shown in Figure 17 , when the block classification is performed on a per 4x4 size block basis, the sum gV, gH, gD1, and gD2 of the gradient values for the vertical, horizontal, first diagonal, and second diagonal directions can be calculated using samples at specific positions within each NxM size block on a sub-sampling basis. v h d1 d2 At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Furthermore, the position where the one-dimensional Laplacian operation is performed can be the position of a subsample. Figure 17 In this example, the block classification index C is assigned to a shaded 4×4 block. In this case, the computational range for calculating the sum of the one-dimensional Laplacian operations can be smaller than the size of the block classification unit. Here, since the computational range for calculating the sum of the one-dimensional Laplacian operations is 1×1, the gradient value can be calculated without calculating the sum of the one-dimensional Laplacian operations. Here, thin solid rectangles represent the reconstructed sample point locations, and thick solid rectangles represent the computational range for calculating the sum of the one-dimensional Laplacian operations.
[0402] like Figure 18 As shown, when performing block classification based on 2×2 size blocks, the sum of gradient values g for the vertical, horizontal, first diagonal, and second diagonal directions can be calculated by using samples at specific locations within each N×M size block based on subsampling. v g h g d1 g d2 At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Furthermore, the positions where the one-dimensional Laplacian operation is performed can be the positions of sub-sampling. Figure 18 In this context, the block classification index C is assigned to a shaded 2×2 block. In this case, the computational range for calculating the sum of the one-dimensional Laplacian operations can be smaller than the size of the block classification unit. Here, since the computational range for calculating the sum of the one-dimensional Laplacian operations is 1×1, the gradient value can be calculated without calculating the sum of the one-dimensional Laplacian operations. Here, thin solid rectangles represent the reconstructed sample point locations, and thick solid rectangles represent the computational range for calculating the sum of the one-dimensional Laplacian operations.
[0403] Figures 19 to 30is a diagram showing a method of determining gradient values with respect to horizontal, vertical, first diagonal, and second diagonal directions at a certain sample position. The certain sample position can be a sub-sampled sample position within a block classification unit, or can be a sub-sampled sample position within a calculation range of a sum of one-dimensional Laplacian operations. In addition, the certain sample position is a sample position within each block. Alternatively, the certain sample position can vary from block to block. Furthermore, the certain sample position can be the same regardless of a direction of one-dimensional Laplacian operations to be calculated. In addition, the certain sample position can be the same for each block regardless of a direction of one-dimensional Laplacian operations.
[0404] As shown in Figure 19 , when block classification is performed based on a block of a 4x4 size, a sum g v , g h , g d1 , g d2 of gradient values is calculated at one or more certain sample positions. Here, V, H, D1, and D2 respectively denote results of sample-based one-dimensional Laplacian operations for vertical, horizontal, first diagonal, and second diagonal directions. That is, one-dimensional Laplacian operations are performed along the vertical, horizontal, first diagonal, and second diagonal directions at positions V, H, D1, and D2, respectively. In addition, the positions at which the one-dimensional Laplacian operations are performed can be sub-sampled positions. In Figure 19 , a block classification index C is assigned to a 4x4 size block which is shaded. In this case, a calculation range for calculating a sum of one-dimensional Laplacians can be greater than a size of a block classification unit. Here, a thin solid line rectangle denotes a reconstructed sample position, and a thick solid line rectangle denotes a calculation range for calculating a sum of one-dimensional Laplacians.
[0405] As shown in Figure 19 , the certain sample positions at which one-dimensional Laplacian operations are performed are the same regardless of a direction of one-dimensional Laplacian operations. In addition, as shown in Figure 19 , a pattern of sample positions at which one-dimensional Laplacian operations are performed can be referred to as a checkerboard pattern or a quincunx pattern. In addition, all of the sample positions at which one-dimensional Laplacian operations are performed are even or odd sample positions in both horizontal (X-axis direction) and vertical (Y-axis direction) directions within a calculation range for calculating a sum of one-dimensional Laplacians within a block classification unit or a block unit.
[0406] As shown in Figure 20 , when block classification is performed based on a block of a 4x4 size, a sum g v , g h , g d1 , g d2At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations performed along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Furthermore, the positions where the one-dimensional Laplacian operation is performed can be the positions of the subsamples. Figure 20 In this context, the block classification index C is assigned to the shaded 4×4 block. In this case, the computational range of the one-dimensional Laplace sum can be larger than the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample point locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplace sum.
[0407] like Figure 20 As shown, regardless of the direction of the one-dimensional Laplace operation, the specific sample point position for performing the one-dimensional Laplace operation is the same. Additionally, as... Figure 20 The pattern of sample point positions for performing a one-dimensional Laplace calculation, as shown, can be referred to as a checkerboard pattern or a cloverleaf pattern. Furthermore, the sample point positions for performing a one-dimensional Laplace operation are the even or odd number of sample point positions in both the horizontal (X-axis) and vertical (Y-axis) directions within the block classification unit or the one-dimensional Laplace operation range of the block unit.
[0408] like Figure 21 As shown, when performing block classification based on 4×4 size blocks, the sum of gradient values g is calculated at one or more specific sample locations. v g h g d1 g d2 At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Furthermore, the position where the one-dimensional Laplacian operation is performed can be the position of a subsample. Figure 21 In this context, the block classification index C is assigned to a shaded 4×4 block. In this case, the computational range for calculating the one-dimensional Laplacian sum can be larger than the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample point locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplacian sum.
[0409] like Figure 22 As shown, when performing block classification based on 4×4 size blocks, the sum of gradient values g is calculated at one or more specific sample locations. v g h, g d1 , g d2 at least one of g v , g h , g d1 , and g d2 . Here, V, H, D1, and D2 respectively denote results of the one-dimensional Laplacian operation based on samples for the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. That is, the one-dimensional Laplacian operation is performed at positions V, H, D1, and D2 respectively along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. In addition, the positions at which the one-dimensional Laplacian operation is performed can be sub-sampled positions. In Figure 22 , the block classification index C is assigned to the 4x4-sized block which is hatched. In this case, the operation range for calculating the one-dimensional Laplacian sum can be larger than the size of the block classification unit. Here, the thin solid rectangle denotes the reconstructed sample positions, and the thick solid rectangle denotes the operation range for calculating the one-dimensional Laplacian sum.
[0410] As shown in Figure 23 , when the block classification is performed based on each 4x4-sized block, at least one of g v , g h , g d1 , and g d2 is calculated at one or more specific sample positions. Here, V, H, D1, and D2 respectively denote results of the one-dimensional Laplacian operation based on samples for the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. That is, the one-dimensional Laplacian operation is performed at positions V, H, D1, and D2 respectively along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. In addition, the positions at which the one-dimensional Laplacian operation is performed can be sub-sampled positions. In Figure 23 , the block classification index C is assigned to the 4x4-sized block which is hatched. In this case, the operation range for calculating the one-dimensional Laplacian sum can be equal to the size of the block classification unit. Here, the thin solid rectangle denotes the reconstructed sample positions, and the thick solid rectangle denotes the operation range for calculating the one-dimensional Laplacian operation sum.
[0411] As shown in Figure 23 , the specific sample positions at which the one-dimensional Laplacian operation is performed are the same regardless of the one-dimensional Laplacian operation direction. In addition, as shown in Figure 23 , the pattern of the sample positions at which the one-dimensional Laplacian operation is performed can be referred to as a checkerboard pattern or a quincunx pattern. In addition, in the one-dimensional Laplacian operation range in the block classification unit or the block unit, all of the sample positions at which the one-dimensional Laplacian operation is performed are even or odd sample positions in two directions or any one direction among the horizontal direction (X-axis direction) and the vertical direction (Y-axis direction).
[0412] As shown in Figure 24As shown in FIG. 10, when the block classification is performed based on each 4x4 size block, at least one of sums g v , g h , g d1 , and g d2 of gradient values are calculated at one or more specific sample positions. Here, V, H, D1, and D2 respectively denote results of sample-based one-dimensional Laplacian operations for vertical, horizontal, first diagonal, and second diagonal directions. That is, one-dimensional Laplacian operations are performed along the vertical, horizontal, first diagonal, and second diagonal directions at positions V, H, D1, and D2, respectively. In addition, the positions at which the one-dimensional Laplacian operations are performed can be sub-sampled positions. In Figure 24 , a block classification index C is assigned to the 4x4 size block which is shaded. In this case, the operation range for calculating the one-dimensional Laplacian sum can be equal to the size of the block classification unit. Here, the thin solid rectangle denotes the reconstructed sample position, and the thick solid rectangle denotes the operation range for calculating the one-dimensional Laplacian sum.
[0413] As shown in Figure 24 , regardless of the one-dimensional Laplacian operation direction, the specific sample positions at which the one-dimensional Laplacian operations are performed are the same. In addition, as shown in Figure 24 , the pattern of the sample positions at which the one-dimensional Laplacian operations are performed can be referred to as a checkerboard pattern or a quincunx pattern. In addition, the sample positions at which the one-dimensional Laplacian operations are performed are even or odd sample positions in both of the horizontal (X-axis direction) and vertical (Y-axis direction) directions or either one of the directions in the one-dimensional Laplacian operation range of the block classification unit or the block unit.
[0414] As shown in Figure 25 , when the block classification is performed based on each 4x4 size block, at least one of sums g v , g h , g d1 , and g d2 of gradient values are calculated at one or more specific sample positions. Here, V, H, D1, and D2 respectively denote results of sample-based one-dimensional Laplacian operations for vertical, horizontal, first diagonal, and second diagonal directions. That is, one-dimensional Laplacian operations are performed along the vertical, horizontal, first diagonal, and second diagonal directions at positions V, H, D1, and D2, respectively. In addition, the positions at which the one-dimensional Laplacian operations are performed can be sub-sampled positions. In Figure 25 , a block classification index C is assigned to the 4x4 size block which is shaded. In this case, the operation range for calculating the one-dimensional Laplacian sum can be equal to the size of the block classification unit. Here, the thin solid rectangle denotes the reconstructed sample position, and the thick solid rectangle denotes the operation range for calculating the one-dimensional Laplacian sum.
[0415] like Figure 26 As shown, when performing block classification based on each 4×4 block, the sum of gradient values g is calculated at one or more specific sample locations. v g h g d1 and g d2 At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Figure 26 In this context, the block classification index C is assigned to a shaded 4×4 block. In this case, the computational range for calculating the one-dimensional Laplace sum can be equal to the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplace sum. A specific sample location can refer to each sample location within the block classification unit.
[0416] like Figure 27 As shown, when performing block classification based on each 4×4 block, the sum of gradient values g is calculated at one or more specific sample locations. v g h g d1 and g d2 At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Furthermore, the position where the one-dimensional Laplacian operation is performed can be the position of a subsample. Figure 27 In this context, the block classification index C is assigned to the shaded 4×4 block. In this case, the computational range for calculating the one-dimensional Laplacian sum can be equal to the size of the block classification unit. Here, the thin solid rectangle represents the location of the reconstructed sample point, and the thick solid rectangle represents the computational range for calculating the one-dimensional Laplacian sum.
[0417] like Figure 28 As shown, when performing block classification based on each 4×4 size block, the sum of gradient values g is calculated at one or more specific sample locations. v g h g d1 and g d2At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Figure 28 In this context, the block classification index C is assigned to a shaded 4×4 block. In this case, the computational range for calculating the one-dimensional Laplace sum can be larger than the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplace sum. A specific sample location can refer to each sample location within the block classification unit.
[0418] like Figure 29 As shown, when performing block classification based on each 4×4 size block, the sum of gradient values g is calculated at one or more specific sample locations. v g h g d1 and g d2 At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Figure 29 In this context, the block classification index C is assigned to a shaded 4×4 block. In this case, the computational range for calculating the one-dimensional Laplace sum can be larger than the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplace sum. A specific sample location can refer to each sample location within the block classification unit.
[0419] like Figure 30 As shown, when performing block classification based on each 4×4 size block, the sum of gradient values g is calculated at one or more specific sample locations. v g h g d1 and g d2 At least one of them. Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplacian operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplacian operations are performed at positions V, H, D1, and D2 along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Furthermore, the position where the one-dimensional Laplacian operation is performed can be the position of a subsample. Figure 30In this case, a block classification index C is assigned to the shaded 4x4 size block. In this case, the calculation range of the one-dimensional Laplacian sum can be greater than the size of the block classification unit. Here, the thin solid line rectangle represents the reconstructed sample position, and the thick solid line rectangle represents the calculation range of the one-dimensional Laplacian sum.
[0420] According to embodiments of the present application, at least one of the methods of calculating gradient values can be performed based on a temporal layer identifier.
[0421] For example, when block classification is performed based on each 2x2 size block, Equations 2 to 5 can be commonly expressed by one equation as shown in Equation 14.
[0422] [Equation 14]
[0423]
[0424] In Equation 14, dir represents a horizontal direction, a vertical direction, a first diagonal direction, and a second diagonal direction, and g dir represents each of the sums of gradient values along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. In addition, i and j represent a horizontal position and a vertical position in the 2x2 size block, respectively, and G dir represents each of the results of one-dimensional Laplacian operations along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction.
[0425] In this case, when block classification is performed based on each 2x2 size block within the current picture (or reconstructed picture) in a case where the temporal layer identifier of the current picture (or reconstructed picture) indicates the top layer, Equation 14 can be expressed as Equation 15.
[0426] [Equation 15]
[0427] g 2×2,dir = |G dir (i0,j0)|
[0428] In Equation 15, G dir (i0,j0) represents gradient values at the upper left position within the 2x2 size block along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction.
[0429] Figure 31 is a diagram illustrating an exemplary method of determining gradient values along horizontal, vertical, first diagonal, and second diagonal directions for a case where the temporal layer identifier indicates the top layer.
[0430] Referring to Figure 31 , the sums gv , g h , g d1 and g d2 The gradient can be simplified by calculating the gradient only at the top-left sample position (i.e., the hatched sample position) within each 2x2-sized block.
[0431] According to embodiments of the present application, a statistical value of gradient values is calculated by calculating a weighted sum while applying weights to results of one-dimensional Laplacian operations performed on one or more samples within a range of a sample on which the one-dimensional Laplacian operation is calculated. In this case, at least one of a weighted average, a median, a minimum value, a maximum value, and a mode value can be used instead of the weighted sum.
[0432] The application of the weights or the calculation of the weighted sum can be determined based on various conditions or encoding parameters associated with the current block and the neighboring blocks.
[0433] For example, the weighted sum can be calculated in units of at least one of a sample, a group of samples, a line, and a block. In this case, the weighted sum can be calculated by varying the weights in units of at least one of a sample, a group of samples, a line, and a block.
[0434] For example, the weights can vary according to at least one of a size of the current block, a shape of the current block, and a position of the sample.
[0435] For example, the weighted sum can be calculated according to conditions preset in the encoder and the decoder.
[0436] For example, the weights are adaptively determined based on at least one of encoding parameters such as a size of a block, a shape of a block, and an intra prediction mode of at least one of the current block and the neighboring blocks.
[0437] For example, whether to calculate the weighted sum is adaptively determined based on at least one of encoding parameters such as a size of a block, a shape of a block, and an intra prediction mode of at least one of the current block and the neighboring blocks.
[0438] For example, when an operation range in which a sum of the one-dimensional Laplacian operations is calculated is greater than a size of a block classification unit, at least one of the weights applied to the samples within the block classification unit can be greater than at least one of the weights applied to the samples outside the block classification unit.
[0439] Alternatively, for example, when an operation range in which a sum of the one-dimensional Laplacian operations is calculated is equal to a size of a block classification unit, the weights applied to the samples within the block classification unit are all the same.
[0440] Information on the weights and / or whether the weighted sum calculation is performed can be entropy-encoded in the encoder and then signaled to the decoder.
[0441] According to an embodiment of the present invention, the sum of gradient values g along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction is calculated. v g h g d1 and g d2 In each of the steps, when one or more unavailable samples exist around the current sample, padding is performed on the unavailable sample, and the padded sample can be used to compute the gradient value. Padding refers to the method of copying the sample value of adjacent available sample to the unavailable sample. Optionally, sample values or statistics obtained based on the values of available sample adjacent to the unavailable sample can be used. Padding can be performed repeatedly for P columns and R rows. Here, P and R are both positive integers.
[0442] Here, an unavailable sample point refers to a sample point located outside the boundaries of a CTU, CTB, strip, parallel block, parallel block group, or screen. Optionally, an unavailable sample point may refer to a sample point belonging to at least one of a CTU, CTB, strip, parallel block, parallel block group, and screen, wherein at least one of the CTU, CTB, strip, parallel block, parallel block group, and screen is different from at least one of the CTU, CTB, strip, parallel block, parallel block group, and screen to which the current sample point belongs.
[0443] According to an embodiment of the present invention, the sum of gradient values g along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction is calculated respectively. v g h g d1 and g d2 When at least one of them is selected, predetermined sample points may not be used.
[0444] For example, in calculating the sum of gradient values g along the vertical direction, horizontal direction, first diagonal direction, and second diagonal direction. v g h g d1 and g d2 When at least one of them is true, filling samples may not be used.
[0445] Alternatively, for example, in calculating the sum of gradient values g along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. v g h g d1 and g d2 In each of the cases, if there are one or more unavailable samples around the current sample, the unavailable sample may not be used to calculate the gradient value.
[0446] Further, alternatively, for example, in calculating the sum of gradient values g along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. v gh , g d1 and g d2 At least one of the above, when a sample around a current sample is located outside of a CTU or a CTB, a neighboring sample adjacent to the current sample can not be used.
[0447] According to an embodiment of the present application, in calculating at least one of the one-dimensional Laplacian values, when there is one or more unavailable samples around a current sample, padding is performed so that a sample value of an available sample adjacent to the unavailable sample is copied to the unavailable sample, and the one-dimensional Laplacian operation is performed using the padded sample.
[0448] According to an embodiment of the present application, in the one-dimensional Laplacian calculation, a predetermined sample can not be used.
[0449] For example, in the one-dimensional Laplacian calculation, a padded sample can not be used.
[0450] Optionally, for example, in calculating at least one of the one-dimensional Laplacian values, when there is one or more unavailable samples around a current sample, the one or more unavailable samples can not be used for the one-dimensional Laplacian operation.
[0451] Further optionally, for example, in calculating at least one of the one-dimensional Laplacian values, when a sample around a current sample is located outside of a CTU or a CTB, a neighboring sample can not be used for the one-dimensional Laplacian operation.
[0452] According to an embodiment of the present application, in calculating each of the sums g v , g h , g d1 and g d2 along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction, or in calculating at least one of the one-dimensional Laplacian values, at least one of samples that have undergone at least one of deblocking filtering, adaptive sample offset (SAO), and adaptive in-loop filtering can be used.
[0453] According to an embodiment of the present application, in calculating at least one of the sums g v , g h , g d1 and g d2 along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction, or in calculating at least one of the one-dimensional Laplacian values, when a sample around a current block is disposed outside of a CTU or a CTB, at least one of deblocking filtering, adaptive sample offset (SAO), and adaptive in-loop filtering can be applied to the corresponding sample.
[0454] Optionally, when at least one of the sums g v , g h , g d1 , and g d2 is calculated or when at least one of the one-dimensional Laplacian operation values is calculated, at least one of deblocking filtering, adaptive sample offset (SAO), and adaptive in-loop filtering can not be applied to the corresponding samples when the samples around the current block are arranged outside of the CTU or the CTB.
[0455] According to embodiments of the present invention, when there are unavailable samples arranged within the operation range for the one-dimensional Laplacian sum operation and arranged outside of the CTU or the CTB, the unavailable samples can be used for the calculation of the one-dimensional Laplacian operation without applying at least one of deblocking filtering, adaptive sample offset, and adaptive in-loop filtering.
[0456] According to embodiments of the present invention, when there are unavailable samples within the block classification unit or outside of the CTU or the CTB, the one-dimensional Laplacian operation can be performed without applying at least one of deblocking filtering, adaptive sample offset, and adaptive in-loop filtering to the unavailable samples.
[0457] On the other hand, when the gradient values are calculated based on sub-sampling, the one-dimensional Laplacian operation is performed not on all samples within the operation range for the one-dimensional Laplacian operation but on sub-samples within the operation range. Therefore, the number of operations (such as multiplication, shift operation, addition, and absolute value operation) required for block classification can be reduced. In addition, the memory access bandwidth required using the reconstructed samples can also be reduced. Therefore, the complexity of the encoder and the decoder can also be reduced. In particular, because the time required for block classification can be reduced, performing the one-dimensional Laplacian operation on the sub-sampled samples is beneficial in terms of hardware complexity of the encoder and the decoder.
[0458] In addition, when the operation range for calculating the sum of the one-dimensional Laplacian operation is equal to or smaller than the size of the block classification unit, the number of additions required for block classification can be reduced. In addition, the memory access bandwidth required using the reconstructed samples can also be reduced. Therefore, the complexity of the encoder and the decoder can also be reduced.
[0459] On the other hand, in the gradient value calculation method based on sub-sampling, the sum g v , g h , g d1 , and gd2 at least one of g
[0460] In addition, in the gradient value calculation method based on sub-sampling, regardless of gradient values for the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction, at least one of g v , g h , g d1 , and g d2 is calculated by using at least one of the sample position, the number of samples, and the direction of the sample position, which perform the one-dimensional Laplacian operation.
[0461] In addition, by using an arbitrary combination of the one or more gradient values calculated above, the one-dimensional Laplacian operation can be performed for the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction, and at least one of g v , g h , g d1 , and g d2 can be calculated.
[0462] According to an embodiment of the present invention, two or more values of g v , g h , g d1 , and g d2 are compared with each other along the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction.
[0463] For example, after the sum of the gradient values is calculated, g v is compared with g h , and the maximum and minimum values of the sum of the gradient values for the vertical direction and the sum of the gradient values for the horizontal direction are derived according to Equation 16.
[0464] [Equation 16]
[0465]
[0466] In this case, in order to compare g v with g h , the values of the sums are compared according to Equation 17.
[0467] [Equation 17]
[0468]
[0469] Optionally, for example, the sum g d1 of gradient values for the first diagonal direction is compared to the sum g d2 of gradient values for the second diagonal direction, and the maximum of the sum of gradient values for the first diagonal direction and the sum of gradient values for the second diagonal direction is derived according to equation 18 and the minimum of the sum of gradient values for the first diagonal direction and the sum of gradient values for the second diagonal direction is derived according to equation 18
[0470] [Equation 18]
[0471]
[0472] In this case, in order to compare the sum g d1 of gradient values for the first diagonal direction to the sum g d2 of gradient values for the second diagonal direction, the values of the sums are compared according to equation 19.
[0473] [Equation 19]
[0474]
[0475] According to one embodiment of the present application, in order to calculate the directionality value D, the maximum is compared to the minimum using two thresholds t1 and t2 as follows.
[0476] The directionality value D is a positive integer or zero. For example, the directionality value D can be a value in the range from 0 to 4. For example, the directionality value D can be a value in the range from 0 to 2.
[0477] In addition, the directionality value D can be determined according to the characteristics of the region. For example, the directionality values Ds 0 to Ds 4 are represented as follows: 0 indicates a texture region; 1 indicates strong horizontal / vertical directionality; 2 indicates weak horizontal / vertical directionality; 3 indicates strong first / second diagonal directionality; and 4 indicates weak first / second diagonal directionality. The directionality value D is determined by the steps described below.
[0478] Step 1 : When and are satisfied, the value D is set to 0
[0479] Step 2: When is satisfied, step 3 is entered, and when it is not satisfied, step 4 is entered
[0480] Step 3: When is satisfied, the value D is set to 2, and when it is not satisfied, the value D is set to 1
[0481] Step 4: When is satisfied, the value D is set to 4, and when it is not satisfied, the value D is set to 3
[0482] wherein the threshold values t1 and t2 are positive integers, and t1 and t2 can be the same value or different values. For example, t1 and t2 are 2 and 9, respectively. In another example, t1 and t2 are both 1. In another example, t1 and t2 are 1 and 9, respectively.
[0483] When the block classification is performed based on the 2x2 size block, the activity value A can be expressed as expression 20.
[0484] [Equation 20]
[0485]
[0486] For example, k and l are the same range. That is, the horizontal length and the vertical length of the operation range for calculating the sum of the one-dimensional Laplacian operation are equal.
[0487] Alternatively, for example, k and l are different ranges from each other. That is, the horizontal length and the vertical length of the operation range for calculating the sum of the one-dimensional Laplacian operation are different.
[0488] Further alternatively, for example, k is a range from i-2 to i+3, and l is a range from j-2 to j+3. In this case, the operation range for calculating the sum of the one-dimensional Laplacian operation is 6x6 size.
[0489] Further alternatively, for example, k is a range from i-1 to i+2, and l is a range from j-1 to j+2. In this case, the operation range for calculating the sum of the one-dimensional Laplacian operation is 4x4 size.
[0490] Further alternatively, for example, k is a range from i to i+1, and l is a range from j to j+1. In this case, the operation range for calculating the sum of the one-dimensional Laplacian operation is 2x2 size. In this case, the operation range for calculating the sum of the one-dimensional Laplacian operation can be equal to the size of the block classification unit.
[0491] For example, the operation range for calculating the sum of the result of the one-dimensional Laplacian operation can have a two-dimensional geometric shape selected from a rhombus, a rectangle, a square, a trapezoid, a diagonal line, a snowflake, a numeral symbol, a shamrock, a cross, a triangle, a pentagon, a hexagon, a decagon, and a dodecagon.
[0492] In addition, when the block classification is performed based on the 4x4 size block, the activity value A can be expressed as expression 21.
[0493] [Equation 21]
[0494]
[0495] For example, k and l are the same range. That is, the horizontal and vertical lengths of the range for calculating the sum of a one-dimensional Laplace operation are equal.
[0496] Optionally, for example, k and l are different ranges from each other. That is, the horizontal length and vertical length of the range for calculating the sum of a one-dimensional Laplace operation are different.
[0497] Further optionally, for example, k is a range from i-2 to i+5, and l is a range from j-2 to j+5. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation is 8×8.
[0498] Further alternatively, for example, k is the range from i to i+3, and l is the range from j to j+3. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation is 4×4. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation can be equal to the size of the block classification unit.
[0499] For example, the range of operations for calculating the sum of the results of a one-dimensional Laplace operation can have two-dimensional geometric shapes selected from rhombuses, rectangles, squares, trapezoids, diagonals, snowflakes, numeral symbols, cloverleaf shapes, crosses, triangles, pentagons, hexagons, decagons, and dodecagons.
[0500] Furthermore, when performing block classification based on 2×2 size blocks, the activity value A can be expressed as expression 22. Here, at least one of the one-dimensional Laplace operation values for the first diagonal direction and the second diagonal direction can be additionally used in the calculation of the activity value A.
[0501] [Equation 22]
[0502]
[0503] For example, k and l are the same range. That is, the horizontal and vertical lengths of the range for calculating the sum of a one-dimensional Laplace operation are equal.
[0504] Optionally, for example, k and l are different ranges from each other. That is, the horizontal length and vertical length of the range for calculating the sum of a one-dimensional Laplace operation are different.
[0505] Further optionally, for example, k is a range from i-2 to i+3, and l is a range from j-2 to j+3. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation is 6×6.
[0506] Further, alternatively, for example, k is a range from i-1 to i+2, and l is a range from j-1 to j+2. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation is 4×4.
[0507] Further optionally, for example, k is the range from i to i+1, and l is the range from j to j+1. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation is 2×2. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation can be equal to the size of the block classification unit.
[0508] For example, the range of operations for calculating the sum of the results of a one-dimensional Laplace operation can have two-dimensional geometric shapes selected from rhombuses, rectangles, squares, trapezoids, diagonals, snowflakes, numeral symbols, cloverleaf shapes, crosses, triangles, pentagons, hexagons, decagons, and dodecagons.
[0509] Additionally, when performing block classification based on 4×4 size blocks, the activity value A can be expressed as expression 23. Here, at least one of the one-dimensional Laplace operation values for the first diagonal direction and the second diagonal direction can be additionally used to calculate the activity value A.
[0510] [Equation 23]
[0511]
[0512] For example, k and l are the same range. That is, the horizontal and vertical lengths of the range for calculating the sum of a one-dimensional Laplace operation are equal.
[0513] Optionally, for example, k and l are different ranges from each other. That is, the horizontal length and vertical length of the range for calculating the sum of a one-dimensional Laplace operation are different.
[0514] Further optionally, for example, k is a range from i-2 to i+5, and l is a range from j-2 to j+5. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation is 8×8.
[0515] Further alternatively, for example, k is the range from i to i+3, and l is the range from j to j+3. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation is 4×4. In this case, the operational range for calculating the sum of a one-dimensional Laplace operation can be equal to the size of the block classification unit.
[0516] For example, the range of operations for calculating the sum of the results of a one-dimensional Laplace operation can have two-dimensional geometric shapes selected from rhombuses, rectangles, squares, trapezoids, diagonals, snowflakes, numeral symbols, cloverleaf shapes, crosses, triangles, pentagons, hexagons, decagons, and dodecagons.
[0517] On the other hand, the activity value A can be quantized to generate a quantized activity value A in the range from I to J. q Here, I and J are both positive integers or zero. For example, I and J are 0 and 4 respectively.
[0518] A predetermined method can be used to determine the quantitative activity value A. q .
[0519] For example, the quantified activity value A can be determined using Equation 24. q In this case, the quantified activity value Aq can be included in the range from a specific minimum value X to a specific maximum value Y.
[0520] [Equation 24]
[0521]
[0522] In Equation 24, the quantified activity value A is calculated by multiplying the activity value A by a specific constant W and then performing a right shift operation R on the product of A and W. q In this case, X, Y, W, and R are all positive integers or zero. For example, W is 24 and R is 13. Alternatively, for example, W is 64 and R is 3+N (bits). For example, N is a positive integer, specifically 8 or 10. In another example, W is 32 and R is 3+N (bits). Alternatively, for example, N is a positive integer, specifically 8 or 10.
[0523] Further, alternatively, for example, a lookup table (LUT) can be used to calculate the quantified activity value A. q And set the activity value A and the quantified activity value A. q The mapping relationship between them. That is, performing operations on the activity value A and using a lookup table to calculate the quantified activity value A. q In this case, the operation may include at least one of multiplication, division, right shift, left shift, addition, and subtraction.
[0524] On the other hand, in the case of chroma components, filtering is performed on each chroma component using K filters, without performing block classification processing. Here, K is a positive integer or zero. For example, K is 1. Furthermore, in the case of chroma components, block classification can be omitted, and filtering can be performed using the block classification index derived from the luminance component at the corresponding position of the chroma component. Additionally, in the case of chroma components, filter information for the chroma components can be transmitted without signal transmission, and fixed-type filters can be used.
[0525] Figure 32 This is a diagram illustrating various computational methods that can be used to replace one-dimensional Laplace operations according to embodiments of the present invention.
[0526] According to embodiments of the present invention, it can be used Figure 32 At least one of the calculation methods shown can be used to replace the one-dimensional Laplace operation. (Refer to...) Figure 32 The computational methods include two-dimensional Laplacian, two-dimensional Sobel, two-dimensional edge extraction, and two-dimensional Laplacian of Gaussian (LoG) operations. Here, the LoG operation represents applying a combination of a Gaussian filter and a Laplacian filter to the reconstructed samples. In addition to these methods, at least one of a one-dimensional edge extraction filter and a two-dimensional edge extraction filter can be used instead of the one-dimensional Laplacian operation. Optionally, the Difference of Gaussian (DoG) operation can be used. Here, the DoG operation represents applying a combination of Gaussian filters with different intrinsic parameters to the reconstructed samples.
[0527] Additionally, to calculate the directionality value D or the mobility value A, an N×M LoG operation can be used. Here, M and L are both positive integers. For example, using... Figure 32 The 5×5 two-dimensional LoG shown in (i) and Figure 32 At least one of the 9×9 two-dimensional LoG operations shown in (j). Alternatively, for example, a one-dimensional LoG operation can be used instead of a two-dimensional LoG operation.
[0528] According to embodiments of the invention, each 2×2-sized block of brightness can be classified based on directionality and two-dimensional Laplacian activity. For example, horizontal / vertical gradient characteristics can be obtained using a Sobel filter. The directionality value D can be obtained using equations 25 to 26.
[0529] A representative vector can be computed such that the gradient vector within a predetermined window size (e.g., a 6×6 block) satisfies the conditions of Equation 25. The direction and deformation can be identified based on θ.
[0530] [Equation 25]
[0531]
[0532] The similarity between the representative vector and each gradient vector within the window can be calculated using the inner product as shown in Equation 26.
[0533] [Equation 26]
[0534]
[0535] The directional value D can be determined using the S value calculated from Equation 26.
[0536] Step 1: When S>th1 is satisfied, set the value of D to 0.
[0537] Step 2: Set the value of D to 2 when θ∈(D0 or D1) and S>th2 are satisfied, and set the value of D to 1 when they are not satisfied.
[0538] Step 3: When θ∈(V or H) and S<th2 are satisfied, set the value of D to 4; when not satisfied, set the value of D to 3.
[0539] Here, the total number of block category indexes can be 25.
[0540] According to an embodiment of the present invention, the block classification of the reconstructed sample point s′(i,j) can be represented by Equation 27.
[0541] [Equation 27]
[0542] For k = 0, ..., K-1
[0543] In Equation 27, I represents the set of sample locations for all reconstructed sample points s′(i,j). D is the classifier that assigns classification indices k∈{0,…,K-1} to sample location (i,j). Furthermore, This is the set of all samples to which the classification index is assigned by the classifier D. Classification supports four different classifiers, and each classifier can provide K = 25 or 27 classes. The classifier used in the decoder can be specified by the syntax element `classification_idx`, which is transmitted as a signal at the stripe level. Given a class with classification index k ∈ {0, ..., K-1}... Perform the following steps.
[0544] When classification_idx = 0, use the block classifier D based on directionality and activity. G The classifier can provide K=25 classes.
[0545] When classification_idx = 1, the feature classifier D based on sample points... S It is used as a classifier. D S (i,j) uses the quantized sample value of each of the sample points s′(i,j) according to Equation 28.
[0546] [Equation 28]
[0547]
[0548] Where B is the sample bit depth, the classification number K is set to 27 (K=27), and the operator... Specify the operation to round to the nearest integer.
[0549] When classification_idx = 2, the sorted sample-based feature classifier can be used. Used as a classifier. Equation 30 represents this. r8(i,j) is a classifier that compares s′(i,j) with its eight neighboring sample points and arranges the sample points in order of their values.
[0550] [Equation 29]
[0551]
[0552] The value of classifier r8(i,j) ranges from 0 to 8. When sample s′(i,j) is the largest sample within a 3×3 block centered at (i,j), the value of r8(i,j) is zero. When s′(i,j) is the second largest sample, the value of r8(i,j) is 1.
[0553] [Equation 30]
[0554]
[0555] In Equation 30, T1 and T2 are predefined thresholds. That is, the dynamic range of the sample points is divided into three bands, and the ranking of local samples in each band is used as an additional criterion. The ranking-based sample-based feature classifier provides 27 classes (K=27).
[0556] When classification_idx = 3, a classifier based on ranking and region variation is used. It can be represented by Equation 31.
[0557] [Equation 31]
[0558]
[0559] In Equation 31, T3 or T4 is a predefined threshold. The local change v(i,j) at each sample location (i,j) can be represented by Equation 32.
[0560] [Equation 32]
[0561] v(i,j)=4*s′(i,j)-(s′(i-1,j)+s′(i+1,j)+s′(i,j+1)+s′(i,j-1))
[0562] In addition to each sample point being first classified into one of three classes based on the local variable |v(i,j)|, Is with The same classifier is used. Next, within each class, the ranking of nearby local samples can be used as an additional criterion to provide the 27 classes.
[0563] According to embodiments of the invention, at the strip level, a filter bank comprising up to 16 filters using three pixel classification methods (such as intensity classifier, histogram classifier, and directional activity classifier) is used for the current strip. At the CTU level, based on control flags in the strip header transmitted with signals, three modes are used for each CTU, including a new filter mode, a spatial filter mode, and a strip filter mode.
[0564] Here, the intensity classifier is similar to the band offset in SAO. The intensity range of the sample points is divided into 32 groups, and the group index for each sample point is determined based on the intensity of the sample point to be processed.
[0565] In the case of a similarity classifier, neighboring samples in a 5×5 diamond filter are compared with the target sample, which is the sample to be filtered. The group index of the sample to be filtered can be initialized to 0. When the difference between a neighboring sample and the target sample is greater than a predefined threshold, the group index is incremented by 1. Furthermore, when the difference between a neighboring sample and the target sample is twice the predefined threshold, the group index is incremented by another 1. In this case, the similarity classifier has 25 groups.
[0566] Furthermore, in the case of the Rot BA classifier, the computational range for calculating the sum of a one-dimensional Laplacian operation over a 2×2 block is reduced from a 6×6 size to a 4×4 size. This classifier has a maximum of 25 groups. Multiple classifiers can have a maximum of 25 or 32 groups. However, the number of filters in a strip filter bank is limited to a maximum of 16 groups. That is, the encoder will continuously combine and merge, ensuring that the number of merged groups remains 16 or fewer.
[0567] According to embodiments of the present invention, when determining a block classification index, the block classification index is determined based on at least one of the coding parameters of the current block and neighboring blocks. The block classification index varies according to at least one of the coding parameters. In this case, the coding parameters include at least one of the following: prediction mode (i.e., whether the prediction is intra-frame prediction or inter-frame prediction), inter-frame prediction mode, intra-frame prediction mode, intra-frame prediction indicator, motion vector, reference frame index, quantization parameter, block size of the current block, block shape of the current block, size of the block classification unit, and coding block flag / style.
[0568] In one example, block classification is determined based on quantization parameters. For instance, when the quantization parameter is less than a threshold T, J block classification indices are used. When the quantization parameter is greater than a threshold R, H block classification indices are used. For other cases, G block classification indices are used. Here, T, R, J, H, and G are positive integers or zero. Furthermore, J is greater than or equal to H. The larger the quantization parameter value, the fewer block classification indices are used.
[0569] In another example, the number of block classifications is determined based on the size of the current block. For instance, J block classification indices are used when the current block size is less than a threshold T. H block classification indices are used when the current block size is greater than a threshold R. For other cases, G block classification indices are used. Here, T, R, J, H, and G are positive integers or zero. Furthermore, J is greater than or equal to H. The larger the block size, the fewer block classification indices are used.
[0570] In another example, the number of block classifications is determined based on the size of the block classification unit. For instance, when the size of the block classification unit is less than a threshold T, J block classification indices are used. When the size of the block classification unit is greater than a threshold R, H block classification indices are used. For other cases, G block classification indices are used. Here, T, R, J, H, and G are positive integers or zero. Additionally, J is greater than or equal to H. The larger the size of the block classification unit, the fewer block classification indices are used.
[0571] According to an embodiment of the present invention, at least one of the sum of gradient values at the same location in the previous frame, the sum of gradient values of neighboring blocks around the current block, and the sum of gradient values of neighboring block classification units around the current block classification unit is determined as at least one of the sum of gradient values of the current block and the sum of gradient values of the current block classification unit. Here, the co-located sample point in the previous frame is the spatial location or neighboring location of the reconstructed sample point in the current frame within the previous frame.
[0572] For example, the sum of the gradient values g in the vertical and horizontal directions for the current block cell. v and g hWhen the difference between at least one of the gradient values in the vertical and horizontal directions of the current block classification unit and at least one of the gradient values in the neighboring block classification units around the current block classification unit is equal to or less than a threshold E, the sum of the gradient values in the first diagonal direction and the second diagonal direction of the neighboring block classification units for the current block classification unit will be g. d1 and g d2 At least one of the values is determined to be at least one of the gradient values of the current block unit. Here, the threshold E is a positive integer or zero.
[0573] In another example, the sum of the gradient values g in the vertical and horizontal directions for the current block cell. v and g h If the difference between the sum of gradient values and the sum of gradient values in the vertical and horizontal directions of neighboring block classification units surrounding the current block classification unit is equal to or less than a threshold E, at least one of the sums of gradient values of neighboring block classification units of the current block classification unit is determined as at least one of the sums of gradient values of the current block unit. Here, the threshold E is a positive integer or zero.
[0574] In another example, when the difference between at least one statistical value of a reconstructed sample within the current block cell and at least one statistical value of a reconstructed sample within neighboring block classification cells surrounding the current block cell is equal to or less than a threshold E, at least one of the sums of gradient values of neighboring block classification cells surrounding the current block cell is determined as at least one of the sums of gradient values of the current block cell. Here, the threshold E is a positive integer or zero. The threshold E is derived from the spatial and / or temporal neighboring blocks of the current block. Furthermore, the threshold E is a value predefined in the encoder and decoder.
[0575] According to an embodiment of the present invention, at least one of the block classification index of the co-position sample point in the previous frame, the block classification index of the neighboring block of the current block, and the block classification index of the neighboring block classification unit of the current block classification unit is determined as at least one of the block classification index of the current block and the block classification index of the current block classification unit.
[0576] For example, the sum of the gradient values g in the vertical and horizontal directions for the current block cell. v and g h If the difference between at least one of the gradient values and at least one of the sums of the vertical and horizontal gradient values of the neighboring block classification units surrounding the current block classification unit is equal to or less than a threshold E, then the block classification index of the neighboring block classification units surrounding the current block classification unit is determined as the block classification index of the current block unit. Here, the threshold E is a positive integer or zero.
[0577] Alternatively, for example, the sum of the gradient values g in the vertical and horizontal directions for the current block cell. v and g hWhen the difference between the sum of the values of the current block classification unit and the sum of the sums of the gradient values of the neighboring block classification units in the vertical and horizontal directions is equal to or less than the threshold E, the block classification index of the neighboring block classification units surrounding the current block classification unit is determined as the block classification unit of the current block classification unit. Here, the threshold E is a positive integer or zero.
[0578] Further optionally, for example, if the difference between at least one statistical value of a reconstructed sample point within the current block cell and at least one statistical value of a reconstructed sample point within a neighboring block cell surrounding the current block cell is equal to or less than a threshold E, the block classification index of the neighboring block cell surrounding the current block cell is determined as the block classification index of the current block cell. Here, the threshold E is a positive integer or zero.
[0579] Further, alternatively, for example, at least one of the combinations of the block classification index determination methods described above can be used to determine the block classification index.
[0580] The following sections will describe the filtering execution sub-steps.
[0581] According to an exemplary embodiment of the present invention, a filter corresponding to a determined block classification index is used to perform filtering on samples or blocks in the reconstructed / decoded image. When performing filtering, one of L filters is selected. L is a positive integer or zero.
[0582] For example, one of L filters is selected based on each block classification unit, and filtering is performed on the reconstructed / decoded image based on each reconstructed / decoded sample.
[0583] Alternatively, for example, one of L filters can be selected based on each block classification unit, and filtering can be performed on the reconstructed / decoded image based on each block classification unit.
[0584] Further, alternatively, for example, one of the L filters is selected based on each block classification unit, and filtering is performed on the reconstructed / decoded image based on each CU.
[0585] Further, alternatively, for example, one of the L filters is selected based on each block classification unit, and filtering is performed on the reconstructed / decoded image based on each block.
[0586] Further, alternatively, for example, U filters out of L filters are selected based on each block classification unit, and filtering is performed on the reconstructed / decoded image based on each reconstructed / decoded sample. Here, U is a positive integer.
[0587] Further, alternatively, for example, U filters out of L filters are selected based on each block classification unit, and filtering is performed on the reconstructed / decoded image based on each block classification unit. Here, U is a positive integer.
[0588] Further, alternatively, for example, U filters out of L filters are selected based on each block classification unit, and filtering is performed on the reconstructed / decoded image based on each CU. Here, U is a positive integer.
[0589] Further, alternatively, for example, U filters out of L filters are selected based on each block classification unit, and filtering is performed on the reconstructed / decoded image for each block pair. Here, U is a positive integer.
[0590] Here, L filters are called a filter bank.
[0591] According to an embodiment of the invention, the L filters differ from each other in at least one of the following aspects: filter coefficients, number of filter taps (i.e., filter length), filter shape, and filter type.
[0592] For example, in units of block, CU, PU, TU, CTU, strip, parallel block, parallel block group, frame, and sequence, L filters are common in at least one aspect of filter coefficients, number of filter taps (filter length), filter shape, and filter type.
[0593] Optionally, for example, in units of CU, PU, TU, CTU, strip, parallel block, parallel block group, frame, and sequence, L filters are common in at least one aspect of filter coefficients, number of filter taps (filter length), filter shape, and filter type.
[0594] Filtering can be performed using the same or different filters at the unit level of CU, PU, TU, CTU, strip, parallel block, parallel block group, frame, and sequence.
[0595] The decision to perform filtering is based on whether filtering is performed on a per-sample, per-block, per-block, per-unit, per-unit, per-unit, per-unit, per-unit, per-unit, per-frame, or per-sequence basis. This filtering execution information refers to the information transmitted from the encoder to the decoder using signals, per-sample, per-block, per-unit, per-block, per-unit, per-unit, per-unit, per-frame, and per-sequence basis.
[0596] According to embodiments of the present invention, N filters with different numbers of filter taps and the same filter shape (i.e., square or diamond-shaped filter) are used. Here, N is a positive integer. For example, in Figure 33 The diagram shows a diamond filter with 5×5, 7×7, or 9×9 filter taps.
[0597] Figure 33 This is a diagram illustrating a diamond-shaped filter according to an embodiment of the present invention.
[0598] Reference Figure 33 To transmit filter information from the encoder to the decoder using signals from three diamond-shaped filters with 5×5, 7×7, or 9×9 filter taps, the filter index is entropy-encoded / decoded based on each frame / parallel block / parallel block group / strip / sequence. In other words, the filter index is entropy-encoded / decoded in the bitstream using sequence parameter sets, frame parameter sets, strip headers, strip data, parallel block headers, parallel block group headers, and so on.
[0599] According to an embodiment of the present invention, when the number of filter taps in the encoder / decoder is fixed at 1, the encoder / decoder performs filtering using the filter index without entropy encoding / decoding the filter index. Here, a diamond filter with 7×7 filter taps is used for the luminance component, and a diamond filter with 5×5 filter taps is used for the chrominance component.
[0600] According to an embodiment of the present invention, at least one of the three diamond filters is used to filter at least one reconstructed / decoded sample of at least one of the luminance component and the chrominance component.
[0601] For example, in Figure 33 At least one of the three diamond-shaped filters shown is used to filter the luminance samples for reconstruction / decoding.
[0602] Optionally, for example, in Figure 33 The 5×5 diamond-shaped filter shown is used to filter the reconstructed / decoded chroma samples.
[0603] Further alternatively, for example, the filter used to filter the luminance sample can be used to filter the reconstructed / decoded chrominance sample corresponding to the luminance sample.
[0604] In addition, Figure 33 The numbers in each filter shape shown represent filter coefficient indices, and these indices are symmetrical about the filter center. That is, in Figure 33 The filter shown is a point-symmetric filter.
[0605] On the other hand, Figure 33 In the case of the 9×9 diamond filter shown in (a), entropy encoding / decoding is performed on a total of 21 filter coefficients. Figure 33 In the case of the 7×7 diamond filter illustrated in (b), entropy encoding / decoding is performed on a total of 13 filter coefficients, and... Figure 33 In the case of the 5×5 diamond filter illustrated in (c), a total of 7 filter coefficients are entropy encoded / decoded. That is, a maximum of 21 filter coefficients need to be entropy encoded / decoded.
[0606] In addition, regarding Figure 33 The 9×9 diamond filter shown in (a) requires a total of 21 multiplications per sample. Figure 33 The 7×7 diamond filter shown in (b) requires a total of 13 multiplications per sample. Figure 33 The 5×5 diamond filter shown in (c) requires a total of 7 multiplications per sample. That is, filtering is performed using a maximum of 21 multiplications per sample.
[0607] In addition, such as Figure 33 As shown in (a), since the 9×9 diamond filter has a size of 9×9, the hardware implementation requires four line buffers, which are half the length of the vertical filter. That is, a maximum of four line buffers are required.
[0608] According to embodiments of the invention, the filter has the same filter length representing 5×5 filter taps, but may have different filter shapes selected from rhombuses, rectangles, squares, trapezoids, diagonals, snowflakes, digit symbols, cloverleaf shapes, crosses, triangles, pentagons, hexagons, octagons, decagons, and dodecagons. For example, in Figure 33 The image shows square, octagonal, snowflake, and diamond filters with 5×5 filter taps.
[0609] The number of filter taps is not limited to 5×5. A filter with H×V filter taps selected from 3×3, 4×4, 5×5, 6×6, 7×7, 8×8, 9×9, 5×3, 7×3, 9×3, 7×5, 9×5, 9×7, and 11×7 can be used. Here, H and V are positive integers and can be the same or different values. Additionally, at least one of H and V is a predefined value in the encoder / decoder and a value transmitted from the encoder to the decoder via signal transmission. Furthermore, one of H and V is used to define the other of H and V. Moreover, the final value of H or V can be defined using the values of H and V.
[0610] On the other hand, in order to Figure 34 The information of which filter will be used, as shown in the diagram, is transmitted from the encoder to the decoder via a signal. This can be achieved by encoding / decoding the filter index entropy based on each frame / parallel block / parallel block group / strip / sequence. In other words, the filter index entropy is encoded / decoded into the sequence parameter set, frame parameter set, strip header, strip data, parallel block header, and parallel block group header within the bitstream.
[0611] On the other hand, will Figure 34At least one of the square, octagonal, snowflake, and rhombus filters shown is used to filter at least one reconstructed / decoded sample of at least one of the luminance and chrominance components.
[0612] On the other hand, Figure 34 The numbers in each filter shape shown represent filter coefficient indices, and these indices are symmetrical about the filter center. That is, in Figure 34 The filter shown is a point-symmetric filter.
[0613] According to embodiments of the present invention, when filtering the reconstructed image based on each sample point, the appropriate filter shape for each image, strip, parallel block, or group of parallel blocks can be determined in terms of rate-distortion optimization within the encoder. Furthermore, filtering is performed using the determined filter shape. Figure 34 As shown, because the degree of improvement in coding efficiency and the amount of filter information (the number of filter coefficients) vary depending on the filter shape, it is necessary to determine the optimal filter shape for each frame, strip, parallel block, or group of parallel blocks. In other words, the optimal filter shape must be determined based on factors such as video resolution, video characteristics, and bit rate. Figure 34 The optimal filter shape among the filter shapes shown.
[0614] According to embodiments of the present invention, and with the use of, Figure 34 Compared to the filter shown, using such Figure 33 The filter shown has the advantage of reducing the computational complexity of the encoder / decoder.
[0615] For example, in Figure 34 In the case of the 5×5 square filter shown in (a), entropy encoding / decoding is performed on a total of 13 filter coefficients. Figure 34 In the case of the 5×5 octagonal filter shown in (b), entropy encoding / decoding is performed on a total of 11 filter coefficients. Figure 34 In the case of the 5×5 snowflake filter shown in (c), entropy encoding / decoding is performed on a total of 9 filter coefficients, and... Figure 34 In the case of the 5×5 diamond-shaped filter shown in (c), a total of 7 filter coefficients are entropy encoded / decoded. That is, the number of filter coefficients to be entropy encoded / decoded varies depending on the filter shape. Here, in Figure 34 The maximum number of filter coefficients in the example (i.e., 13) is less than in Figure 34 The maximum number of filter coefficients in the example is 21. Therefore, when using Figure 33When using filters in the example, the number of filter coefficients that will be entropy encoded / decoded is reduced. Therefore, in this case, the computational complexity of the encoder / decoder can be reduced.
[0616] Optionally, for example, for Figure 34 The 5×5 square filter shown in (a) requires a total of 13 multiplications per sample. Figure 34 The 5×5 octagonal filter shown in (b) requires a total of 11 multiplications per sample. Figure 34 The 5×5 snowflake filter shown in (c) requires a total of 9 multiplications per sample, and for Figure 34 The 5×5 rhombus filter shown in (d) requires a total of 7 multiplications per sample. Figure 34 The maximum number of filter coefficients in the example (i.e., 13) is less than in Figure 33 The maximum number of filter coefficients in the example is 21. Therefore, when using Figure 35a When using filters in the example, the number of multiplications per sample is reduced. Therefore, in this case, the computational complexity of the encoder / decoder can be reduced.
[0617] Alternatively, for example, due to in Figure 35b All filters in the example are 5×5 in size, so the hardware implementation requires two line buffers, each half the length of a vertical filter. Here, using... Figure 35a The filter in the example requires fewer line buffers (i.e., two line buffers) than using a filter like... Figure 35b The example filter requires the number of line buffers (i.e., four line buffers). Therefore, when using Figure 35a When using filters as an example, the size of the line buffer, the hardware complexity of the encoder / decoder, memory capacity requirements, and memory access bandwidth can be reduced.
[0618] According to an embodiment of the present invention, as the filter used in the above-described filtering process, a filter having at least one shape selected from rhombus, rectangle, square, trapezoid, diagonal, snowflake, numeral symbol, four-leaf clover, cross, triangle, pentagon, hexagon, octagon, decagon, and dodecagon is used. For example, such as Figure 35b and / or Figure 35a As shown, the filter can have a shape selected from square, octagon, snowflake, rhombus, hexagon, rectangle, cross, number symbol, clover, and diagonal.
[0619] For example, using Figure 35b and / or Figure 35aThe filter bank is constructed by using at least one of the filters shown, which has a vertical length of 5, and then filtering is performed using the filter bank.
[0620] Alternatively, for example, using Figure 35b and Figure 35a The filter bank is constructed using at least one of the vertical filter lengths of 3 shown in the filter diagram, and then filtering is performed using the filter bank.
[0621] Further, alternatively, for example, using Figure 35b and / or Figure 35a The filter bank is constructed using at least one of the vertical filter lengths of 3 or 5 shown, and filtering is performed using the filter bank.
[0622] exist Figure 35b and Figure 35a The filter shown is designed with a vertical filter length of 3 or 5. However, the filter shape used in embodiments of the invention is not limited to this. The filter can be designed with any vertical filter length M. Here, M is a positive integer.
[0623] On the other hand, used in Figure 35b and / or Figure 33 The filters shown are used to prepare H filter banks, and information about which filter to use is transmitted from the encoder to the decoder via a signal. In this case, the filter index is entropy encoded / decoded based on each frame, parallel block, parallel block group, stripe, or sequence. Here, H is a positive integer. That is, the filter index is entropy encoded / decoded into the sequence parameter set, frame parameter set, stripe header, stripe data, parallel block header, and parallel block group header within the bitstream.
[0624] At least one of the following filters—rhombus, rectangle, square, trapezoid, diagonal, snowflake, number symbol, four-leaf clover, cross, triangle, pentagon, hexagon, octagon, and decagon—is used to filter the reconstructed / decoded samples of at least one of the luminance and chrominance components.
[0625] On the other hand, Figure 35a and / or Figure 35b The numbers in each filter shape shown represent filter coefficient indices, and these indices are symmetrical about the filter center. That is, Figure 35a and / or Figure 35b The filter shape shown is a point-symmetric filter.
[0626] According to embodiments of the present invention, and with the use of, Figure 33 Compared to the filters in the examples, using such Figure 35a and / orFigure 35b The filters in the example have the advantage of reducing the computational complexity of the encoder / decoder.
[0627] For example, when used in Figure 33 and / or Figure 35a When at least one of the filters shown is used, it is similar to the filter used in Figure 35b Compared to the case of one of the 9×9 diamond filters shown, the number of filter coefficients to be entropy encoded / decoded is reduced. Therefore, the computational complexity of the encoder / decoder can be reduced.
[0628] Alternatively, for example, when used in Figure 33 and / or Figure 36 When at least one of the filters shown is used, it is similar to the filter used in Figure 36 Compared to one of the 9×9 diamond filters shown, the number of multiplications required to filter the filter coefficients is reduced. Therefore, the computational complexity of the encoder / decoder can be reduced.
[0629] Further, alternatively, for example, when used in Figure 36 and / or Figure 36 When at least one of the filters shown is used, it is similar to the filter used in Figure 36 Compared to one of the 9×9 diamond filters shown, the number of lines in the line buffer required to filter the filter coefficients is reduced. Furthermore, hardware complexity, memory requirements, and memory access bandwidth are also reduced.
[0630] According to an embodiment of the present invention, it can be derived from Figure 36 At least one filter selected from the horizontal / vertical symmetric filters shown is used in place of the point symmetric filter for filtering. Alternatively, in addition to the point symmetric filter and the horizontal / vertical symmetric filter, a diagonal symmetric filter may also be used. Figure 36 In the diagram, the numbers in each filter shape represent the filter coefficient indices.
[0631] For example, using in Figure 33 The filter bank is constructed by using at least one of the vertical filters with a length of 5 shown, and then the filter bank is used to perform filtering.
[0632] Optionally, for example, using in Figure 36 At least one of the vertical filters with a length of 3 shown in the filter diagram constitutes a filter group, and then the filter group is used for filtering.
[0633] Further, alternatively, for example, using in Figure 36The filter bank is constructed using at least one of the vertical filter lengths of 3 or 5 shown, and the filter bank is used to perform filtering.
[0634] exist Figure 33 The filter shape shown is designed with a vertical filter length of 3 or 5. However, the filter shape used in embodiments of the invention is not limited to this. The filter can be designed with any vertical filter length M. Here, M is a positive integer.
[0635] In order to prepare for including Figure 36 The diagram shows a filter bank of H filters, and information about which filter in the filter bank will be used is transmitted from the encoder to the decoder via a signal. The filter indices are entropy-encoded / decoded based on each frame, parallel block, parallel block group, stripe, or sequence. Here, H is a positive integer. That is, the filter indices are entropy-encoded / decoded into sequence parameter sets, frame parameter sets, stripe headers, stripe data, parallel block headers, and parallel block group headers within the bitstream.
[0636] At least one of the following filters—rhombus, rectangle, square, trapezoid, diagonal, snowflake, number symbol, four-leaf clover, cross, triangle, pentagon, hexagon, octagon, and decagon—is used to filter reconstructed / decoded samples of at least one of the luminance and chrominance components.
[0637] According to embodiments of the present invention, and with the use of, Figure 33 Compared to the filter shown, using such Figure 36 The filter shown has the advantage of reducing the computational complexity of the encoder / decoder.
[0638] For example, when used in Figure 33 When at least one of the filters shown is used, it is similar to the filter used in Figure 37 Compared to the case of one of the 9×9 diamond filters shown, the number of filter coefficients to be entropy encoded / decoded is reduced. Therefore, the computational complexity of the encoder / decoder can be reduced.
[0639] Alternatively, for example, when used in Figure 37 When at least one of the filters shown is used, it is similar to the filter used in Figure 38 Compared to one of the 9×9 diamond filters shown, the number of multiplications required to filter the filter coefficients is reduced. Therefore, the computational complexity of the encoder / decoder can be reduced.
[0640] Further, alternatively, for example, when used in Figures 39 to 5 When at least one of the filters shown is used, it is similar to the filter used in Figures 39 to 5Compared to one of the 9×9 diamond filters shown, the number of lines in the line buffer required to filter the filter coefficients is reduced. Furthermore, hardware complexity, memory requirements, and memory access bandwidth are also reduced.
[0641] According to an embodiment of the present invention, before performing filtering based on each block classification unit, the sum of gradient values calculated based on each block classification unit (i.e., the sum of gradient values in the vertical direction, horizontal direction, first diagonal direction, and second diagonal direction, g) is used. v g h g d1 and g d2 At least one of the following performs a geometric transformation on the filter coefficients f(k,l). In this case, the geometric transformation of the filter coefficients is achieved by performing a 90° rotation, a 180° rotation, a 270° rotation, a second diagonal flip, a first diagonal flip, a vertical flip, a horizontal flip, a vertical and horizontal flip, or a scaling up / down operation on the filter, thereby producing a geometrically transformed filter.
[0642] On the other hand, after performing a geometric transformation on the filter coefficients, the reconstructed / decoded samples are filtered using the geometrically transformed filter coefficients. In this case, a geometric transformation is performed on at least one of the reconstructed / decoded samples that are the filtering targets, and then the reconstructed / decoded samples are filtered using the filter coefficients.
[0643] According to an embodiment of the present invention, geometric transformations are performed according to equations 33 to 35.
[0644] [Equation 33]
[0645] f D (k,l)=f(l,k)
[0646] [Equation 34]
[0647] f V (k,l)=f(k,Kl-1)
[0648] [Equation 35]
[0649] f R (k,l)=f(Kl-1,k)
[0650] Here, Equation 33 is an example of the equation used for the second diagonal flip, Equation 34 is an example of the vertical flip, and Equation 35 is an example of the 90° rotation. In Equations 34 to 35, K is the number of filter taps (filter length) in the horizontal and vertical directions, and "0 ≤ K and 1 ≤ K-1" represents the coordinates of the filter coefficients. For example, (0,0) represents the top left corner, and (K-1,K-1) represents the bottom right corner.
[0651] Table 1 shows an example of the geometric transformation applied to the filter coefficients f(k,l) based on the sum of the gradient values.
[0652] [Table 1]
[0653]
[0654]
[0655] Figure 39 This is a diagram illustrating filters obtained by performing geometric transformations on square filters, octagonal filters, snowflake filters, and rhombus filters according to embodiments of the present invention.
[0656] Reference Figure 39 The filter coefficients of square, octagonal, snowflake, and diamond filters undergo at least one geometric transformation, including a second diagonal flip, a vertical flip, and a 90° rotation. The filter coefficients obtained through this geometric transformation can then be used for filtering. Alternatively, after performing a geometric transformation on the filter coefficients, the reconstructed / decoded samples are filtered using the geometrically transformed filter coefficients. In this case, a geometric transformation is performed on at least one of the reconstructed / decoded samples that are the filtering targets, and then the reconstructed / decoded samples are filtered using the filter coefficients.
[0657] According to one embodiment of the present invention, filtering is performed on the reconstructed / decoded sample R(i,j) to generate a filtered decoded sample R′(i,j). The filtered decoded sample can be represented by Equation 36.
[0658] [Equation 36]
[0659]
[0660] In Equation 36, L is the number of filter taps (filter length) in the horizontal or vertical direction, and f(k,l) are the filter coefficients.
[0661] On the other hand, when filtering is performed, the offset value Y can be added to the filtered decoded sample R′(i,j). Entropy encoding / decoding can be performed on the offset value Y. Furthermore, the offset value Y is calculated using at least one statistical value from the current reconstructed / decoded sample value and the neighboring reconstructed / decoded sample values. Additionally, the offset value Y is determined based on at least one encoding parameter from the current reconstructed / decoded sample and the neighboring reconstructed / decoded sample. Here, the threshold E is a positive integer or zero.
[0662] Additionally, the filtered decoded samples can be truncated to represent N bits. Here, H is a positive integer. For example, when the filtered decoded samples, generated by filtering the reconstructed / decoded samples, are truncated to 10 bits, the final decoded sample value can be a value in the range of 0 to 1023.
[0663] According to an embodiment of the present invention, filtering of the chrominance component is performed based on filter information of the luminance component.
[0664] For example, filtering of the reconstructed image of the luminance component can only be performed if filtering of the reconstructed image of the luminance component was performed in a previous stage. Here, chrominance component reconstructed image filtering can be performed on U(Cr), V(Cb), or both of these components.
[0665] Alternatively, for example, in the case of the chroma component, filtering can be performed using at least one of the filter coefficients of the corresponding luminance component, the number of filter taps, the filter shape, and whether filtering is performed.
[0666] According to an exemplary embodiment of the present invention, when filtering is performed, if an unavailable sample exists near the current sample, padding is performed, and then filtering is performed using the padded sample. Padding refers to a method of copying the sample value of an adjacent available sample to an unavailable sample. Optionally, sample values or statistical values obtained based on the values of available samples adjacent to the unavailable sample are used. Padding can be performed repeatedly for P columns and R rows. Here, M and L are both positive integers.
[0667] Here, an unavailable sample point refers to a sample point located outside the boundaries of a CTU, CTB, strip, parallel block, parallel block group, or screen. Optionally, an unavailable sample point refers to a sample point belonging to at least one of the CTU, CTB, strip, parallel block, parallel block group, and screen that is different from the current sample point's CTU, CTB, strip, parallel block, parallel block group, and screen.
[0668] In addition, when performing filtering, predetermined samples may not be used.
[0669] For example, when performing filtering, it is possible to omit the filler samples.
[0670] Alternatively, for example, when performing filtering, if there are unavailable samples near the current sample, the unavailable samples may not be used during filtering.
[0671] Further alternatively, for example, when performing filtering, if a sample near the current sample is located outside the CTU or CTB, the neighboring samples near the current sample may not be used during filtering.
[0672] Additionally, when performing filtering, samples that have been subjected to at least one of deblocking filtering, adaptive sample offsetting, and adaptive in-loop filtering can be used.
[0673] Additionally, when performing filtering, if at least one of the samples present near the current sample is outside the CTU or CTB boundary, at least one of deblocking filtering, adaptive sample offset, and adaptive in-loop filtering may not be applied.
[0674] In addition, the target samples for filtering include unusable samples located outside the CTU or CTB boundaries. At least one of deblocking filtering, adaptive sample offsetting, and adaptive in-loop filtering is not performed on unusable samples, and the unusable samples are used for filtering as is.
[0675] According to an embodiment of the present invention, when filtering is performed, filtering is performed on at least one sample point located near the boundary of at least one of CU, PU, TU, block, block classification unit, CTU, and CTB. In this case, the boundary includes at least one of a vertical boundary, a horizontal boundary, and a diagonal boundary. Additionally, the sample point located near the boundary can be at least one of U rows, U columns, and U sample points adjacent to the boundary. Here, U is a positive integer.
[0676] According to an embodiment of the present invention, when filtering is performed, filtering is performed on at least one sample point located within a block, and filtering is not performed on sample points located outside the boundaries of at least one of CU, PU, TU, block, block classification unit, CTU, and CTB. In this case, the boundary includes at least one of a vertical boundary, a horizontal boundary, and a diagonal boundary. Additionally, sample points located near the boundary can be at least one of U rows, U columns, and U sample points adjacent to the boundary. Here, U is a positive integer.
[0677] According to an embodiment of the present invention, when filtering is performed, it is determined whether to perform filtering based on at least one of the coding parameters of the current block and neighboring blocks. In this case, the coding parameters include at least one of the following: prediction mode (i.e., whether the prediction is intra-frame prediction or inter-frame prediction), inter-frame prediction mode, intra-frame prediction mode, intra-frame prediction indicator, motion vector, reference frame index, quantization parameter, block size of the current block, block shape of the current block, size of block classification unit, and coding block flag / style.
[0678] Additionally, when performing filtering, at least one of the following is determined based on at least one of the coding parameters of the current block and neighboring blocks: filter coefficients, the number of filter taps (filter length), filter shape, and filter type. At least one of the following—filter coefficients, the number of filter taps (filter length), filter shape, and filter type—varies according to at least one of the coding parameters.
[0679] For example, the number of filters used for filtering is determined based on the quantization parameter. For instance, J filters are used when the quantization parameter is less than a threshold T. H filters are used when the quantization parameter is greater than a threshold R. In other cases, G filters are used. Here, T, R, J, H, and G are positive integers or zero. Furthermore, J is greater than or equal to H. Generally, the larger the quantization parameter value, the fewer filters are used.
[0680] Optionally, for example, the number of filters used for filtering can be determined based on the size of the current block. For instance, J filters are used when the current block size is less than a threshold T. H filters are used when the current block size is greater than a threshold R. In other cases, G filters are used. Here, T, R, J, H, and G are positive integers or zero. Furthermore, J is greater than or equal to H. The larger the block size, the fewer block filters are used.
[0681] Optionally, for example, the number of filters used for filtering is determined based on the size of the block classification unit. For instance, J filters are used when the size of the block classification unit is less than a threshold T. H filters are used when the size of the block classification unit is greater than a threshold R. In other cases, G filters are used. Here, T, R, J, H, and G are positive integers or zero. Additionally, J is greater than or equal to H. The larger the size of the block classification unit, the fewer block filters are used.
[0682] Further, alternatively, filtering can be performed, for example, by using any combination of the filtering methods described above.
[0683] The following section will describe the filter information encoding / decoding steps.
[0684] According to an embodiment of the present invention, filter information is entropy encoded / decoded to be placed between the strip header and the first CTU syntax element of the strip data within the bitstream.
[0685] In addition, filter information is entropy encoded / decoded to be arranged in the sequence parameter set, frame parameter set, strip header, strip data, parallel block header, parallel block group header, CTU or CTB within the bitstream.
[0686] On the other hand, the filter information includes at least one piece of information selected from the following: information on whether luma component filtering is performed, information on whether chroma component filtering is performed, filter coefficient values, number of filters, number of filter taps (filter length), filter shape information, filter type information, information on whether filtering is performed based on each strip, parallel block, parallel block group, frame, CTU, CTB, block, or CU, information on the number of times CU-based filtering is performed, CU maximum depth filtering information, information on whether CU-based filtering is performed, information on whether filters from previous reference frames are used, information on the filter index of previous reference frames, information on whether fixed filter information is used for block classification index information, index information for fixed filters, filter merging information, information on whether different filters are used for luma and chroma components respectively, and filter symmetry shape information.
[0687] Here, the number of filter taps refers to at least one of the following: the horizontal length of the filter, the vertical length of the filter, the first diagonal length of the filter, the second diagonal length of the filter, the horizontal and vertical lengths of the filter, and the number of filter coefficients within the filter.
[0688] On the other hand, the filter information includes up to L luminance filters. Here, L is a positive integer, specifically 25. Additionally, the filter information includes up to L chromatic aberration filters. Here, L is a positive integer, specifically 1.
[0689] On the other hand, a filter includes at most K luminance filter coefficients. Here, K is a positive integer, specifically 13. Additionally, the filter information includes at most K chrominance filter coefficients. Here, K is a positive integer, specifically 7.
[0690] For example, information about the symmetry shape of a filter is information about filter shapes such as point symmetry, horizontal symmetry, vertical symmetry, or combinations of point symmetry, horizontal symmetry, and vertical symmetry.
[0691] On the other hand, only some of the filter coefficients are transmitted using a signal. For example, when the filter is symmetrical, information about only one filter coefficient group in the symmetrical filter coefficient set is transmitted using a signal. Alternatively, for example, since the filter coefficients at the center of the filter can be implicitly derived, the filter coefficients at the center of the filter are not transmitted using a signal.
[0692] According to an embodiment of the invention, filter coefficient values in filter information are quantized in the encoder, and the quantized filter coefficient values are entropy encoded as a result. Similarly, the quantized filter coefficient values are entropy decoded in the decoder, and the quantized filter coefficient values are dequantized to recover the original filter coefficient values. The filter coefficient values are quantized to a range of values that can be represented by a fixed number of M bits, and then dequantized. Additionally, at least one filter coefficient is quantized to different bits and dequantized. Conversely, at least one of the filter coefficients can be quantized to the same number of bits and dequantized. The number of M bits is determined according to the quantization parameters. Furthermore, M in the M bits is a constant predefined in the encoder and decoder. Here, M can be a positive integer, specifically 8 or 10. The number of M bits can be less than or equal to the number of bits required to represent a sample in the encoder / decoder. For example, when the number of bits required to represent a sample is 10, then M can be 8. The first filter coefficient in the filter coefficients within the filter can be from -2... M Up to 2 M The values are in the range of -1, and the second filter coefficients can be from 0 to 2. M Values within the range of -1. Here, the first filter coefficient refers to the filter coefficients excluding the center filter coefficient, and the second filter coefficient refers to the center filter coefficient.
[0693] The filter coefficient values in the filter information can be cropped by at least one of the encoder and decoder, and at least one of the minimum and maximum values associated with the cropping can be entropy encoded / decoded. The filter coefficient values can be cropped to fall within the range of minimum to maximum. For each filter coefficient, at least one of the minimum and maximum values can be different. On the other hand, for each filter coefficient, at least one of the minimum and maximum values can be the same. At least one of the minimum and maximum values can be determined based on quantization parameters. At least one of the minimum and maximum values can be a constant value predefined in the encoder and decoder.
[0694] According to an embodiment of the present invention, at least one piece of filter information is entropy encoded / decoded based on at least one of the encoding parameters of the current block and neighboring blocks. In this case, the encoding parameters include at least one of the following: prediction mode (i.e., whether the prediction is intra-frame prediction or inter-frame prediction), inter-frame prediction mode, intra-frame prediction mode, intra-frame prediction indicator, motion vector, reference frame index, quantization parameter, block size of the current block, block shape of the current block, size of block classification unit, and encoding block flag / style.
[0695] For example, the number of filters in multiple filter information is determined based on the quantization parameters of the frame, strip, parallel block group, parallel block, CTU, CTB, or block. Specifically, when the quantization parameter is less than a threshold T, J filters are entropy encoded / decoded. When the quantization parameter is greater than a threshold R, H filters are entropy encoded / decoded. In other cases, G filters are entropy encoded / decoded. Here, T, R, J, H, and G are positive integers or zero. Furthermore, J is greater than or equal to H. The larger the quantization parameter value, the fewer filters are entropy encoded.
[0696] According to an exemplary embodiment of the present invention, filtering execution information (flags) is used to indicate whether filtering is performed on at least one of the luminance component and the chrominance component.
[0697] For example, filtering execution information (flags) is used based on each CTU, CTB, CU, or block to indicate whether filtering is performed on at least one of the luma and chroma components. For example, when the filtering execution information is a first value, filtering is performed based on each CTB, and when the filtering execution information is a second value, no filtering is performed on the corresponding CTB. In this case, entropy encoding / decoding can be performed on the information regarding whether filtering is performed on each CTB. Alternatively, for example, entropy encoding / decoding can be performed on information regarding the maximum depth or minimum size of the CU (maximum depth filter information of the CU), and entropy encoding / decoding can be performed on CU-based filtering execution information regarding the CU with the maximum depth or the CU with the minimum size.
[0698] For example, when a block can be partitioned into smaller square sub-blocks and non-square sub-blocks based on its block structure, CU-based flag entropy encoding / decoding can be performed until the block has a partition depth that allows it to be partitioned into smaller square sub-blocks. Conversely, CU-based flag entropy encoding / decoding can be performed until the block has a partition depth that allows it to be partitioned into smaller non-square sub-blocks.
[0699] Optionally, for example, the information regarding whether filtering is performed on at least one of the luminance and chrominance components can be based on block-based flags (i.e., flags based on each block). For example, filtering is performed on a block when the block-based flag for the corresponding block is a first value, and filtering is not performed when the block-based flag for the corresponding block is a second value. The block size is N×M, where N and M are positive integers.
[0700] Further optionally, for example, the information regarding whether filtering is performed on at least one of the luminance and chrominance components can be based on a CTU flag (i.e., based on the flag of each CTU). For example, filtering is performed on a CTU when the CTU-based flag of the corresponding CTU is a first value, and filtering is not performed when the CTU-based flag of the corresponding CTU is a second value. The size of the CTU is N×M, where N and M are positive integers.
[0701] Further, alternatively, for example, the determination of whether to perform filtering on at least one of the luminance and chrominance components can be based on the frame, strip, parallel block group, or parallel block type. Information regarding whether to perform filtering on at least one of the luminance and chrominance components can be based on flags for each frame, strip, parallel block group, or parallel block.
[0702] According to embodiments of the present invention, filter coefficients belonging to different block classifications can be merged to reduce the amount of filter coefficients that will be entropy encoded / decoded. In this case, entropy encoding / decoding is performed on filter merging information regarding whether or not filter coefficients are merged.
[0703] Furthermore, to reduce the amount of filter coefficients that will be entropy encoded / decoded, the filter coefficients of a reference frame can be used as the filter coefficients of the current frame. In this case, the method of using the filter coefficients of the reference frame is called temporal filter coefficient prediction. For example, temporal filter coefficient prediction is used for inter-frame prediction frames (B / P frames, stripes, parallel block groups, or parallel blocks). On the other hand, the filter coefficients of the reference frame are stored in memory. Additionally, when the filter coefficients of the reference frame are used for the current frame, entropy encoding / decoding of the filter coefficients of the current frame is omitted. In this case, entropy encoding / decoding is performed on the filter index of the previous reference frame that indicates which reference frame's filter coefficients are used.
[0704] For example, when using temporal filter coefficient prediction, a filter bank candidate list is constructed. The filter bank candidate list is empty until a new sequence is decoded. However, each time a frame is decoded, the frame's filter coefficients are added to the filter bank candidate list. When the number of filters in the filter bank candidate list reaches the maximum number G, new filters can replace the oldest filters in the decoding order. That is, the filter bank candidate list is updated in a first-in, first-out (FIFO) manner. Here, G is a positive integer, specifically 6. To prevent duplicate filters in the filter bank candidate list, filter coefficients from frames that are not used for temporal filter coefficient prediction can be added to the filter bank candidate list.
[0705] Optionally, for example, when using temporal filter coefficient prediction, a filter bank candidate list is constructed for multiple temporal layer indices to support temporal scalability. That is, a filter bank candidate list is constructed for each temporal layer. For example, the filter bank candidate list for a given temporal layer contains filter banks for decoding frames, where the temporal layer index of the decoded frame is equal to or less than the temporal layer index of previously decoded frames. Additionally, after decoding each frame, the filter coefficients for the current frame are added to a filter bank candidate list with a temporal layer index equal to or greater than the current frame's temporal layer index.
[0706] According to an embodiment of the present invention, a fixed filter bank is used to perform filtering.
[0707] Although temporal filter coefficient prediction cannot be used in intra-frame prediction frames (I-frames, stripes, parallel block groups, or parallel blocks), at least one of the up to 16 fixed filters within the filter bank can be used for filtering based on the block classification index. To transmit information from the encoder to the decoder regarding whether a fixed filter bank is used, entropy encoding / decoding is performed on information regarding whether a fixed filter is used for each block classification index. When a fixed filter is used, entropy encoding / decoding is also performed on the index information regarding the fixed filter. Even when a fixed filter is used for a specific block classification index, the filter coefficients are entropy encoded / decoded, and the reconstructed frame is filtered using the entropy-encoded / decoded filter coefficients and the fixed filter coefficients.
[0708] In addition, fixed filter banks are also used in inter-frame prediction frames (B / P frames, stripes, parallel block groups, or parallel blocks).
[0709] Alternatively, adaptive in-loop filtering can be performed using fixed filters without entropy encoding / decoding the filter coefficients. Here, a fixed filter may represent a filter bank predefined in the encoder and decoder. In this case, without entropy encoding / decoding the filter coefficients, the encoder and decoder entropy encode / decode the fixed filter index information, which indicates which filter in the filter bank or which filter bank predefined in the encoder and decoder is used. In this case, filtering is performed using fixed filters that differ in at least one aspect of filter coefficient values, filter taps (i.e., the number of filter taps or the filter length), and filter shape, based on at least one of block classification, block, CU, stripe, parallel block, parallel block group, and frame.
[0710] On the other hand, at least one filter within a fixed filter bank can be transformed in terms of filter taps and / or filter shape. For example, such as Figures 40a to 40dThe diagram illustrates the transformation of coefficients in a 9×9 diamond filter into coefficients in a 5×5 square filter. Specifically, the coefficients in a 9×9 diamond filter can be transformed into coefficients in a 5×5 square filter.
[0711] For example, the sum of the filter coefficients corresponding to filter coefficient indices 0, 2 and 6 in the 9×9 rhombus shape is assigned to filter coefficient index 2 in the 5×5 square shape.
[0712] Alternatively, for example, the sum of the filter coefficients corresponding to filter coefficient indices 1 and 5 in the 9×9 rhombus shape can be assigned to filter coefficient index 1 in the 5×5 square shape.
[0713] Further alternatively, for example, the sum of the filter coefficients corresponding to filter coefficient indices 3 and 7 in the 9×9 rhombus shape is assigned to filter coefficient index 3 in the 5×5 square shape.
[0714] Further, alternatively, for example, the filter coefficients corresponding to filter coefficient index 4 in the 9×9 rhombus shape are assigned to filter coefficient index 0 in the 5×5 square shape.
[0715] Further, alternatively, for example, the filter coefficients corresponding to filter coefficient index 8 in the 9×9 rhombus shape are assigned to filter coefficient index 4 in the 5×5 square shape.
[0716] Further alternatively, for example, the sum of the filter coefficients corresponding to filter coefficient indices 9 and 10 in the 9×9 rhombus shape is assigned to filter coefficient index 5 in the 5×5 square shape.
[0717] Further alternatively, for example, the filter coefficients corresponding to filter coefficient index 11 in the 9×9 rhombus shape are assigned to filter coefficient index 6 in the 5×5 square shape.
[0718] Further, alternatively, for example, the filter coefficients corresponding to filter coefficient index 12 in the 9×9 rhombus shape are assigned to filter coefficient index 7 in the 5×5 square shape.
[0719] Further, alternatively, for example, the filter coefficients corresponding to filter coefficient index 13 in the 9×9 rhombus shape are assigned to filter coefficient index 8 in the 5×5 square shape.
[0720] Further, alternatively, for example, the sum of the filter coefficients corresponding to filter coefficient indices 14 and 15 in the 9×9 rhombus shape is assigned to filter coefficient index 9 in the 5×5 square shape.
[0721] Further optionally, for example, the sum of the filter coefficients corresponding to filter coefficient indices 16, 17 and 18 in the 9×9 rhombus shape is assigned to filter coefficient index 10 in the 5×5 square shape.
[0722] Further, alternatively, for example, the filter coefficients corresponding to filter coefficient index 19 in the 9×9 rhombus shape are assigned to filter coefficient index 11 in the 5×5 square shape.
[0723] Further, alternatively, for example, the filter coefficients corresponding to filter coefficient index 20 in the 9×9 rhombus shape are assigned to filter coefficient index 12 in the 5×5 square shape.
[0724] Table 2 illustrates an exemplary method for generating filter coefficients by transforming 9×9 diamond filter coefficients into 5×5 square filter coefficients.
[0725] [Table 2]
[0726]
[0727]
[0728]
[0729] In Table 2, the sum of at least one of the filter coefficients of a 9×9 diamond filter is equal to the sum of at least one of the filter coefficients of the corresponding 5×5 square filter.
[0730] On the other hand, when using a maximum of 16 fixed filter banks for the coefficients of a 9×9 diamond filter, a maximum of 21 filter coefficients × 25 filters × 16 filter types need to be stored in memory. When a maximum of 16 fixed filter banks are used for the filter coefficients of a 5×5 square filter, a maximum of 13 filter coefficients × 25 filters × 16 filter types need to be stored in memory. Here, since the memory size required to store fixed filter coefficients in a 5×5 square filter is smaller than the memory size required to store fixed filter coefficients in a 9×9 diamond filter, the memory capacity requirement and memory access bandwidth are reduced.
[0731] On the other hand, the chroma component of the reconstructed / decoded component can be filtered using a filter obtained by transforming the filter for the co-occurring luminance component in terms of filter taps and / or filter shape.
[0732] According to an embodiment of the present invention, it is prohibited to predict filter coefficients from the filter coefficients of a predefined fixed filter.
[0733] According to an embodiment of the invention, multiplication is replaced by shift operations. First, the filter coefficients used to perform filtering on the luminance and / or chrominance blocks are divided into two groups. For example, the filter coefficients are divided into a first group including coefficients {L0, L1, L2, L3, L4, L5, L7, L8, L9, L10, L14, L15, L16, and L17} and a second group including the remaining coefficients. The first group is limited to including only coefficient values {-64, -32, -16, -8, -4, 0, 4, 8, 16, 32, and 64}. In this case, the multiplication of the filter coefficients included in the first group and the reconstructed / decoded samples can be performed with a single bit shift operation. Therefore, the filter coefficients included in the first group are mapped to pre-bindified bit shift values to reduce the overhead of signal transmission.
[0734] According to an embodiment of the present invention, as a result of determining whether to perform block classification and / or filtering on the chroma component, the result of determining whether to perform block classification and / or filtering on the corresponding luminance component is used as is. Furthermore, as filter coefficients for the chroma component, filter coefficients already used for the corresponding luminance component are used. For example, a predetermined 5×5 diamond filter is used.
[0735] As an example, the filter coefficients in a 9×9 filter used for the luminance component can be transformed into the filter coefficients in a 5×5 filter used for the chrominance component. In this case, the outermost filter coefficients are set to zero.
[0736] As another example, when filter coefficients in the form of a 5×5 filter are used for the luminance component, the filter coefficients used for the luminance component are the same as those used for the chrominance component. That is, the filter coefficients used for the luminance component can be used as filter coefficients for the chrominance component without modification.
[0737] As another example, in order to maintain the shape of the 5×5 filter used to filter the chromaticity components, the filter coefficients outside the 5×5 diamond filter are replaced by coefficients arranged at the boundaries of the 5×5 diamond filter.
[0738] On the other hand, intra-loop filtering for the luma block and intra-loop filtering for the chroma block can be performed separately. Control flags are transmitted via signals at the picture, strip, parallel block group, parallel block, CTU, or CTB level to indicate whether adaptive intra-loop filtering for the chroma component is supported independently. Flags indicating whether adaptive intra-loop filtering is performed jointly for both luma and chroma blocks, or individually for both luma and chroma blocks, can be transmitted via signals.
[0739] According to embodiments of the present invention, when entropy encoding / decoding at least one piece of filter information, at least one of the following binarization methods can be used:
[0740] Truncation Rice binarization method;
[0741] K-order exponential Columbus binarization method;
[0742] Finite K-order exponential Columbus binarization method;
[0743] Fixed-length binarization method;
[0744] Univariate binarization methods; and
[0745] A truncated univariate binarization method.
[0746] As an example, entropy encoding / decoding of the filter coefficients of the luminance filter and the filter coefficients of the chrominance filter are performed using different binarization methods for the luminance filter and the chrominance filter.
[0747] As another example, entropy encoding / decoding of the filter coefficients of a luminance filter is performed using different binarization methods. As yet another example, entropy encoding / decoding of the filter coefficients of a luminance filter is performed using the same binarization method.
[0748] As another example, entropy encoding / decoding of the filter coefficients of a chroma filter is performed using different binarization methods. As yet another example, entropy encoding / decoding of the filter coefficients of a chroma filter is performed using the same binarization method.
[0749] When entropy encoding / decoding at least one filter information, as an example, at least one filter information of at least one neighboring block in a neighboring block, or at least one previously encoded / decoded filter information, or the encoded / decoded filter information in a previous frame, is used to determine the context model.
[0750] As another example, when entropy encoding / decoding at least one filter piece of information, the context model is determined using at least one filter piece of information with different components.
[0751] As another example, when entropy encoding / decoding filter coefficients, at least one of the filter coefficients in the filter is used to determine the context model.
[0752] As another example, when entropy encoding / decoding at least one filter information, at least one filter information of at least one neighboring block in a neighboring block, or at least one previously encoded / decoded filter information, or the encoded / decoded filter information in a previous frame, is used to determine the context model.
[0753] As another example, when entropy encoding / decoding is performed on at least one filter information, at least one filter information with different components is used as the predicted value of the filter information to perform entropy encoding / decoding.
[0754] As another example, when entropy encoding / decoding filter coefficients, at least one of the filter coefficients within the filter is used as a prediction value to perform entropy encoding / decoding.
[0755] As another example, any combination of filter information entropy encoding / decoding methods can be used to entropy encode / decode filter information.
[0756] According to embodiments of the present invention, adaptive in-loop filtering is performed on a unit consisting of at least one of the following: block, CU, PU, TU, CB, PB, TB, CTU, CTB, strip, parallel block, parallel block group, and frame. When adaptive in-loop filtering is performed on each of the above units, it means that a block classification step, a filtering execution step, and a filter information encoding / step are performed on a unit consisting of at least one of the following: block, CU, PU, TU, CB, PB, TB, CTU, CTB, strip, parallel block, parallel block group, and frame.
[0757] According to an embodiment of the present invention, whether to perform adaptive in-loop filtering is determined based on whether at least one of deblocking filtering, sample adaptive offsetting, and bidirectional filtering is performed.
[0758] As an example, adaptive in-loop filtering is performed on samples in the current frame that have undergone at least one of deblocking filtering, adaptive sample offsetting, and bidirectional filtering.
[0759] As another example, adaptive in-loop filtering is not performed on samples in the current frame that have already undergone at least one of deblocking filtering, sample adaptive offsetting, and bidirectional filtering.
[0760] As another example, for reconstructed / decoded samples in the current frame that have already undergone at least one of deblocking filtering, adaptive sample offsetting, and bidirectional filtering, adaptive in-loop filtering is performed on the reconstructed / decoded samples in the current frame using L filters, without performing block classification. Here, L is a positive integer.
[0761] According to an embodiment of the present invention, whether to perform adaptive in-loop filtering is determined based on the strip or parallel block group type of the current frame.
[0762] As an example, adaptive in-loop filtering is only performed when the current frame's strip or parallel block group type is I strip or I parallel block group.
[0763] As another example, adaptive in-loop filtering is performed when the current frame's strip or parallel block group type is at least one of I strip, B strip, P strip, I parallel block group, B parallel block group, and P parallel block group.
[0764] As an example, when the current frame's stripe or parallel block group type is at least one of I-strip, B-strip, P-strip, I-parallel block group, B-parallel block group, and P-parallel block group, adaptive intra-loop filtering is performed on the reconstructed / decoded samples within the current frame using L filters, without performing block classification. Here, L is a positive integer.
[0765] As another example, when the current frame's strip or parallel block group type is at least one of I strip, B strip, P strip, I parallel block group, B parallel block group, and P parallel block group, a filter shape is used to perform adaptive in-loop filtering.
[0766] As another example, when the current frame's strip or parallel block group type is at least one of I strip, B strip, P strip, I parallel block group, B parallel block group, and P parallel block group, a filter tap is used to perform adaptive in-loop filtering.
[0767] As another example, when the current frame's stripe or parallel block group type is at least one of I-strip, B-strip, P-strip, I-parallel block group, B-parallel block group, and P-parallel block group, at least one of block classification and adaptive in-loop filtering is performed based on each M×N block. In this case, both M and N are positive integers. Specifically, both M and N are 4.
[0768] According to an embodiment of the present invention, whether to perform adaptive in-loop filtering is determined based on whether the current frame is used as a reference frame.
[0769] For example, when using the current frame as a reference frame when encoding / decoding subsequent frames, adaptive in-loop filtering is performed on the current frame.
[0770] As another example, when the current frame is not used as a reference frame during the encoding / decoding of subsequent frames, adaptive in-loop filtering is not performed on the current frame.
[0771] As another example, when the current frame is not used in processing subsequent frames, adaptive intra-loop filtering is performed on the reconstructed / decoded samples within the current frame using L filters, without performing block classification. Here, L is a positive integer.
[0772] As another example, when the current frame is not used when encoding / decoding subsequent frames, adaptive in-loop filtering is performed using a filter shape.
[0773] As another example, when the current image is not used when encoding / decoding subsequent frames, an adaptive in-loop filtering is performed using a filter tap.
[0774] As another example, when the current frame is not used during encoding / decoding of subsequent frames, at least one of block classification and filtering is performed based on each N×M block. In this case, both M and N are positive integers. Specifically, both M and N are 4.
[0775] According to an embodiment of the present invention, whether to perform adaptive in-loop filtering is determined based on the time layer identifier.
[0776] As an example, when the time layer identifier of the current frame is zero, representing the underlying layer, adaptive in-loop filtering is performed on the current frame.
[0777] As another example, when the time layer identifier of the current frame is 4, representing the top layer, adaptive in-loop filtering is performed.
[0778] As another example, the temporal layer identifier for the current frame is 4, representing the top layer. When performing adaptive intra-loop filtering on the current frame, L filters are used to perform adaptive intra-loop filtering on the reconstructed / decoded samples within the current frame, without performing block classification. Here, L is a positive integer.
[0779] As another example, when the time layer identifier of the current frame is 4, representing the top layer, adaptive in-loop filtering is performed using a filter shape.
[0780] As another example, when the time layer identifier of the current frame is 4, representing the top layer, an adaptive in-loop filtering is performed using a filter tap.
[0781] As another example, when the time layer identifier of the current frame is 4, representing the top layer, at least one of block classification and adaptive in-loop filtering is performed based on each N×M block. In this case, both M and N are positive integers. Specifically, both M and N are 4.
[0782] According to an embodiment of the present invention, at least one of the block classification methods is performed based on a time layer identifier.
[0783] For example, when the time layer identifier of the current frame is zero, representing the underlying layer, at least one of the block classification methods described above is performed on the current frame.
[0784] Optionally, when the time layer identifier of the current frame is 4, representing the top layer, at least one of the block classification methods described above is performed on the current frame.
[0785] According to an embodiment of the present invention, at least one of the above-described block classification methods is performed based on the value of the time layer identifier.
[0786] As another example, when the time layer identifier of the current frame is 4, representing the top layer, adaptive intra-loop filtering is performed on the reconstructed / decoded samples within the current frame using L filters, without performing block classification. Here, L is a positive integer.
[0787] As another example, when the time layer identifier of the current frame is 4, representing the top layer, adaptive in-loop filtering is performed using a filter shape.
[0788] As another example, when the time layer identifier of the current frame is 4, representing the top layer, an adaptive in-loop filtering is performed using a filter tap.
[0789] As another example, when the time layer identifier of the current frame is 4, representing the top layer, at least one of block classification and adaptive in-loop filtering is performed based on each N×M block. In this case, both M and N are positive integers. Specifically, both M and N are 4.
[0790] As another example, when performing adaptive intra-loop filtering on the current frame, L filters are used to perform adaptive intra-loop filtering on the reconstructed / decoded samples within the current frame, without performing block classification. Here, L is a positive integer. Optionally, in this case, adaptive intra-loop filtering is performed on the reconstructed / decoded samples within the current frame using L filters, without performing block classification and regardless of the temporal layer identifier.
[0791] On the other hand, when performing adaptive intra-loop filtering on the current frame, L filters are used to perform adaptive intra-loop filtering on the reconstructed / decoded samples within the current frame, regardless of whether block classification is performed. Here, L is a positive integer. In this case, L filters can be used to perform adaptive intra-loop filtering on the reconstructed / decoded samples within the current frame without performing block classification, and regardless of the temporal layer identifier or whether block classification is performed.
[0792] On the other hand, an adaptive intra-loop filtering can be performed using a filter shape. In this case, an adaptive intra-loop filtering can be performed on the reconstructed / decoded samples within the current image using a filter shape, without requiring block classification. Alternatively, an adaptive intra-loop filtering can be performed on the reconstructed / decoded samples within the current image using a filter shape, regardless of whether block classification is performed.
[0793] On the other hand, an adaptive intra-loop filtering can be performed using a single filter tap. In this case, adaptive intra-loop filtering can be performed on the reconstructed / decoded samples within the current image using a single filter tap, without requiring block classification. Alternatively, adaptive intra-loop filtering can be performed on the reconstructed / decoded samples within the current image using a single filter tap, regardless of whether block classification is performed.
[0794] On the other hand, adaptive in-loop filtering can be performed based on specific units. For example, a specific unit can be at least one of a frame, strip, parallel block, parallel block group, CTU, CTB, CU, PU, TU, CB, PB, TB, and M×N sized blocks. Here, M and N are both positive integers. M and N can be the same integer or different integers. Furthermore, M, N, or both M and N are predefined values in the encoder / decoder. Optionally, M, N, or both M and N can be values transmitted from the encoder to the decoder via signals.
[0795] Figures 41a to 41d 5 is a diagram illustrating an exemplary method for determining the sum of gradient values for the vertical direction, horizontal direction, first diagonal direction, and second diagonal direction based on subsampling.
[0796] Reference Figures 42a to 42d 5. Perform filtering based on each 4×4 lumen block. In this case, different filter coefficients can be used to perform filtering for each 4×4 lumen block. A subsampled Laplacian operation can be performed to classify the 4×4 lumen blocks. Furthermore, the filter coefficients used for filtering vary for each 4×4 lumen block. Additionally, the 4×4 lumen blocks are classified into up to 25 categories. Furthermore, a classification index corresponding to the filter index of the 4×4 lumen block can be derived based on the block's directionality value and / or quantization activity value. Here, to calculate the directionality value and / or quantization activity value for each 4×4 lumen block, the sum of gradient values for the vertical direction, horizontal direction, first diagonal direction, and second diagonal direction is calculated by summing the results of the one-dimensional Laplacian operation calculated at the subsampled positions within the 8×8 block.
[0797] Specifically, refer to Figure 43 In the case of block classification based on each 4×4 block, the sum of gradient values g for the vertical, horizontal, first diagonal, and second diagonal directions is calculated based on subsampling. v g h g d1 and g d2At least one of the following (hereinafter referred to as the "first method"). Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplace operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplace operations are performed at positions V, H, D1, and D2 along the horizontal, vertical, first diagonal, and second diagonal directions, respectively. Furthermore, the position where the one-dimensional Laplace operation is performed can be the position of a subsample. Figure 43 In this approach, a block classification index C is assigned based on each 4×4 block (i.e., the shaded area). In this case, the computational range for calculating the one-dimensional Laplace sum can be larger than the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample point locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplace sum.
[0798] here, Figure 43 An exemplary block-based encoding / decoding process using the first method is shown. Figures 44a to 44d Another exemplary block-based encoding / decoding process using the first method is shown in one dimension. Figures 45a to 45d Another exemplary block-based encoding / decoding process using the first method is shown in two dimensions.
[0799] Reference Figures 46a to 46d In the case of block classification based on each 4×4 block, the sum of gradient values g for the vertical, horizontal, first diagonal, and second diagonal directions is calculated based on subsampling. v g h g d1 and g d2 At least one of the following (hereinafter referred to as the "second method"). Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplace operations for the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplace operations are performed at positions V, H, D1, and D2 along the horizontal, vertical, first diagonal, and second diagonal directions, respectively. Furthermore, the position where the one-dimensional Laplace operation is performed can be the position of a subsample. Figures 47a to 47d In this approach, a block classification index C is assigned based on each 4×4 block (i.e., the shaded area). In this case, the computational range for calculating the one-dimensional Laplace sum can be larger than the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplace sum.
[0800] Specifically, the second method means that when both x and y values are even, or when both x and y values are odd, a one-dimensional Laplace operation is performed at position (x, y). When neither x nor y value is even, or neither x nor y value is odd, the result of the one-dimensional Laplace operation at position (x, y) is assigned zero. In other words, it means performing a one-dimensional Laplace operation using a chessboard pattern based on the x and y values.
[0801] Reference Figure 48 The position for performing the one-dimensional Laplace operation is the same for the horizontal, vertical, first diagonal, and second diagonal directions. In other words, regardless of the direction of the vertical, horizontal, first diagonal, and second diagonal directions, a uniform subsampled position is used to perform the one-dimensional Laplace operation for each direction.
[0802] here, Figure 48 An exemplary block-based encoding / decoding process using the second method is shown. Figure 48 Another exemplary block-based encoding / decoding process using the second method is shown. Figures 49a to 49d Another exemplary block-based encoding / decoding process using the first method is shown. Figures 50a to 50d Another exemplary block-based encoding / decoding process using the first method is shown.
[0803] Reference Figures 51a to 51d In the case of block classification based on each 4×4 block, the sum of gradient values g for the vertical, horizontal, first diagonal, and second diagonal directions is calculated based on subsampling. v g h g d1 and g d2 At least one of the following (hereinafter referred to as the "third method"). Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplace operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplace operations are performed at positions V, H, D1, and D2 along the horizontal, vertical, first diagonal, and second diagonal directions, respectively. Furthermore, the position where the one-dimensional Laplace operation is performed can be the position of a subsample. Figure 52 In this context, a block classification index C is assigned based on each 4×4 block (i.e., the shaded area). In this case, the computational range for calculating the one-dimensional Laplacian sum can be larger than the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample point locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplacian sum.
[0804] Specifically, the third method involves performing a one-dimensional Laplace operation at position (x, y) when either the x-coordinate or the y-coordinate is even and the other is odd. When both the x-coordinate and the y-coordinate are even or odd, the result of the one-dimensional Laplace operation at position (x, y) is assigned zero. In other words, it means performing a one-dimensional Laplace operation using a chessboard pattern based on the x-coordinate and the y-coordinate.
[0805] Reference Figure 52 The positions for performing one-dimensional Laplace operations in the horizontal, vertical, first diagonal, and second diagonal directions are the same. In other words, regardless of the direction, a uniform subsampled one-dimensional Laplace operation position is used to perform the one-dimensional Laplace operation for each direction.
[0806] here, Figure 52 An exemplary block-based encoding / decoding process using a third method is shown. Figures 53a to 53d This illustrates another exemplary block-based encoding / decoding process using a third method. Figures 54a to 54d Another exemplary block-based encoding / decoding process using a third method is shown in one dimension.
[0807] Reference Figures 55a to 55d In the case of block classification based on each 4×4 block, the sum of gradient values g for the vertical, horizontal, first diagonal, and second diagonal directions is calculated based on subsampling. v g h g d1 and g d2 At least one of the following (hereinafter referred to as the "fourth method"). Here, V, H, D1, and D2 represent the results of sample-based one-dimensional Laplace operations along the vertical, horizontal, first diagonal, and second diagonal directions, respectively. That is, one-dimensional Laplace operations are performed at positions V, H, D1, and D2 along the horizontal, vertical, first diagonal, and second diagonal directions, respectively. Furthermore, the position where the one-dimensional Laplace operation is performed can be the position of a subsample. Figures 39 to 5 In this approach, a block classification index C is assigned based on each 4×4 block (i.e., the shaded area). In this case, the computational range for calculating the one-dimensional Laplacian sum can be larger than the size of the block classification unit. Here, thin solid rectangles represent the reconstructed sample point locations, and thick solid rectangles represent the computational range for calculating the one-dimensional Laplacian sum.
[0808] Specifically, the fourth method means performing a one-dimensional Laplace operation along the vertical direction at the subsampled position (x,y), instead of performing subsampling along the horizontal direction. In other words, it means skipping one row when performing a one-dimensional Laplace operation.
[0809] Reference The positions for performing one-dimensional Laplace operations in the horizontal, vertical, first diagonal, and second diagonal directions are the same. In other words, regardless of the direction, a uniform subsampled one-dimensional Laplace operation position is used to perform the one-dimensional Laplace operation for each direction.
[0810] here, An exemplary block-based encoding / decoding process using the fourth method is shown. Another exemplary block-based encoding / decoding process using the fourth method is shown. Another exemplary block-based encoding / decoding process using the fourth method is shown.
[0811] On the other hand, by using The gradient value calculated by the method shown in Figure 5 is used to derive the quantized activity value A of the directionality value and the activity value A. q At least one of the methods is similar to the in-loop filtering method described above.
[0812] On the other hand, in subsampling-based gradient value calculation methods, the one-dimensional Laplacian operation is not calculated for all samples within the operational range (e.g., an 8×8 block) used to compute the one-dimensional Laplacian sum, but rather for the positions of subsamples within that operational range. This reduces the number of computations required for block classification (e.g., multiplication, shift operations, addition, and absolute value calculations). Consequently, the computational complexity in both the encoder and decoder is reduced.
[0813] According to methods one through four, in the case of 4×4 block classification based on one-dimensional Laplacian operations using subsampling, the results V, H, D1, and D2 of the one-dimensional Laplacian operations calculated at the sample locations within an 8×8 block are added to a 4×4 lumen block to derive gradient values for the vertical, horizontal, first diagonal, and second diagonal directions, respectively. Therefore, to calculate all gradient values within the 8×8 range, 720+240 additions, 288 comparisons, and 144 shifts are required.
[0814] On the other hand, according to the traditional in-loop filtering method, in the case of 4×4 block classification, V, H, D1, and D2, which are the results of the one-dimensional Laplacian operation calculated at all positions within an 8×8 range, are used for the 4×4 lumen block to derive the gradient values in the vertical, horizontal, first diagonal, and second diagonal directions. Therefore, to calculate all gradient values within the 8×8 range, 1586+240 additions, 576 comparisons, and 144 shifts are required.
[0815] Then, the gradient value is used to derive the quantized activity value A of the directionality value D and the activity value A. q The processing requires 8 additions, 28 comparisons, 8 multiplications, and 20 shifts.
[0816] Therefore, the block classification methods using the first to fourth methods within an 8×8 area require a total of 968 additions, 316 comparisons, 8 multiplications, and 164 shifts. Consequently, each sample point requires 15.125 additions, 4.9375 comparisons, 0.125 multiplications, and 2.5625 shifts.
[0817] On the other hand, block classification methods using traditional in-loop filter...
Claims
1. A video decoding method, comprising: Decode the filter information; The basic blocks in the coding tree unit are classified into one of several classes; The block classification index is assigned to the basic block in the coding tree unit; and By using the filter information and the block classification index, a filter is applied to the samples of the basic blocks in the coding tree unit, and The block classification index is determined based on directional and activity information. Wherein, at least one of the directional information and the activity information is determined based on a gradient value for at least one of the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. The gradient value is obtained by applying the Laplace operation to the basic block. The Laplace operation is performed only on specific samples included in the basic block. Wherein, the horizontal and vertical positions of the specific sample point are both even-numbered positions or both odd-numbered positions, and The Laplace operation is a one-dimensional Laplace operation.
2. A video encoding method, comprising: The basic blocks in the coding tree unit are classified into one of several classes; Assign the block classification index to the basic block in the coding tree unit; A filter is applied to samples of the basic blocks in the coding tree unit using filter information and the block classification index; and The filter information is encoded, and The block classification index is determined based on directional and activity information. Wherein, at least one of the directional information and the activity information is determined based on a gradient value for at least one of the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. The gradient value is obtained by applying the Laplace operation to the basic block. The Laplace operation is performed only on specific samples included in the basic block. Wherein, the horizontal and vertical positions of the specific sample point are both even-numbered positions or both odd-numbered positions, and The Laplace operation is a one-dimensional Laplace operation.
3. An apparatus for transmitting a bit stream, the apparatus comprising: The processor is configured to generate the bitstream by encoding the image; as well as A transmitter is configured to transmit the bit stream. In order to generate the bitstream, the processor is configured to: The basic blocks in the coding tree unit are classified into one of several classes; Assign the block classification index to the basic block in the coding tree unit; A filter is applied to samples of the basic blocks in the coding tree unit using filter information and the block classification index; and The filter information is encoded, and The block classification index is determined based on directional and activity information. Wherein, at least one of the directional information and the activity information is determined based on a gradient value for at least one of the vertical direction, the horizontal direction, the first diagonal direction, and the second diagonal direction. The gradient value is obtained by applying the Laplace operation to the basic block. The Laplace operation is performed only on specific samples included in the basic block. Wherein, the horizontal and vertical positions of the specific sample points are both even positions or both are odd positions, and The Laplace operation is a one-dimensional Laplace operation.
Citation Information
Patent Citations
Video coding and decoding methods and video coding and decoding devices using adaptive loop filtering
CN102474615A
Flexible Region Based Sample Adaptive Offset (SAO) and Adaptive Loop Filter (ALF)
US20130051455A1