Image encoding, decoding method, and recording medium storing bitstream
By selecting an appropriate filtering method based on image block features during image encoding/decoding, the encoding/decoding efficiency is improved, the problem of low efficiency in traditional methods is solved, and image quality is improved.
Patent Information
- Application Number
- CN202211278534.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-10-20
- Filing Date
- 2018-10-19
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2038-10-19
AI Technical Summary
Traditional image encoding/decoding methods are limited in the types and applications of filtering methods, resulting in low encoding/decoding efficiency, which increases costs, especially in the transmission and storage of high-resolution and high-quality images.
Multiple filtering methods are employed, and different types of filters (smoothing filters, edge-preserving filters, and noise filters) are selected based on the characteristics of image blocks (uniform regions, edge regions, and pseudo-edge regions). The filter length and filtering operation are determined by combining factors such as block size, shape, and intra-frame prediction mode.
It improves the efficiency of image encoding/decoding, reduces computational complexity, and improves the ringing effect in the target boundary region of the image and the contour effect in orientation prediction.
Smart Images

Figure CN115484458B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of October 19, 2018, the application number of "201880068209.8", and the title of "Image encoding / decoding method and apparatus and recording medium storing bitstream". TECHNICAL FIELD
[0002] The present invention relates to an image encoding / decoding method and apparatus and a recording medium storing a bitstream. More particularly, the present invention relates to an image encoding / decoding method and apparatus using various filtering methods. BACKGROUND
[0003] Recently, the demand for high resolution and high quality images such as high definition (HD) images and ultra high definition (UHD) images has increased in various application fields. However, higher resolution and higher quality image data has an increased amount of data compared to conventional image data. Therefore, when image data is transmitted by using a medium such as a conventional wired and wireless broadband network, or when image data is stored by using a conventional storage medium, the cost of transmission and storage increases. In order to solve these problems that occur as the resolution and quality of image data increase, an efficient image encoding / decoding technique is required for higher resolution and higher quality images.
[0004] Image compression techniques include various techniques including an inter prediction technique of predicting pixel values included in a current picture from previous or subsequent pictures of the current picture, an intra prediction technique of predicting pixel values included in a current picture by using pixel information in the current picture, a transform and quantization technique for compressing the energy of a residual signal, an entropy encoding technique of assigning a short code to a value having a high frequency of occurrence and a long code to a value having a low frequency of occurrence, and the like. Image data can be efficiently compressed by using such image compression techniques, and can be transmitted or stored.
[0005] The type of filtering method used in a conventional image encoding / decoding method and apparatus and the method to which the type is applied are limited, and thus encoding / decoding is limited. SUMMARY
[0006] TECHNICAL PROBLEM
[0007] The present invention provides various filtering methods performed at each step when image encoding / decoding is performed, to improve the encoding / decoding efficiency of images.
[0008] TECHNICAL SOLUTION
[0009] A method of decoding an image according to the present invention can include determining reference samples of a current block, performing filtering for the reference samples based on a characteristic of a region including the reference samples, and performing intra prediction by using the reference samples for which the filtering is performed.
[0010] In the method of decoding an image, the characteristic of the region including the reference samples is any one of a uniform region, an edge region, and a pseudo edge region.
[0011] In the method of decoding an image, in the step of performing filtering for the reference samples, when the characteristic of the region including the reference samples is the uniform region, the filtering is performed by using a smoothing filter.
[0012] In the method of decoding an image, in the step of performing filtering for the reference samples, when the characteristic of the region including the reference samples is the edge region, the filtering is performed by using an edge preserving filter.
[0013] In the method of decoding an image, in the step of performing filtering for the reference samples, when the characteristic of the region including the reference samples is the pseudo edge region, the filtering is performed by excluding samples determined to be noise.
[0014] In the method of decoding an image, the characteristic of the region including the reference samples is determined based on a uniformity of the region.
[0015] In the method of decoding an image, further including determining whether to perform filtering for the reference samples based on at least one of a size of the current block, a shape of the current block, an intra prediction mode of the current block, a partition depth of the current block, and a pixel component of the current block, wherein the operation of performing filtering for the reference samples is performed based on a result of the determination.
[0016] In the method of decoding an image, the step of performing filtering for the reference samples includes determining a filter length based on at least one of a size of the current block, a shape of the current block, an intra prediction mode of the current block, a partition depth of the current block, and a pixel component of the current block, and performing filtering for the reference samples based on the determined filter length.
[0017] In the method of decoding an image, the reference samples of the current block are at least one of at least one line of reconstructed samples located at a left side of the current block and at least one line of reconstructed samples located at an upper side of the current block.
[0018] A method of encoding an image according to the present invention can include determining reference samples of a current block, performing filtering for the reference samples based on a characteristic of a region including the reference samples, and performing intra prediction by using the reference samples for which the filtering is performed.
[0019] In the method of encoding an image, the characteristic of the region including the reference samples is any one of a uniform region, an edge region, and a pseudo edge region.
[0020] In the method of encoding an image, in the step of performing filtering for the reference samples, when the characteristic of the region including the reference samples is the uniform region, the filtering is performed by using a smoothing filter.
[0021] In the method of encoding an image, in the step of performing filtering for the reference samples, when the characteristic of the region including the reference samples is the edge region, the filtering is performed by using an edge preserving filter.
[0022] In the method of encoding an image, in the step of performing filtering for the reference samples, when the characteristic of the region including the reference samples is the pseudo edge region, the filtering is performed by excluding pixels determined as noise.
[0023] In the method of encoding an image, the characteristic of the region including the reference samples is determined based on a uniformity of the region.
[0024] In the method of encoding an image, further including determining whether to perform filtering for the reference samples based on at least one of a size of the current block, a shape of the current block, an intra prediction mode of the current block, a partition depth of the current block, and a pixel component of the current block, wherein the filtering for the reference samples is performed based on the determined result.
[0025] In the method of encoding an image, the step of performing filtering for the reference samples includes determining a filter length based on at least one of a size of the current block, a shape of the current block, an intra prediction mode of the current block, a partition depth of the current block, and a pixel component of the current block, and performing filtering for the reference samples based on the determined filter length.
[0026] In the method of encoding an image, the reference samples of the current block are at least one of at least one line of reconstructed samples located at a left side of the current block and at least one line of reconstructed samples located at an upper side of the current block.
[0027] A recording medium according to the present invention stores a bitstream generated by performing an image coding method, wherein the image coding method includes: determining reference samples of a current block; performing filtering on the reference samples based on features of a region including the reference samples; and performing intra-frame prediction by using the filtered reference samples.
[0028] Beneficial effects
[0029] This invention provides various filtering methods to be performed at each step during image encoding / decoding to improve image encoding / decoding efficiency.
[0030] This invention can improve prediction efficiency by generating a predicted image using reference samples that are close to the original image.
[0031] This invention can improve the ringing effect in the target boundary region of an image and the contour effect that occurs when performing orientation prediction.
[0032] According to the present invention, the efficiency of image encoding and decoding can be improved.
[0033] According to the present invention, the computational complexity of image encoders and decoders can be reduced. Attached Figure Description
[0034] Figure 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.
[0035] Figure 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment of the present invention.
[0036] Figure 3 It is a schematic diagram illustrating the partitioning structure of an image during encoding and decoding.
[0037] Figure 4 This is a diagram illustrating intra-frame prediction processing.
[0038] Figure 5 This is a diagram illustrating inter-frame prediction processing.
[0039] Figure 6 This is a diagram illustrating the transformation and quantization processes.
[0040] Figure 7 This is a diagram illustrating an example embodiment of configuring reference samples using multiple reconstructed sample lines.
[0041] Figure 8 This is a diagram illustrating the image features of a region including reference samples according to an embodiment of the present invention.
[0042] Figure 9This is a diagram illustrating a method for deriving the uniformity of an image according to an embodiment of the present invention.
[0043] Figure 10 This is a diagram illustrating a method for deriving the uniformity of an image using gradients according to an embodiment of the present invention.
[0044] Figure 11 This is a diagram illustrating the direction of applied filtering according to an embodiment of the present invention.
[0045] Figure 12 This is a diagram illustrating a pixel region used for filtering according to an embodiment of the present invention.
[0046] Figure 13 This is a diagram illustrating an example of an embodiment where filtering is applied in 1 / 4 (or quarter pel) units.
[0047] Figure 14 This is a diagram illustrating an example of an embodiment where filtering is performed when a portion of the region used for filtering is outside a boundary.
[0048] Figure 15 This is a diagram illustrating a 1D filter according to an embodiment of the present invention.
[0049] Figure 16 This is a diagram illustrating a 2D filter according to an embodiment of the present invention.
[0050] Figure 17 This is a flowchart illustrating an image decoding method according to an embodiment of the present invention.
[0051] Figure 18 This is a flowchart illustrating an image decoding method according to another embodiment of the present invention. Detailed Implementation
[0052] Various modifications can be made to this invention, and various embodiments of the invention exist, wherein examples of these various embodiments will now be provided with reference to the accompanying drawings, and examples of these various embodiments will be described in detail. However, the invention is not limited thereto, although exemplary embodiments may be interpreted as including all modifications, equivalents, or substitutions within the technical concept and scope of the invention. Similar reference numerals indicate the same or similar functions in various respects. In the drawings, the shapes and sizes of elements may be exaggerated for clarity. In the following detailed description of the invention, reference is made to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice this disclosure. It should be understood that the various embodiments of this disclosure, though different, are not necessarily mutually exclusive. For example, specific features, structures, and characteristics described herein in connection with one embodiment may be implemented in other embodiments without departing from the spirit and scope of this disclosure. Furthermore, it should be understood that the position or arrangement of various elements within each disclosed embodiment may be modified without departing from the spirit and scope of this disclosure. Therefore, the following detailed description should not be construed in a limiting sense, and the scope of this disclosure is defined only by the appended claims (and, where appropriate, the full scope of the equivalents claimed in the claims).
[0053] The terms "first," "second," etc., used in this specification may be used to describe various components, but these components are not to be construed as being limited to these terms. These terms are only used to distinguish one component from others. For example, without departing from the scope of the invention, a "first" component may be referred to as a "second" component, and a "second" component may similarly be referred to as a "first" component. The term "and / or" includes a combination of multiple items or any one of multiple items.
[0054] It will be understood that, in this specification, when an element is referred to only as "connected to" or "joined to" another element rather than "directly connected to" or "directly joined to" another element, the element may be "directly connected to" or "directly joined to" the other element, or may be connected to or joined to the other element if there are other elements between the element and the other element. Conversely, it should be understood that when an element is referred to as "directly joined" or "directly connected" to another element, there are no intermediate elements.
[0055] Furthermore, the components shown in the embodiments of the present invention are illustrated independently to present distinct functionalities. Therefore, this does not imply that each component is composed as a separate hardware or software unit. In other words, for convenience, each component includes every one of the enumerated components. Thus, at least two components in each component may be combined to form a single component, or a single component may be divided into multiple components for performing each function. Embodiments where each component is combined and embodiments where a component is divided are also included within the scope of the invention without departing from its spirit.
[0056] The terminology used in this specification is for describing particular embodiments only and is not intended to limit the invention. Expressions used in the singular include plural expressions unless they have a distinct meaning in the context. In this specification, it will be understood that terms such as “comprising,” “having,” etc., are intended to indicate the presence of the features, quantities, steps, actions, elements, components, or combinations thereof disclosed in the specification, and are not intended to exclude the possibility that one or more other features, quantities, steps, actions, elements, components, or combinations thereof may be present or added. In other words, when a particular element is referred to as “comprising,” elements other than the corresponding element are not excluded; rather, additional elements may be included in embodiments of the invention or within the scope of the invention.
[0057] Furthermore, some components may not be essential for performing the necessary functions of the invention, but rather optional components that merely enhance its performance. The invention can be implemented by including only the essential components necessary for carrying out the invention, excluding components used to enhance performance. Structures that include only the essential components and exclude optional components used solely for enhancing performance are also included within the scope of the invention.
[0058] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In describing exemplary embodiments of the invention, well-known functions or structures will not be described in detail, as they would unnecessarily obscure the understanding of the invention. The same constituent elements in the drawings are denoted by the same reference numerals, and repeated descriptions of the same elements will be omitted.
[0059] In the following text, an image may refer to a frame that constitutes a video, or it may refer to the video itself. For example, "encoding or decoding an image or both" may refer to "encoding or decoding a moving image or both," and may also refer to "encoding or decoding an image within an image of a moving image or both."
[0060] In the following text, the terms "moving images" and "video" may be used to mean the same thing and may be used interchangeably.
[0061] In the following text, the target image can be an encoded target image that serves as an encoding target and / or a decoded target image that serves as a decoding target. Furthermore, the target image can be an input image input to an encoding device and an input image input to a decoding device. Here, the target image may have the same meaning as the current image.
[0062] In the following text, the terms “image,” “picture,” “frame,” and “screen” may be used as having the same meaning and interchangeably with each other.
[0063] In the following text, a target block can be an encoded target block that serves as the encoding target and / or a decoded target block that serves as the decoding target. Furthermore, a target block can be the current block that serves as the target of the current encoding and / or decoding. For example, the terms "target block" and "current block" can be used to mean the same thing and are interchangeable.
[0064] In the following text, the terms “block” and “unit” may be used to mean the same thing and are interchangeable. Alternatively, “block” may refer to a specific unit.
[0065] In the following text, the terms “region” and “fragment” are used interchangeably.
[0066] In the following text, a specific signal can be a signal representing a specific block. For example, the original signal can be a signal representing the target block. The prediction signal can be a signal representing the prediction block. The residual signal can be a signal representing the residual block.
[0067] In this embodiment, each of the specific information, data, flags, indexes, elements, and attributes may have a value. A value equal to "0" for information, data, flags, indexes, elements, and attributes may represent logical false or a first predefined value. In other words, the values "0", false, logical false, and the first predefined value can be interchanged with each other. A value equal to "1" for information, data, flags, indexes, elements, and attributes may represent logical true or a second predefined value. In other words, the values "1", true, logical true, and the second predefined value can be interchanged with each other.
[0068] When variables i or j are used to represent columns, rows, or indices, the value of i can be an integer equal to or greater than 0, or an integer equal to or greater than 1. That is, columns, rows, indices, etc., can be counted starting from 0, or they can be counted starting from 1.
[0069] Terminology Description
[0070] Encoder: Represents the device that performs encoding. In other words, it refers to the encoding device.
[0071] Decoder: Represents the device that performs decoding. In other words, it refers to the decoding device.
[0072] A block is an M×N sample array. Here, M and N can represent positive integers, and a block can represent a two-dimensional sample array. A block can refer to a unit. The current block can represent a coding target block that becomes the target during encoding, or a decoding target block that becomes the target during decoding. Furthermore, the current block can be at least one of a coding block, a prediction block, a residual block, and a transform block.
[0073] Samples are the basic units that make up a block. Based on the bit depth (Bd), samples can be represented as numbers from 0 to 2. Bd The value is -1. In this invention, a sample point can be used to represent a pixel. That is, a sample point, a pel, and a pixel can have the same meaning.
[0074] Unit: Refers to an encoding and decoding unit. When encoding and decoding an image, a unit can be a region generated by partitioning a single image. Furthermore, a unit can represent a sub-partitioning unit when a single image is partitioned into sub-partitioning units during encoding or decoding. That is, an image can be partitioned into multiple units. When encoding and decoding an image, predetermined processing can be performed for each unit. A single unit can be partitioned into sub-units smaller than the unit's size. Depending on the function, a unit can represent a block, macroblock, coding tree unit, coding tree block, coding unit, coding block, prediction unit, prediction block, residual unit, residual block, transform unit, transform block, etc. Furthermore, to distinguish a unit from a block, a unit can include a luma component block, a chroma component block associated with the luma component block, and syntax elements for each chroma component block. Units can have various sizes and shapes; specifically, the shape of a unit can be a two-dimensional geometric figure, such as a square, rectangle, trapezoid, triangle, pentagon, etc. In addition, the cell information may include at least one of the following: cell type indicating coding cell, prediction cell, transform cell, etc., cell size, cell depth, and the order of encoding and decoding of the cell.
[0075] A coding tree unit is a single coding tree block configured with the luminance component Y and two coding tree blocks associated with the chrominance components Cb and Cr. Furthermore, a coding tree unit can represent a block and the syntax elements of each block. Each coding tree unit can be partitioned using at least one of quadtree partitioning, binary tree partitioning, and ternary tree partitioning methods to configure lower-level units such as coding units, prediction units, transform units, etc. A coding tree unit can be used as a term to specify a sample block that becomes a processing unit when encoding / decoding an image as an input image. Here, a quadtree can represent a quaternion tree.
[0076] Encoding block: Can be used as a term to specify any one of the Y encoding block, Cb encoding block, and Cr encoding block.
[0077] Neighboring blocks: These can represent blocks adjacent to the current block. A block adjacent to the current block can be a block that touches the boundary of the current block, or a block located within a predetermined distance from the current block. A neighboring block can also represent a block adjacent to a vertex of the current block. Here, a block adjacent to a vertex of the current block can be a block that is vertically adjacent to a block horizontally adjacent to the current block, or a block that is horizontally adjacent to a block vertically adjacent to the current block.
[0078] Reconstructed neighboring blocks: These can represent neighboring blocks that are adjacent to the current block and have already been spatially / temporally encoded or decoded. Here, reconstructed neighboring blocks can represent reconstructed neighboring units. Reconstructed spatial neighboring blocks can be blocks within the current frame that have already been reconstructed through encoding or decoding, or both. Reconstructed temporally neighboring blocks are blocks within a reference image that are located at the position corresponding to the current block in the current frame, or neighboring blocks of said block.
[0079] Cell depth: Represents the degree of cell partitioning. In a tree structure, the highest node (root node) corresponds to the first cell that is not partitioned. Furthermore, the highest node can have the minimum depth value. In this case, the depth of the highest node can be level 0. A node with a depth of level 1 represents a cell generated by partitioning the first cell once. A node with a depth of level 2 represents a cell generated by partitioning the first cell twice. A node with a depth of level n represents a cell generated by partitioning the first cell n times. Leaf nodes can be the lowest-level nodes and cannot be further partitioned. The depth of a leaf node can be the highest-level. For example, the predefined value for the highest-level can be 3. The root node can have the lowest depth, and the leaf nodes can have the deepest depth. Additionally, when cells are represented as a tree structure, the level at which the cell exists can represent the cell depth.
[0080] Bitstream: A bitstream that can represent encoded image information.
[0081] Parameter set: Corresponds to the header information in the configuration within the bitstream. At least one of the video parameter set, sequence parameter set, frame parameter set, and adaptive parameter set may be included in the parameter set. In addition, the parameter set may include slice header and tile header information.
[0082] Explanation: This could mean determining the value of a syntax element by performing entropy decoding, or it could mean entropy decoding itself.
[0083] Symbols: can represent at least one of the syntax elements, encoding parameters, and transform coefficient values of the encoding / decoding target unit. Additionally, symbols can represent entropy encoding targets or entropy decoding results.
[0084] Prediction mode: This can be information indicating the mode that is encoded / decoded using intra-frame prediction or the mode that is encoded / decoded using inter-frame prediction.
[0085] Prediction Unit: A basic unit that can be represented when performing prediction (such as inter-frame prediction, intra-frame prediction, inter-frame compensation, intra-frame compensation, and motion compensation). A single prediction unit can be partitioned into multiple partitions with smaller sizes, or it can be partitioned into multiple lower-level prediction units. Multiple partitions can be the basic units when performing prediction or compensation. Partitions generated by dividing prediction units can also be prediction units.
[0086] Prediction cell partitioning: can represent the shape obtained by partitioning prediction cells.
[0087] Reference frame list: This can refer to a list of one or more reference frames used for inter-frame prediction or motion compensation. Several types of available reference frame lists exist, including LC (list combination), L0 (list 0), L1 (list 1), L2 (list 2), and L3 (list 3).
[0088] Inter-frame prediction indicator: This can indicate the direction of inter-frame prediction (one-way prediction, two-way prediction, etc.) for the current block. Optionally, the inter-frame prediction indicator can indicate the number of reference frames used to generate the prediction blocks for the current block. Optionally, the inter-frame prediction indicator can indicate the number of prediction blocks used when performing inter-frame prediction or motion compensation on the current block.
[0089] Prediction list utilization flag: Indicates whether at least one reference frame from a specific reference frame list is used to generate the prediction block. The prediction list utilization flag can be used to derive the inter-frame prediction indicator, and conversely, the inter-frame prediction indicator can be used to derive the prediction list utilization flag. For example, when the prediction list utilization flag has a first value of zero (0), it indicates that reference frames from the reference frame list are not used to generate the prediction block. On the other hand, when the prediction list utilization flag has a second value of one (1), it indicates that the reference frame list is used to generate the prediction block.
[0090] Reference screen index: can refer to the index of a specific reference screen in the reference screen list.
[0091] Reference frame: This can refer to a reference frame referenced by a specific block for the purpose of performing inter-frame prediction or motion compensation on that specific block. Optionally, a reference frame can be a frame that includes a reference block referenced by the current block for inter-frame prediction or motion compensation. In the following text, the terms "reference frame" and "reference image" have the same meaning and are interchangeable.
[0092] Motion vector: This can be a two-dimensional vector used for inter-frame prediction or motion compensation. The motion vector represents the offset between the encoded / decoded target block and the reference block. For example, (mvX, mvY) can represent a motion vector. Here, mvX can represent the horizontal component, and mvY can represent the vertical component.
[0093] Search range: This can be a two-dimensional region searched during inter-frame prediction to retrieve motion vectors. For example, the size of the search range can be M×N. Here, M and N are both integers.
[0094] Motion vector candidate: This can refer to a prediction candidate block or the motion vector of a prediction candidate block when predicting motion vectors. Furthermore, motion vector candidates can be included in a motion vector candidate list.
[0095] Candidate list of motion vectors: can represent a list consisting of one or more candidates of motion vectors.
[0096] Motion vector candidate index: This can represent an indicator that points to a motion vector candidate in the motion vector candidate list. Optionally, the motion vector candidate index can be an index of the motion vector predictor.
[0097] Motion information: may represent information including at least one of the following: motion vector, reference frame index, inter-frame prediction indicator, prediction list utilization flag, reference frame list information, reference frame, motion vector candidate, motion vector candidate index, merge candidate and merge index.
[0098] Merge candidate list: can represent a list consisting of one or more merge candidates.
[0099] Merge Candidates: These can represent spatial merge candidates, temporal merge candidates, combined merge candidates, combined double prediction merge candidates, or zero merge candidates. Merge candidates may include motion information such as inter-frame prediction indicators, reference frame indices for each list, motion vectors, prediction list utilization flags, and inter-frame prediction indicators.
[0100] Merge Index: This can represent an indicator that points to a merge candidate in the merge candidate list. Optionally, the merge index can indicate a block among the reconstructed blocks that are spatially / temporally adjacent to the current block and from which a merge candidate has been derived. Optionally, the merge index can indicate at least one piece of motion information for the merge candidate.
[0101] Transform unit: This can represent the basic unit used when performing encoding / decoding (such as transform, inverse transform, quantization, dequantization, transform coefficient encoding / decoding) on the residual signal. A single transform unit can be partitioned into multiple lower-level transform units with smaller sizes. Here, the transform / inverse transform may include at least one of a first transform / first inverse transform and a second transform / second inverse transform.
[0102] Scaling: This refers to the process of multiplying the quantization level by a factor. Transform coefficients can be generated by scaling the quantization level. Scaling can also be called inverse quantization.
[0103] Quantization parameters: These represent values used when transform coefficients are used to generate quantization levels during quantization. Quantization parameters can also represent values used when transform coefficients are generated during dequantization by scaling the quantization levels. Quantization parameters can be values mapped to the quantization step size.
[0104] Incremental quantization parameter: can represent the difference between the predicted quantization parameter and the quantization parameter of the encoding / decoding target unit.
[0105] Scan: This can refer to a method of sorting coefficients within a cell, block, or matrix. For example, changing a two-dimensional matrix of coefficients into a one-dimensional matrix can be called a scan, and changing a one-dimensional matrix of coefficients into a two-dimensional matrix can be called a scan or inverse scan.
[0106] Transform coefficients: These represent the coefficient values generated after a transform is performed in the encoder. Transform coefficients can also represent the coefficient values generated after at least one of entropy decoding and dequantization is performed in the decoder. The quantization level obtained by quantizing the transform coefficients or residual signal, or the quantized transform coefficient level, can also fall within the meaning of transform coefficients.
[0107] Quantization level: This can represent the value generated in the encoder by quantizing the transform coefficients or residual signal. Optionally, the quantization level can represent the value of the dequantization target that will undergo dequantization in the decoder. Similarly, the transform coefficient level, as a result of transform and quantization, can also fall within the meaning of quantization level.
[0108] Non-zero transform coefficients: can represent transform coefficients with values other than zero, or transform coefficient levels or quantization levels with values other than zero.
[0109] Quantization matrix: A matrix used in quantization or dequantization processes to improve subjective or objective image quality. The quantization matrix can also be referred to as a scaling list.
[0110] Quantization matrix coefficients: These represent each element within the quantization matrix. Quantization matrix coefficients can also be called matrix coefficients.
[0111] Default matrix: can represent a predefined quantization matrix in the encoder or decoder.
[0112] Non-default matrix: can represent a quantization matrix that is not predefined in the encoder or decoder but is sent by the user via signal.
[0113] Statistical value: For at least one of the following variables, coding parameters, constant values, etc., that have a computable specific value, a statistical value can be one or more of the following: mean, weighted average, weighted sum, minimum, maximum, most frequent value, median, interpolation value.
[0114] Figure 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.
[0115] Encoding device 100 may be an encoder, a video encoding device, or an image encoding device. The video may include at least one image. Encoding device 100 may encode at least one image sequentially.
[0116] Reference Figure 1 The encoding device 100 may include a motion prediction unit 111, a motion compensation unit 112, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.
[0117] Encoding device 100 can encode the input image using intra-frame mode, inter-frame mode, or both. Furthermore, encoding device 100 can generate a bitstream including encoding information by encoding the input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or streamed via a wired / wireless transmission medium. When intra-frame mode is used as the prediction mode, switcher 115 can switch to intra-frame mode. Optionally, when inter-frame mode is used as the prediction mode, switcher 115 can switch to inter-frame mode. Here, intra-frame mode can refer to intra-frame prediction mode, and inter-frame mode can refer to inter-frame prediction mode. Encoding device 100 can generate prediction blocks for input blocks of the input image. Furthermore, encoding device 100 can encode residual blocks using the residual between the input block and the prediction block after generating the prediction blocks. The input image can be referred to as the current image as the current encoding target. The input block can be referred to as the current block as the current encoding target, or as the encoding target block.
[0118] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use samples from blocks that have been encoded / decoded and are adjacent to the current block as reference samples. The intra-frame prediction unit 120 can perform spatial prediction on the current block using the reference samples, or generate prediction samples for the input block by performing spatial prediction. Here, intra-frame prediction can refer to prediction within a frame.
[0119] When the prediction mode is inter-frame mode, the motion prediction unit 111 can retrieve the region that best matches the input block from the reference image during motion prediction and derive the motion vector using the retrieved region. In this case, the search region can be used as the region. The reference image can be stored in the reference frame buffer 190. Here, the reference image can be stored in the reference frame buffer 190 when encoding / decoding the reference image is performed.
[0120] The motion compensation unit 112 can generate a prediction block by performing motion compensation on the current block using motion vectors. Here, inter-frame prediction can refer to prediction or motion compensation between frames.
[0121] When the value of the motion vector is not an integer, motion prediction unit 111 and motion compensation unit 112 can generate prediction blocks by applying an interpolation filter to a portion of the reference frame. To perform inter-frame prediction or motion compensation on the coding unit, it can be determined which of the following modes—skip mode, merge mode, Advanced Motion Vector Prediction (AMVP) mode, and current frame reference mode—is used for motion prediction and motion compensation on the prediction unit included in the corresponding coding unit. Then, depending on the determined mode, inter-frame prediction or motion compensation can be performed differently.
[0122] Subtractor 125 can generate a residual block by using the residual between the input block and the prediction block. The residual block can be referred to as a residual signal. The residual signal can represent the difference between the original signal and the prediction signal. Furthermore, the residual signal can be a signal generated by transforming or quantizing, or transforming and quantizing, the difference between the original signal and the prediction signal. The residual block can be the residual signal of a block cell.
[0123] Transform unit 130 can generate transform coefficients by performing a transform on the residual block and output the generated transform coefficients. Here, the transform coefficients can be coefficient values generated by performing a transform on the residual block. When a transform skip mode is applied, transform unit 130 can skip the transform on the residual block.
[0124] The level of quantization can be generated by applying quantization to the transform coefficients or to the residual signal. In the following examples, the level of quantization may also be referred to as the transform coefficients.
[0125] The quantization unit 140 can generate a quantization level by quantizing the transform coefficients or residual signal according to parameters, and output the generated quantization level. Here, the quantization unit 140 can quantize the transform coefficients using a quantization matrix.
[0126] Entropy coding unit 150 can generate a bitstream by performing entropy coding on the values calculated by quantization unit 140 according to a probability distribution or on the coding parameter values calculated during encoding, and output the generated bitstream. Entropy coding unit 150 can perform entropy coding on sample information of the image and information used for decoding the image. For example, the information used for decoding the image may include syntax elements.
[0127] When entropy coding is applied, symbols are represented such that fewer bits are allocated to symbols with high generation probability and more bits are allocated to symbols with low generation probability, thus reducing the size of the bitstream used to encode the symbols. The entropy coding unit 150 can use coding methods for entropy coding such as Exponential Golomb, Context Adaptive Variable Length Coding (CAVLC), and Context Adaptive Binary Arithmetic Coding (CABAC). For example, the entropy coding unit 150 can perform entropy coding by using a variable-length code (VLC) table. Furthermore, the entropy coding unit 150 can derive a binarization method for the target symbol and a probability model for the target symbol / bits, and perform arithmetic coding by using the derived binarization method and context model.
[0128] In order to encode the transform coefficient levels (quantization levels), the entropy coding unit 150 can change the coefficients in two-dimensional block form into one-dimensional vector form by using a transform coefficient scanning method.
[0129] Encoding parameters may include information such as syntax elements (flags, indexes, etc.) encoded in the encoder and signaled to the decoder, as well as information derived during encoding or decoding. Encoding parameters can represent the information required when encoding or decoding an image. For example, at least one value or combination of the following may be included in the encoding parameters: cell / block size, cell / block depth, cell / block partitioning information, cell / block shape, cell / block partitioning structure, whether quadtree partitioning is performed, whether binary tree partitioning is performed, binary tree partitioning direction (horizontal or vertical), binary tree partitioning type (symmetric or asymmetric), whether the current encoding unit is partitioned via ternary tree partitioning, the direction of ternary tree partitioning (horizontal or vertical), the type of ternary tree partitioning (symmetric or asymmetric), whether the current encoding unit is partitioned via multi-type tree partitioning, and the direction of multi-type tree partitioning. This includes the following: orientation (horizontal or vertical), type of multi-type tree partitions (symmetric or asymmetric), tree structure of multi-type tree partitions (binary or ternary), prediction mode (intra-frame prediction or inter-frame prediction), intra-frame prediction mode / direction for luma, intra-frame prediction mode / direction for chroma, intra-frame partition information, inter-frame partition information, coded block partition flag, prediction block partition flag, transform block partition flag, reference sample filtering method, reference sample filter taps, reference sample filter coefficients, prediction block filtering method, prediction block filter taps, prediction block filter coefficients, prediction block boundary filtering method, prediction block boundary filter taps, prediction block boundary filter coefficients, and intra-frame prediction mode. Inter-frame prediction mode, motion information, motion vector, motion vector difference, reference frame index, inter-frame prediction angle, inter-frame prediction indicator, prediction list utilization flag, reference frame list, reference frame, motion vector predictor index, motion vector predictor candidate, motion vector candidate list, whether to use merge mode, merge index, merge candidate, merge candidate list, whether to use skip mode, interpolation filter type, interpolation filter taps, interpolation filter coefficients, motion vector magnitude, motion vector representation accuracy, transform type, transform size, information on whether the first (first) transform is used, information on whether the second transform is used, first transform index, second transform index Information on the presence of residual signals, code block style, code block flag (CBF), quantization parameters, quantization parameter residuals, quantization matrix, whether an intra-loop filter is applied, intra-loop filter coefficients, intra-loop filter taps, intra-loop filter shape / form, whether a deblocking filter is applied, deblocking filter coefficients, deblocking filter taps, deblocking filter strength, deblocking filter shape / form, whether adaptive sample offset is applied, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, whether an adaptive loop filter is applied, adaptive loop filter coefficients, adaptive loop filter taps, adaptive loop filter shape / form.Binarization / debinarization method, context model determination method, context model update method, whether to execute normal mode, whether to execute bypass mode, context binary bits, bypass binary bits, valid coefficient flag, last valid coefficient flag, encoding flag for the unit of the coefficient group, position of the last valid coefficient, flag indicating whether the coefficient value is greater than 1, flag indicating whether the coefficient value is greater than 2, flag indicating whether the coefficient value is greater than 3, information about the remaining coefficient values, symbol information, reconstructed luminance samples, reconstructed chrominance samples, residual luminance samples, residual chrominance samples, luminance transformation coefficient, chrominance transformation coefficient, quantized luminance level, quantized chrominance level, transformation coefficient level scanning method. Information includes: the size and shape of the motion vector search region on the decoder side; the number of motion vector searches on the decoder side; information about the CTU size; information about the minimum block size; information about the maximum block size; information about the maximum block depth; information about the minimum block depth; image display / output order; strip identification information; strip type; strip partition information; parallel block identification information; parallel block type; parallel block partition information; picture type; bit depth of input samples; bit depth of reconstructed samples; bit depth of residual samples; bit depth of transform coefficients; bit depth of quantization levels; and information about the luminance signal or the chrominance signal.
[0130] Here, sending a flag or index with a signal can represent the encoder entropy encoding the corresponding flag or index and including it in the bitstream, and can also represent the decoder entropy decoding the corresponding flag or index from the bitstream.
[0131] When the encoding device 100 performs encoding via inter-frame prediction, the encoded current image can be used as a reference image for another image to be processed subsequently. Therefore, the encoding device 100 can reconstruct or decode the encoded current image, or store the reconstructed or decoded image as a reference image in the reference frame buffer 190.
[0132] The quantization level can be dequantized in dequantization unit 160 or inverse transformed in inverse transform unit 170. The coefficients that have undergone dequantization or inverse transform, or both, can be added to the prediction block by adder 175. A reconstructed block can be generated by adding the coefficients that have undergone dequantization or inverse transform, or both, to the prediction block. Here, the coefficients that have undergone dequantization or inverse transform, or both, can represent coefficients for which at least one of dequantization and inverse transform has been performed, and can represent the reconstructed residual block.
[0133] The reconstructed block can be passed through filter unit 180. Filter unit 180 can apply at least one of deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) to the reconstructed sample, reconstructed block, or reconstructed image. Filter unit 180 may be referred to as an in-loop filter.
[0134] Deblocking filters remove block distortion generated at the boundaries between blocks. To determine whether to apply a deblocking filter, the number of samples included in several rows or columns within the block can be used. When a deblocking filter is applied to a block, another filter can be applied based on the desired deblocking intensity.
[0135] To compensate for coding errors, a suitable offset value can be added to the sample value using a sample-adaptive offset. The sample-adaptive offset corrects the offset between the deblocked image and the original image on a sample-by-sample basis. This can be achieved by considering edge information about each sample point when applying the offset, or by dividing the image's samples into a predetermined number of regions, determining the regions where the offset will be applied, and then applying the offset to those regions.
[0136] Adaptive loop filters (ALFs) can perform filtering based on a comparison between the filtered reconstructed image and the original image. Samples included in the image can be partitioned into predetermined groups, the filter to be applied to each group can be determined, and differential filtering can be performed on each group. Information regarding whether to apply an ALF can be transmitted via a coding unit (CU), and the form and coefficients of the ALF to be applied to each block can vary.
[0137] The reconstructed blocks or reconstructed image that have passed through filter unit 180 can be stored in reference frame buffer 190. The reconstructed blocks processed by filter unit 180 can be a portion of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks processed by filter unit 180. The stored reference image can be used later in inter-frame prediction or motion compensation.
[0138] Figure 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment and to which the present invention is applied.
[0139] Decoding device 200 can be a decoder, video decoding device, or image decoding device.
[0140] Reference Figure 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 225, a filter unit 260, and a reference frame buffer 270.
[0141] Decoding device 200 can receive bitstreams output from encoding device 100. Decoding device 200 can receive bitstreams stored on a computer-readable recording medium, or bitstreams streamed via wired / wireless transmission media. Decoding device 200 can decode the bitstreams using intra-frame mode or inter-frame mode. Furthermore, decoding device 200 can generate and output reconstructed or decoded images generated through decoding.
[0142] When the prediction mode used during decoding is intra-frame mode, the switcher can be switched to intra-frame mode. Optionally, when the prediction mode used during decoding is inter-frame mode, the switcher can be switched to inter-frame mode.
[0143] Decoding device 200 obtains a reconstructed residual block and generates a prediction block by decoding the input bitstream. Once the reconstructed residual block and prediction block are obtained, decoding device 200 generates a reconstructed block that becomes the decoding target by adding the reconstructed residual block and the prediction block. The decoding target block can be referred to as the current block.
[0144] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream according to a probability distribution. The generated symbols may include symbols in quantized hierarchical form. Here, the entropy decoding method can be the inverse process of the entropy encoding method described above.
[0145] In order to decode the transform coefficient levels (quantization levels), the entropy decoding unit 210 can change the coefficients in unidirectional vector form into two-dimensional block form by using a transform coefficient scanning method.
[0146] The quantization level can be dequantized in the dequantization unit 220, or the quantization level can be inversely transformed in the inverse transform unit 230. The quantization level can be the result of dequantization, inverse transform, or both, and can be generated as a reconstructed residual block. Here, the dequantization unit 220 can apply the quantization matrix to the quantization level.
[0147] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the current block, wherein the spatial prediction uses sample values of blocks that are adjacent to the target block and have already been decoded.
[0148] When using inter-frame mode, motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block, wherein the motion compensation uses motion vectors and a reference image stored in reference frame buffer 270.
[0149] Adder 225 generates a reconstructed block by adding the reconstructed residual block to the prediction block. Filter unit 260 can apply at least one of a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the reconstructed block or reconstructed image. Filter unit 260 can output a reconstructed image. The reconstructed block or reconstructed image can be stored in a reference frame buffer 270 and used when performing inter-frame prediction. The reconstructed block processed by filter unit 260 can be a portion of the reference image. That is, the reference image is a reconstructed image composed of the reconstructed blocks processed by filter unit 260. The stored reference image can be used later in inter-frame prediction or motion compensation.
[0150] Figure 3 It is a schematic diagram illustrating the partitioning structure of an image when it is encoded and decoded. Figure 3 An example of partitioning a single cell into multiple lower-level cells is illustrated.
[0151] To effectively partition an image, coding units (CUs) are used during encoding and decoding. A coding unit can serve as the basic unit when encoding / decoding an image. Furthermore, a coding unit can be used to distinguish between intra-frame prediction modes and inter-frame prediction modes during image encoding / decoding. A coding unit can be the basic unit used for prediction, transform, quantization, inverse transform, inverse quantization, or encoding / decoding processing of transform coefficients.
[0152] Reference Figure 3 Image 300 is partitioned sequentially according to the Largest Coding Unit (LCU), and the LCU unit is determined as the partition structure. Here, LCU can be used with the same meaning as Coding Tree Unit (CTU). Unit partitioning can represent partitioning of the block associated with that unit. The block partitioning information may include information about the unit depth. The depth information may represent the number or degree to which the unit is partitioned, or both the number and degree to which the unit is partitioned. A single unit can be partitioned into multiple lower-level units hierarchically associated with the depth information based on a tree structure. In other words, the unit and the lower-level units generated by partitioning that unit may correspond to a node and the child nodes of that node, respectively. Each of the partitioned lower-level units may have depth information. The depth information may be information representing the size of the CU and may be stored in each CU. The unit depth represents the number and / or degree associated with partitioning the unit. Therefore, the partitioning information of the lower-level units may include information about the size of the lower-level units.
[0153] The partitioning structure represents the distribution of coding units (CUs) within the LCU 310. This distribution can be determined by whether a single CU is partitioned into multiple CUs (positive integers equal to or greater than 2, including 2, 4, 8, 16, etc.). The horizontal and vertical dimensions of the CUs generated by partitioning can be half the horizontal and vertical dimensions of the CUs before partitioning, respectively, or they can have dimensions smaller than the horizontal and vertical dimensions before partitioning, depending on the number of partitions. CUs can be recursively partitioned into multiple CUs. Through recursive partitioning, at least one of the height and width of the CU after partitioning can be reduced compared to at least one of the height and width of the CU before partitioning. CU partitioning can be performed recursively until a predefined depth or a predefined size is reached. For example, the depth of the LCU can be 0, and the depth of the minimum coding unit (SCU) can be a predefined maximum depth. Here, as mentioned above, the LCU can be the coding unit with the maximum coding unit size, and the SCU can be the coding unit with the minimum coding unit size. Partitioning begins at the LCU 310, and the CU depth increases by 1 when the horizontal or vertical dimension of the CU, or both the horizontal and vertical dimensions, decrease through partitioning. For example, for each depth, the size of an unpartitioned CU can be 2N×2N. Furthermore, in the case of partitioned CUs, a CU of size 2N×2N can be partitioned into four CUs of size N×N. As the depth increases by 1, the size of N can be halved.
[0154] Furthermore, partition information of a CU can be used to indicate whether a CU is partitioned. Partition information can be 1 bit. All CUs except SCUs can include partition information. For example, when the partition information value is 1, the CU may not be partitioned; when the partition information value is 2, the CU may be partitioned.
[0155] Reference Figure 3 An LCU with depth 0 can be a 64×64 block. 0 can be the minimum depth. An SCU with depth 3 can be an 8×8 block. 3 can be the maximum depth. CUs with 32×32 blocks and 16×16 blocks can be represented as depth 1 and depth 2, respectively.
[0156] For example, when a single coding unit is partitioned into four coding units, the horizontal and vertical dimensions of the four partitioned coding units can be half the horizontal and vertical dimensions of the CU before partitioning. In one embodiment, when a 32×32 coding unit is partitioned into four coding units, each of the four partitioned coding units can have a size of 16×16. When a single coding unit is partitioned into four coding units, the coding unit can be said to be partitioned into a quadtree form.
[0157] For example, when a coding unit is partitioned into two sub-coding units, the horizontal or vertical dimension (width or height) of each of the two sub-coding units can be half the horizontal or vertical dimension of the original coding unit. For example, when a coding unit of size 32×32 is vertically partitioned into two sub-coding units, each of the two sub-coding units can have a size of 16×32. For example, when a coding unit of size 8×32 is horizontally partitioned into two sub-coding units, each of the two sub-coding units can have a size of 8×16. When a coding unit is partitioned into two sub-coding units, it can be said that the coding unit is binary partitioned, or partitioned according to a binary tree partitioning structure.
[0158] For example, when a coding unit is divided into three sub-coding units, the horizontal or vertical dimensions of the coding unit can be divided in a 1:2:1 ratio, resulting in three sub-coding units with a horizontal or vertical dimension ratio of 1:2:1. For instance, when a 16×32 coding unit is horizontally divided into three sub-coding units, these three sub-coding units, in order from the top to the bottom, can have dimensions of 16×8, 16×16, and 16×8, respectively. Similarly, when a 32×32 coding unit is vertically divided into three sub-coding units, these three sub-coding units, in order from the left to the right, can have dimensions of 8×32, 16×32, and 8×32, respectively. When a coding unit is divided into three sub-coding units, it can be said that the coding unit is tri-partitioned or partitioned according to a ternary tree partitioning structure.
[0159] exist Figure 3 In the example, the coding tree unit (CTU) 320 is an example of a CTU in which quadtree partitioning, binary tree partitioning, and ternary tree partitioning structures are all applied.
[0160] As described above, to partition the CTU, at least one of a quadtree partitioning structure, a binary tree partitioning structure, and a ternary tree partitioning structure can be applied. Various tree partitioning structures can be applied sequentially to the CTU according to a predetermined priority order. For example, a quadtree partitioning structure can be preferentially applied to the CTU. Encoding units that cannot be further partitioned using a quadtree partitioning structure can correspond to leaf nodes of a quadtree. Encoding units corresponding to leaf nodes of a quadtree can be used as root nodes of binary and / or ternary tree partitioning structures. That is, encoding units corresponding to leaf nodes of a quadtree can be further partitioned according to a binary or ternary tree partitioning structure, or they can be left unpartitioned. Therefore, by preventing the encoded blocks obtained from binary or ternary tree partitioning of encoding units corresponding to leaf nodes of a quadtree from undergoing further quadtree partitioning, block partitioning operations and / or the operation of signaling partitioning information can be effectively performed.
[0161] The fact that a coding unit corresponding to a node in a quadtree is partitioned can be signaled using four-partition information. Four-partition information with a first value (e.g., "1") indicates that the current coding unit is partitioned according to the quadtree partitioning structure. Four-partition information with a second value (e.g., "0") indicates that the current coding unit is not partitioned according to the quadtree partitioning structure. The four-partition information can be a flag with a predetermined length (e.g., one bit).
[0162] There may be no priority between binary tree partitions and ternary tree partitions. That is, the coding unit corresponding to the leaf node of the quadtree can further undergo any partition in either binary tree or ternary tree partitions. Furthermore, the coding unit generated by binary tree partitions or ternary tree partitions may undergo further binary tree partitions or further ternary tree partitions, or it may not be further partitioned.
[0163] A tree structure in which there is no priority between binary tree partitions and ternary tree partitions is called a multi-type tree structure. The coding unit corresponding to the leaf node of a quadtree can be used as the root node of a multi-type tree. At least one of multi-type tree partition indication information, partition direction information, and partition tree information can be used to signal whether to partition the coding unit corresponding to a node in the multi-type tree. To partition the coding unit corresponding to a node in the multi-type tree, the multi-type tree partition indication information, partition direction information, and partition tree information can be signaled sequentially.
[0164] A multi-type tree partitioning indication with a first value (e.g., "1") indicates that the current coding unit will undergo a multi-type tree partition. A multi-type tree partitioning indication with a second value (e.g., "0") indicates that the current coding unit will not undergo a multi-type tree partition.
[0165] When the coding unit corresponding to a node of a multi-type tree is further partitioned according to the multi-type tree partitioning structure, the coding unit may include partitioning direction information. The partitioning direction information may indicate in which direction the current coding unit will be partitioned according to the multi-type tree partitioning. Partitioning direction information with a first value (e.g., "1") may indicate that the current coding unit will be vertically partitioned. Partitioning direction information with a second value (e.g., "0") may indicate that the current coding unit will be horizontally partitioned.
[0166] When the coding unit corresponding to a node of a multi-type tree is further partitioned according to the multi-type tree partitioning structure, the current coding unit may include partitioning tree information. The partitioning tree information may indicate the tree partitioning structure that will be used to partition the nodes of the multi-type tree. Partitioning tree information with a first value (e.g., "1") may indicate that the current coding unit will be partitioned according to a binary tree partitioning structure. Partitioning tree information with a second value (e.g., "0") may indicate that the current coding unit will be partitioned according to a ternary tree partitioning structure.
[0167] Partition indication information, partition tree information, and partition direction information can all be flags with a predetermined length (e.g., one bit).
[0168] At least one of the following—quadtree partitioning indication information, multi-type tree partitioning indication information, partitioning direction information, and partitioning tree information—can be entropy encoded / decoded. To entropy encode / decode those types of information, information about neighboring coding units adjacent to the current coding unit can be used. For example, there is a high probability that the partitioning type (partitioned or unpartitioned, partitioning tree, and / or partitioning direction) of the left-hand neighboring coding unit and / or the upper-hand neighboring coding unit of the current coding unit is similar to the partitioning type of the current coding unit. Therefore, contextual information for entropy encoding / decoding of information about the current coding unit can be derived from the information about neighboring coding units. Information about neighboring coding units may include at least one of the following: quadtree partitioning information, multi-type tree partitioning indication information, partitioning direction information, and partitioning tree information.
[0169] As another example, in binary tree partitioning and ternary tree partitioning, binary tree partitioning can be performed first. That is, the current coding unit can first undergo binary tree partitioning, and then the coding unit corresponding to the leaf node of the binary tree can be set as the root node for ternary tree partitioning. In this case, for the coding unit corresponding to the node of the ternary tree, neither quadtree partitioning nor binary tree partitioning can be performed.
[0170] Encoding units that cannot be partitioned according to quadtree, binary tree, and / or ternary tree partitioning structures become the basic units for encoding, prediction, and / or transformation. In other words, these encoding units cannot be further partitioned for prediction and / or transformation. Therefore, partitioning structure information and partitioning information for dividing encoding units into prediction and / or transformation units may not exist in the bitstream.
[0171] However, when the size of the coding unit (i.e., the basic unit used for partitioning) is larger than the size of the maximum transform block, the coding unit can be partitioned recursively until the size of the coding unit is reduced to be equal to or smaller than the size of the maximum transform block. For example, when the size of the coding unit is 64×64 and the size of the maximum transform block is 32×32, the coding unit can be partitioned into four 32×32 blocks for transformation. For example, when the size of the coding unit is 32×64 and the size of the maximum transform block is 32×32, the coding unit can be partitioned into two 32×32 blocks for transformation. In this case, the partitions of the coding unit for transformation are not sent separately by signal, and the partitions of the coding unit for transformation can be determined by comparing the horizontal or vertical dimensions of the coding unit with the horizontal or vertical dimensions of the maximum transform block. For example, when the horizontal dimension (width) of the coding unit is greater than the horizontal dimension (width) of the maximum transform block, the coding unit can be vertically bisected. For example, when the vertical dimension (length) of the coding unit is greater than the vertical dimension (length) of the maximum transform block, the coding unit can be horizontally bisected.
[0172] Information regarding the maximum and / or minimum size of the coding unit and the maximum and / or minimum size of the transform block can be transmitted or determined at a higher level than the coding unit. This higher level can be, for example, the sequence level, the frame level, the strip level, etc. For example, the minimum size of the coding unit can be determined to be 4×4. For example, the maximum size of the transform block can be determined to be 64×64. For example, the minimum size of the transform block can be determined to be 4×4.
[0173] Information regarding the minimum size of the coding unit corresponding to the leaf node of the quadtree (minimum size of the quadtree) and / or the maximum depth of the multi-type tree from the root node to the leaf node (maximum depth of the multi-type tree) can be transmitted or determined at a higher level of the coding unit. For example, the higher level could be the sequence level, the frame level, the strip level, etc. Information regarding the minimum size of the quadtree and / or the maximum depth of the multi-type tree can be transmitted or determined for each of the intra-frame stripes and inter-frame stripes.
[0174] The difference information between the size of the CTU and the maximum size of the transform block can be signaled or determined at a higher level of the coding unit. For example, the higher level could be the sequence level, the frame level, the stripe level, etc. The maximum size of the coding unit corresponding to each node of the binary tree (hereinafter referred to as the maximum size of the binary tree) can be determined based on the size of the coding tree unit and the difference information. The maximum size of the coding unit corresponding to each node of the ternary tree (hereinafter referred to as the maximum size of the ternary tree) can vary depending on the type of stripe. For example, for an intra-frame stripe, the maximum size of the ternary tree could be 32×32. For example, for an inter-frame stripe, the maximum size of the ternary tree could be 128×128. For example, the minimum size of the coding unit corresponding to each node of the binary tree (hereinafter referred to as the minimum size of the binary tree) and / or the minimum size of the coding unit corresponding to each node of the ternary tree (hereinafter referred to as the minimum size of the ternary tree) can be set as the minimum size of the coding block.
[0175] As another example, the maximum size of a binary tree and / or the maximum size of a ternary tree can be signaled or determined at the stripe level. Alternatively, the minimum size of a binary tree and / or the minimum size of a ternary tree can be signaled or determined at the stripe level.
[0176] Based on the size and depth information of the various blocks mentioned above, the four-partition information, multi-type tree partition indication information, partition tree information and / or partition direction information may or may not be included in the bitstream.
[0177] For example, when the size of the coding unit is no greater than the minimum size of the quadtree, the coding unit does not contain four-partition information. Therefore, the four-partition information can be derived from the second value.
[0178] For example, when the size (horizontal and vertical dimensions) of the coding unit corresponding to a node of a multi-type tree is greater than the maximum size (horizontal and vertical dimensions) of a binary tree and / or the maximum size (horizontal and vertical dimensions) of a ternary tree, the coding unit may not be divided into two or three partitions. Therefore, multi-type tree partitioning indication information can be derived from a second value instead of being sent by signal.
[0179] Optionally, when the size (horizontal and vertical dimensions) of the coding unit corresponding to a node of a multi-type tree is the same as the maximum size (horizontal and vertical dimensions) of a binary tree, and / or twice the maximum size (horizontal and vertical dimensions) of a ternary tree, the coding unit may not be further divided into two or three partitions. Therefore, multi-type tree partitioning indication information can be derived from a second value instead of being sent via signal transmission. This is because when the coding unit is partitioned according to the binary tree partitioning structure and / or the ternary tree partitioning structure, coding units smaller than the minimum size of the binary tree and / or the minimum size of the ternary tree are generated.
[0180] Optionally, when the depth of the coding unit corresponding to a node in the multi-type tree is equal to the maximum depth of the multi-type tree, the coding unit may not be further divided into two and / or three partitions. Therefore, multi-type tree partition indication information can be derived from the second value without sending a signal.
[0181] Optionally, multi-type tree partitioning indication information may be signaled only if at least one of the vertical binary tree partitioning, horizontal binary tree partitioning, vertical ternary tree partitioning, and horizontal ternary tree partitioning is feasible for the coding unit corresponding to the node of the multi-type tree. Otherwise, it may not be possible to partition the coding unit into two and / or three partitions. Therefore, multi-type tree partitioning indication information may not be signaled, but can be derived from a second value.
[0182] Optionally, partition direction information may be signaled only if both vertical binary tree partitioning and horizontal binary tree partitioning, or both vertical ternary tree partitioning and horizontal ternary tree partitioning, are feasible for the coding units corresponding to the nodes of the multi-type tree. Otherwise, partition direction information may not be signaled; instead, it may be derived from values indicating possible partition directions.
[0183] Optionally, partition tree information may be signaled only if both vertical binary tree partitions and vertical ternary tree partitions, or both horizontal binary tree partitions and horizontal ternary tree partitions, are feasible for the encoded tree corresponding to the nodes of the multi-type tree. Otherwise, partition tree information may not be signaled; instead, it may be derived from values indicating possible partition tree structures.
[0184] Figure 4 This is a diagram illustrating intra-frame prediction processing.
[0185] Figure 4 The arrows from the center to the outside in the image can indicate the prediction direction of the intra-frame prediction mode.
[0186] Intra-frame coding and / or decoding can be performed using reference samples from neighboring blocks of the current block. A neighboring block can be a reconstructed neighboring block. For example, intra-frame coding and / or decoding can be performed using coding parameters or values of reference samples included in the reconstructed neighboring block.
[0187] A prediction block can represent a block generated by performing intra-frame prediction. A prediction block can correspond to at least one of CU, PU, and TU. The cells of a prediction block can have the size of one of CU, PU, and TU. A prediction block can be a square block with dimensions such as 2×2, 4×4, 16×16, 32×32, or 64×64, or a rectangular block with dimensions such as 2×8, 4×8, 2×16, 4×16, and 8×16.
[0188] Intra-prediction can be performed based on the intra-prediction mode for the current block. The number of intra-prediction modes that the current block can have can be a fixed value, or it can be a value determined differently depending on the attributes of the predicted block. For example, the attributes of the predicted block can include the size and shape of the predicted block.
[0189] Regardless of the block size, the number of intra-prediction modes can be fixed at N. Alternatively, the number of intra-prediction modes can be 3, 5, 9, 17, 34, 35, 36, 65, or 67, etc. Optionally, the number of intra-prediction modes can vary depending on the block size or the color component type, or both. For example, the number of intra-prediction modes can vary depending on whether the color component is a luma signal or a chrominance signal. For example, the number of intra-prediction modes can increase as the block size increases. Optionally, the number of intra-prediction modes for the luma component block can be greater than the number of intra-prediction modes for the chrominance component block.
[0190] Intra-frame prediction modes can be non-angular or angular. Non-angular modes can be DC or planar modes, and angular modes can be prediction modes with a specific direction or angle. Intra-frame prediction modes can be represented by at least one of mode number, mode value, mode number, mode angle, and mode direction. The number of intra-frame prediction modes can be greater than or equal to 1, M, including non-angular and angular modes.
[0191] To perform intra-frame prediction on the current block, a step can be performed to determine whether a sample included in a reconstructed neighboring block can be used as a reference sample for the current block. When there is a sample that cannot be used as a reference sample for the current block, a value obtained by copying or interpolating at least one sample value included in a reconstructed neighboring block, or by performing both copying and interpolation, can be used to replace the unavailable sample value of the sample, and thus the replaced sample value is used as a reference sample for the current block.
[0192] When performing intra-frame prediction, filters can be applied to at least one of the reference samples and the prediction samples based on the intra-frame prediction mode and the current block size.
[0193] In planar mode, when generating the prediction block for the current block, the sample value of the target sample is generated by using a weighted sum of the upper and left reference samples of the current sample, and the upper right and lower left reference samples of the current block, based on the position of the target sample within the prediction block. Furthermore, in DC mode, the average value of the upper and left reference samples of the current block can be used when generating the prediction block. Additionally, in angled mode, the prediction block can be generated using the upper, left, upper right, and / or lower left reference samples of the current block. Real-valued interpolation can be performed to generate the prediction sample values.
[0194] The intra-prediction mode of the current block can be entropy-coded / decoded by predicting the intra-prediction modes of adjacent blocks. When the intra-prediction modes of the current block and its neighboring blocks are the same, information indicating that the intra-prediction modes of the current block and its neighboring blocks are the same can be signaled using predetermined flag information. Furthermore, an indicator of the intra-prediction mode among multiple neighboring blocks that is the same as the intra-prediction mode of the current block can be signaled. When the intra-prediction modes of the current block and its neighboring blocks are different, the intra-prediction mode information of the current block can be entropy-coded / decoded by performing entropy coding / decoding based on the intra-prediction modes of neighboring blocks.
[0195] Figure 5 This is a diagram illustrating an embodiment of inter-screen prediction processing.
[0196] exist Figure 5 In this context, rectangles can represent the image. Figure 5 In the image, the arrow indicates the prediction direction. Based on the encoding type of the frame, frames can be classified into intra-frame frames (I-frames), predictive frames (P-frames), and dual-predictive frames (B-frames).
[0197] I-frames can be encoded via intra-frame prediction without requiring inter-frame prediction. P-frames can be encoded via inter-frame prediction using a reference frame present in one direction (i.e., forward or backward) relative to the current block. B-frames can be encoded via inter-frame prediction using reference frames present in both directions (i.e., forward and backward) relative to the current block. When using inter-frame prediction, the encoder can perform inter-frame prediction or motion compensation, and the decoder can perform the corresponding motion compensation.
[0198] The following section will describe in detail an embodiment of inter-screen prediction.
[0199] Reference frames and motion information can be used to perform inter-frame prediction or motion compensation.
[0200] Motion information of the current block can be derived by each of the encoding device 100 and the decoding device 200 during inter-frame prediction. The motion information of the current block can be derived using motion information of reconstructed neighboring blocks, motion information of co-located blocks (also called col blocks or co-position blocks), and / or motion information of blocks adjacent to the co-position block. A co-position block can represent a block within a previously reconstructed co-located frame (also called a col frame or co-position frame) that is spatially located at the same position as the current block. A co-position frame can be one of one or more reference frames included in a list of reference frames.
[0201] The method for deriving motion information of the current block can be based on changes in the prediction mode of the current block. For example, prediction modes used for inter-frame prediction may include AMVP mode, merge mode, skip mode, and current frame reference mode. The merge mode can be called the motion merge mode.
[0202] For example, when AMVP is used as a prediction mode, at least one of the motion vectors of reconstructed neighboring blocks, co-located blocks, blocks adjacent to co-located blocks, and (0,0) motion vectors can be identified as motion vector candidates for the current block, and a motion vector candidate list is generated using these motion vector candidates. Motion vector candidates for the current block can be derived using the generated motion vector candidate list. Motion information for the current block can be determined based on the derived motion vector candidates. The motion vectors of co-located blocks or blocks adjacent to co-located blocks can be referred to as temporal motion vector candidates, and the motion vectors of reconstructed neighboring blocks can be referred to as spatial motion vector candidates.
[0203] Encoding device 100 can calculate the motion vector difference (MVD) between the motion vector of the current block and motion vector candidates, and can perform entropy encoding on the motion vector difference (MVD). Furthermore, encoding device 100 can perform entropy encoding on the motion vector candidate index and generate a bitstream. The motion vector candidate index indicates the best motion vector candidate among the motion vector candidates included in the motion vector candidate list. Decoding device 200 can perform entropy decoding on the motion vector candidate index included in the bitstream, and can select motion vector candidates for the target block to be decoded from the motion vector candidates included in the motion vector candidate list by using the entropy-decoded motion vector candidate index. Furthermore, decoding device 200 can add the entropy-decoded MVD to the motion vector candidate extracted by entropy decoding to derive the motion vector of the target block.
[0204] The bitstream may include a reference frame index indicating a reference frame. The reference frame index may be entropy encoded by the encoding device 100 and subsequently transmitted as a bitstream to the decoding device 200. The decoding device 200 may generate a predicted block for the decoded target block based on the derived motion vectors and the reference frame index information.
[0205] Another example of a method for deriving motion information for the current block could be a merge pattern. A merge pattern can represent a method for merging the motion of multiple blocks. A merge pattern can represent a pattern for deriving motion information for the current block from motion information of neighboring blocks. When a merge pattern is applied, the motion information of reconstructed neighboring blocks and / or co-located blocks can be used to generate a list of merge candidates. Motion information may include at least one of motion vectors, reference frame indices, and inter-frame prediction indicators. Prediction indicators may indicate unidirectional prediction (L0 prediction or L1 prediction) or bidirectional prediction (L0 prediction and L1 prediction).
[0206] The merge candidate list can be a list of stored motion information. The motion information included in the merge candidate list can be at least one of zero merge candidates and new motion information, wherein the new motion information is a combination of motion information of a neighboring block adjacent to the current block (spatial merge candidate), motion information of a co-located block of the current block included in the reference frame (temporal merge candidate), and motion information existing in the merge candidate list.
[0207] Encoding device 100 can generate a bitstream by performing entropy encoding on at least one of a merge flag and a merge index, and can transmit the bitstream as a signal to decoding device 200. The merge flag may be information indicating whether a merge mode is performed for each block, and the merge index may be information indicating which neighboring block among the current block's neighboring blocks is the target block for merging. For example, neighboring blocks of the current block may include a left-hand neighboring block to the left of the current block, an upper-hand neighboring block positioned above the current block, and a time-adjacent neighboring block.
[0208] Skip mode can be a mode in which motion information of neighboring blocks is applied to the current block as is. When skip mode is applied, encoding device 100 can perform entropy encoding on information about which block's motion information will be used as the current block's motion information to generate a bitstream, and can send the bitstream to decoding device 200 as a signal. Encoding device 100 may not send syntax elements regarding at least one of motion vector difference information, coded block flags, and transform coefficient levels to decoding device 200 as a signal.
[0209] The current frame reference mode can represent the prediction mode in which the previously reconstructed region within the current frame to which the current block belongs was used for prediction. Here, a vector can be used to specify the previously reconstructed region. Information indicating whether the current block will be encoded in the current frame reference mode can be encoded using the current block's reference frame index. A flag or index indicating whether the current block is a block encoded in the current frame reference mode can be signaled, and the flag or index can be derived based on the current block's reference frame index. When the current block is encoded in the current frame reference mode, the current frame can be added to the reference frame list for the current block so that the current frame is located at a fixed position or an arbitrary position in the reference frame list. The fixed position can be, for example, the position indicated by reference frame index 0, or the last position in the list. When the current frame is added to the reference frame list so that the current frame is located at an arbitrary position, a reference frame index indicating the arbitrary position can be signaled.
[0210] Figure 6 This is a diagram illustrating the transformation and quantization processes.
[0211] like Figure 6 As shown, transform and / or quantization are performed on the residual signal to generate a quantized level signal. The residual signal is the difference between the original block and the predicted block (i.e., an intra-frame predicted block or an inter-frame predicted block). The predicted block is generated through intra-frame prediction or inter-frame prediction. The transform can be a first transform, a second transform, or both. The first transform of the residual signal produces transform coefficients, and the second transform of the transform coefficients produces second transform coefficients.
[0212] At least one scheme selected from a variety of predefined transform schemes is used to perform the initial transform. Examples of the predefined transform schemes include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), and Karhunen-Loève Transform (KLT). The transform coefficients generated by the initial transform may undergo a secondary transform. The transform scheme used for the initial and / or secondary transforms can be determined based on the coding parameters of the current block and / or its neighboring blocks. Alternatively, the transform scheme can be determined via signaling of the transform information.
[0213] Since the residual signal is quantized through the first and second transforms, a quantized level signal (quantization coefficients) is generated. Depending on the intra-frame prediction mode or block size / shape, the quantized level signal can be scanned using at least one of diagonal top-right scan, vertical scan, and horizontal scan. For example, when scanning coefficients according to the diagonal top-right scan, the block-form coefficients become a one-dimensional vector form. In addition to the diagonal top-right scan, horizontal scans of two-dimensional block-form coefficients and vertical scans of two-dimensional block-form coefficients can be used depending on the intra-frame prediction mode and / or size of the transform block. The scanned quantized level coefficients can be entropy-encoded for insertion into the bitstream.
[0214] The decoder performs entropy decoding on the bitstream to obtain quantization level coefficients. These coefficients can be arranged in a two-dimensional block format via inverse scanning. For inverse scanning, at least one of the following methods can be used: diagonal top-right scan, vertical scan, and horizontal scan.
[0215] The quantization level coefficients can then be dequantized, then subjected to a second inverse transform as needed, and finally subjected to an initial inverse transform as needed to generate the reconstructed residual signal.
[0216] The reference will be described in detail below. Figure 7 The method described is for configuring reference samples for intra-frame prediction.
[0217] When performing intra-prediction based on a derived intra-prediction mode for the current block or for a sub-block whose size or shape is smaller than the current block, the encoder / decoder can configure reference samples for performing the prediction. In the following description, the current block may refer to the current sub-block.
[0218] Available Figure 7 The reference sample is configured by including at least one sample in at least one reconstructed sample line shown, or by combining multiple samples. Here, the encoder / decoder may use each reconstructed sample in multiple reconstructed sample lines as is, or may use it as a reference sample after filtering is performed between samples on the same reconstructed sample line or between samples on different reconstructed sample lines.
[0219] When through Figure 7 When selecting at least one line from multiple reconstructed sample lines to configure a reference sample, the encoder can send the indicator of the selected reconstructed sample line to the decoder using a signal.
[0220] Optionally, the distance from the current block or the intra-prediction mode of the current block can be used to calculate the distance from the current block. Figure 7 The statistical values of multiple reconstructed sample points selected from multiple reconstructed sample point lines can be used as reference sample points.
[0221] In the example, when calculating statistics using a weighted sum, the weights of the weighted sum are adaptively determined based on the distance from the current block to the reference sample line.
[0222] In the example, when calculating statistics using a weighted sum, the weights of the weighted sum can be adaptively determined based on the intra-prediction mode of the current block.
[0223] Additionally, at least one of the following can be determined based on whether the upper or left boundary of the current block corresponds to at least one of the image, strip, parallel block, and coding tree block (CTB): the number, position, and configuration method of the reconstructed sample lines used to configure the reference samples.
[0224] In the example, when the upper boundary of the current block corresponds to at least one of the frame, strip, parallel block, and CTB, the reference sample can be configured as described in Table 1 below.
[0225] [Table 1]
[0226] Selected reconstructed sample line Top side Left side 1,2 1 1,2 1,2,3,4 1,2 1,2,3,4 2 1 2
[0227] Additionally, information about configuring reference samples can be sent using signals.
[0228] For example, a signal can be used to send at least one of the following information: whether multiple reconstructed sample lines are used and information about the selected reconstructed sample lines.
[0229] The neighboring reconstructed samples used for intra-frame prediction can be configured as reference samples by determining whether the neighboring reconstructed samples are available.
[0230] In the example, a neighboring reconstruction sample is determined to be unavailable when it is not located outside the area that includes at least one of the current block, the strip, the parallel block, and the CTU.
[0231] In the example, when constrained intra-frame prediction is performed for the current block or when a neighboring reconstructed sample is located in a block that is encoded / decoded via inter-frame prediction, it can be determined that the neighboring reconstructed sample is unavailable.
[0232] Additionally, when a neighboring reconstructed sample is determined to be unavailable, the encoder / decoder can replace the unavailable sample with an available neighboring reconstructed sample.
[0233] In the example, the operation of replacing an unavailable sample point can be performed by using an available reconstructed sample point adjacent to the unavailable sample point or by using the statistical values of multiple available reconstructed sample points. Here, when consecutive unavailable sample points exist, the available reconstructed sample point used for replacement can be at least one available reconstructed sample point adjacent to the consecutive unavailable sample points.
[0234] When the current block is divided into multiple sub-blocks and each sub-block has a separate intra-frame prediction mode, reference samples can be configured for each sub-block. Here, depending on the scanning order in which the multiple sub-blocks are predicted, at least one reconstructed sub-block adjacent to the left, top, upper right, and lower left sides of the sub-block to be predicted can be used. Here, the scanning order can be at least one of raster scanning, zigzag scanning, vertical scanning, and horizontal scanning.
[0235] The following section will describe in detail the filtering performed on the reference samples used for intra-frame prediction.
[0236] Whether to perform filtering can be determined based on at least one of block size, block shape, intra-frame prediction mode, segmentation depth, and pixel components.
[0237] According to an embodiment of the present invention, whether to perform filtering on reference samples can be determined based on the size of the current block. Here, the size N of the current block (where N is a positive integer) can be defined by at least one of the block's horizontal size (W), the block's vertical size (H), the sum of the block's horizontal and vertical sizes (W+H), and the number of pixels within the block (W×H).
[0238] In the example, filtering can be performed when the size N of the current block is equal to or greater than a predetermined value T (here, T is a positive integer).
[0239] In another example, filtering can be performed when the size N of the current block is equal to or less than a predetermined value T (here, T is a positive integer).
[0240] In another example, filtering can be performed when the size N of the current block is equal to or greater than a predetermined value T1 and equal to or less than T2. (Here, T1 and T2 are positive integers, and T2 > T1.)
[0241] In another example, filtering can be performed when the size N of the current block is equal to or less than a predetermined value T1 and equal to or greater than T2. (Here, T1 and T2 are positive integers, and T2 > T1.)
[0242] According to embodiments of the present invention, it can be determined whether to perform filtering on reference samples based on the shape of the current block. Here, the block shape can include square blocks and non-square blocks. Furthermore, non-square blocks can be classified into horizontally elongated non-square blocks and vertically elongated non-square blocks.
[0243] In the example, filtering can be performed when the current block is a square block.
[0244] In another example, filtering can be performed when the current block is a non-square block.
[0245] Additionally, when the current block is a non-square block, filtering can be performed on the top and left reference points based on the horizontal value (W) or the vertical value (H) of the current block.
[0246] In the example, it can be determined whether to perform filtering for the upper reference sample based on the horizontal value (W) of the current block, and whether to perform filtering for the left reference sample based on the vertical value (H) of the current block.
[0247] In another example, filtering for the upper and left reference samples can be determined based on the larger of the current block's horizontal value (W) and vertical value (H).
[0248] In another example, filtering for the upper and left reference samples can be determined based on the smaller of the horizontal value (W) and the vertical value (H) of the current block.
[0249] According to an embodiment of the present invention, it can be determined whether to perform filtering on reference samples based on the intra-prediction mode of the current block.
[0250] In the example, filtering can be performed on either the PLANA mode or the DC mode, or both of the PLANA and DC modes, which are non-directional modes.
[0251] In another example, filtering may not be performed for either the PLANA mode or the DC mode, which are non-directional modes, or for both the PLANA mode and the DC mode.
[0252] In another example, filtering may not be performed for all block sizes for either the vertical or horizontal mode, or both, within the orientation pattern.
[0253] Reference sample filtering can be performed on a CurMode that satisfies min{abs(CurMode-Hor_Idx), abs(CurMode-Ver_Idx)}>Th when the intra-prediction mode of the current block is defined as CurMode, the horizontal mode number or index is defined as Hor_Idx, and the vertical mode number or index is defined as Ver_Idx. Here, the threshold Th can be any positive integer and can be a value adaptively determined according to the size of the current block. In the example, the threshold Th can be reduced as the size of the current block increases. Reference sample filtering is performed when min{abs(CurMode-Hor_Idx), abs(CurMode-Ver_Idx)}>Th, and therefore min{abs(CurMode-Hor_Idx), abs(CurMode-Ver_Idx)}>Th can represent the condition for performing reference sample filtering based on the intra-prediction mode. In other words, reference sample filtering can be performed when the above condition is met.
[0254] Whether to perform reference sample filtering can be determined based on the current block partitioning depth.
[0255] Whether to perform reference sample filtering can be determined based on the pixel components of the current block. Here, pixel components may include at least one of a luminance component and a chrominance component (Cb and Cr in the example).
[0256] In the example, reference sample filtering can be performed on the luminance component, but not on the chrominance component.
[0257] In addition, reference sample filtering can be performed on all components, whether it is the luminance component or the chrominance component.
[0258] As described above, it can be determined whether to perform the final filtering for the upper reference sample or the left reference sample, or both, of the current block by combining each filtering execution condition based on at least one of the current block's size, shape, intra-frame prediction mode, segmentation depth, and pixel components.
[0259] The filter type can be determined based on at least one of the following: image features, block size, block shape, intra-prediction mode, segmentation depth, whether the conditions for performing reference sample filtering according to the intra-prediction mode are met, and pixel components. Here, the filter type can represent the filter shape.
[0260] The filter type can be defined as any one of the following: n-tap filter, linear filter, nonlinear filter, bilateral filter, smoothing filter, edge-preserving filter, and sequential statistical filter. At least one of the following can be preset according to the filter type: filter length, number of filter taps, and filter coefficients. N can represent a positive integer.
[0261] The filter type can be determined based on image features of the region including reference samples. Here, the image features can be defined as any one of a uniform region, an edge region, and a pseudo-edge region. Image features can be determined based on the uniformity or texture of the image. Image uniformity and image texture can be indicators with opposite meanings, and image uniformity can be calculated as K*[1 / image texture] (K is a positive integer).
[0262] Figure 8 This is a diagram showing the image features of the region including the reference sample points.
[0263] Reference Figure 8 When the image feature of the region including the reference sample points is (1) an edge region, the filter type can be determined as an edge-preserving filter. Here, the edge region can be a boundary region.
[0264] When the image features of the region including the reference sample points are (2) uniform regions, the filter type can be determined as a smoothing filter.
[0265] When the image features of the region including the reference sample point are (3) pseudo-edge regions, the reference sample point can be identified as noise and filtering anomaly processing can be performed.
[0266] One of the following methods can be used to identify noisy pixels among the pixels to be filtered.
[0267] In the method for determining noise according to an embodiment of the present invention, if the absolute value of the difference between the statistical values of N target pixel values adjacent to the current target pixel value is greater than a predetermined threshold Th, the current target pixel can be determined as a noise pixel. Here, N and Th can be positive integers, and the statistical value can be any one of the average, median, maximum, and minimum values.
[0268] In the example, when the reference sample (or current target pixel) value is defined as Vcur, the value of the target pixel immediately preceding the reference sample (or current target pixel) is defined as Vpre, and the value of the target pixel immediately following the reference sample (or current target pixel) is defined as Vaft, and filtering is performed on a 1D line, the reference sample (or current target pixel) can be identified as a noise pixel if the following conditions are met: (Vcur-Vpre)*(Vafr-Vcur)<0 and max{abs(Vcur-Vpre),abs(Vafr-Vcur)}>=Th. Here, Th can be a positive integer satisfying Th>=0.
[0269] In the example, when the reference sample (or current target pixel) value is defined as Vcur, and the values of the N target pixels adjacent to the reference sample (or current target pixel) are each defined as Vi (here, i = 1, 2, ..., N, and N is a positive integer), and filtering is performed on a 2D line, the reference sample (or current target pixel) can be identified as a noise pixel if the following conditions are met: for all N, Vcur - Vi > or Vcur - Vi < 0, and max{abs(Vcur - V1), abs(Vcur - V2), ..., abs(Vafr - VN)} >= Th. Here, Th can be a positive integer satisfying Th >= 0.
[0270] Additionally, reference samples (or target pixels to be filtered) identified as noise during filtering can be processed in any of the following ways. In one example, filtering may not be performed on noisy target pixels. In another example, filtering may be performed after excluding noisy target pixels from the target region to be filtered.
[0271] As described above, when performing filtering by considering image features of regions including reference samples, prediction efficiency can be improved by generating prediction blocks using reference samples that are close to those of the original image. Specifically, residual signal values can be reduced since edge-preserving filters may be applied to edge regions. Furthermore, ringing artifacts in target boundary regions of the image and contour artifacts that appear when performing orientation prediction can be improved.
[0272] The image uniformity or image texture used when determining image features can be derived as follows.
[0273] In the example, this can be achieved by using at least one of the following formulas. Figure 9 The image uniformity for the upper or left reference point of the image can be derived independently within the block, or the image uniformity for the upper or left reference point of the image can be derived for both the left and upper reference points.
[0274] The uniformity of the upper (top) reference sample points can be derived using Formula 1 or Formula 2.
[0275] [Formula 1]
[0276]
[0277] [Formula 2]
[0278]
[0279] The uniformity of the left reference sample can be derived using formula 3 or formula 4.
[0280] [Formula 3]
[0281]
[0282] [Formula 4]
[0283]
[0284] Alternatively, the overall uniformity of the current block can be derived by using a weighted sum of the left-side uniformity and the top-side uniformity calculated as above.
[0285] As another example of calculating image uniformity, the degree of change or gradient of a reference pixel can be used.
[0286] In the example, such as Figure 10 As shown, the gradient (Pixel_Gradient) of the currently filtered reference pixel (Cur) can be derived according to Equation 5 or Equation 6. Here, in Equation 5 or Equation 6, N can be any positive integer, and W in Equation 6... k It can be any real number.
[0287] [Formula 5]
[0288] Pixel_Gradient=abs(Prev_N-Aft_N)
[0289] [Formula 6]
[0290]
[0291] When the number of upper reference points or left reference points, or the number of upper reference points and left reference points, is M (i.e., M is W, H, W+H or W×H), the average gradient of the entire reference point group is calculated according to Formula 7.
[0292] [Formula 7]
[0293] Where M = W or H or W + H
[0294] When filtering is applied to multiple reference sample lines, the uniformity can be derived for each line, or a weighted sum of the uniformities calculated in each sample line can be used as the overall uniformity.
[0295] The filter type can be determined based on the pixel components of the current block.
[0296] In an example, the filter type for the chrominance component can be set to be the same as the filter type for the luminance component.
[0297] Alternatively, the filter type for the luminance component and the filter type for the chrominance component can be determined independently.
[0298] At least one of the filter length and the filter coefficients can be determined according to the filter type. However, even if the filter type for reference sample filtering is determined, at least one of the filter length and the filter coefficients can be adaptively changed.
[0299] The filter length can be determined based on at least one of the image characteristics, block size, block shape, intra prediction mode and partition depth, whether reference sample filtering is performed, whether the conditions for performing reference sample filtering according to the intra prediction mode are satisfied, and pixel components. Here, the filter length can represent the number of filter taps.
[0300] According to an embodiment of the present invention, the filter length applied to reference sample filtering can be determined based on the size of the current block. Here, the size N of the current block (where N is a positive integer) can be defined as at least one of the horizontal size (W) of the block, the vertical size (H) of the block, the sum of the horizontal size and the vertical size of the block (W + H), and the number of pixels in the block (W × H). Here, the reference sample can represent at least one of the upper reference sample and the left reference sample.
[0301] The filter length can be determined adaptively according to the value of the block size N.
[0302] In an example, when the value of N is less than Th_1, filtering with length L_1 can be applied, when the value of N is equal to or greater than Th_1 and less than Th_2, filtering with length L_2 can be applied, and when the value of N is equal to or greater than Th_(K - 1) and less than Th_K, filtering with length L_K can be applied. Here, L_1 to L_K can be positive integers satisfying L_1 < L_2 <... < L_K, and Th_1 to Th_K can be positive integers satisfying Th_1 < Th_2 <... < Th_K. Alternatively, L_1 to L_K can be positive integers satisfying L_1 < L_2 <... < L_K, and Th_1 to Th_K can be positive integers satisfying Th_K < Th_K - 1 <... < Th_1.
[0303] Additionally, a fixed filter length can be used regardless of the value of the block size N.
[0304] When filtering is applied to multiple reference sample lines, the filter length determined according to the above conditions can be applied identically to all reference sample lines, or independent filter lengths can be applied to each sample line.
[0305] In an example, the filter length to be applied to the first upper reference sample line or the first left reference sample line or both the first upper reference sample line and the first left reference sample line can be determined according to the above conditions, and the filter length to be applied to the second upper reference sample line or the second left reference sample line or both the second upper reference sample line and the second left reference sample line can be determined as a filter length reduced compared to the filter length applied to the first reference sample line. Conversely, the filter length of the second and subsequent upper reference sample lines or left reference sample lines or both the second and subsequent upper reference sample lines and left reference sample lines can be determined as a filter length increased compared to the filter length applied to the first reference sample line.
[0306] According to an embodiment of the present invention, the filter length applied to reference sample filtering can be determined based on the image characteristics of the region including the reference sample. Since the image characteristics have been described in detail above, the description of the image characteristics will be omitted.
[0307] Specifically, the filter length applied to the upper reference sample or the left reference sample or both the upper reference sample and the left reference sample can be adaptively determined according to the uniformity of the region including the reference sample.
[0308] In an example, when the value of the uniformity is less than Th_1, filtering with length L_1 can be applied; when the value of the uniformity is equal to or greater than Th_1 and less than Th_2, filtering with length L_2 can be applied; when the value of the uniformity is equal to or greater than Th_(K - 1) and less than Th_K, filtering with length L_K can be applied. Here, L_1 to L_K can be positive integers satisfying L_1 < L_2 <... < L_K, and Th_1 to Th_K can be positive integers satisfying Th_1 < Th_2 <... < Th_K. Alternatively, L_1 to L_K can be positive integers satisfying L_1 < L_2 <... < L_K, and Th_1 to Th_K can be positive integers satisfying Th_K < Th_K - 1 <... < Th_1. K can be a predetermined positive integer.
[0309] Additionally, a fixed filter length can be used regardless of the value of the uniformity of the region including the reference sample.
[0310] According to an embodiment of the present invention, the filter length applied to the upper reference sample or the left reference sample or both the upper reference sample and the left reference sample can be determined based on the shape of the current block.
[0311] In the example, when the current block is square (i.e., when the horizontal dimension (W) and vertical dimension (H) of the current block are the same), filters of the same length can be applied to the top and left reference samples. Conversely, when the current block is not square, filters of different lengths can be applied to the top and left reference samples.
[0312] In the example, when the horizontal size (W) of the current block is greater than the vertical size (H) of the current block, a filter length larger than that applied to the left reference sample can be applied to the upper reference sample, and when the horizontal size (W) of the current block is smaller than the vertical size (H) of the current block, a filter length larger than that applied to the upper reference sample can be applied to the left reference sample.
[0313] Furthermore, the filter lengths applied to the upper and left reference samples can be determined independently based on the horizontal dimension (W) and vertical dimension (H) of the current block.
[0314] Furthermore, even if the current block is square, filters with different lengths can be applied to the upper and left reference samples.
[0315] Alternatively, the same filter length can be applied to both the upper and left reference samples, regardless of the block shape.
[0316] According to an embodiment of the present invention, the filter length applied to the upper reference sample or the left reference sample, or both the upper reference sample and the left reference sample, can be determined based on the intra-prediction mode of the current block.
[0317] In the example, when the intra-prediction mode of the current block is one of the vertical modes, a filter length larger than that applied to the left reference sample can be applied to the upper reference sample. Similarly, when the intra-prediction mode of the current block is one of the horizontal modes, a filter length larger than that applied to the upper reference sample can be applied to the left reference sample.
[0318] Conversely, when the intra-prediction mode of the current block is one of the vertical modes, a filter length larger than that applied to the upper reference sample can be applied to the left reference sample. Furthermore, when the intra-prediction mode of the current block is one of the horizontal modes, a filter length larger than that applied to the left reference sample can be applied to the upper reference sample.
[0319] Alternatively, the same filter length can be applied to both the upper and left reference samples, regardless of the intra-prediction mode of the current block.
[0320] The filter length can be determined based on the pixel components of the current block.
[0321] In the example, the filter length for the chroma component can be set to the same as the filter length for the luminance component.
[0322] In addition, the filter lengths for the luminance component and the chrominance component can be determined independently.
[0323] The filter coefficients can be determined based on at least one of the following: image features, block size, block shape, intra-prediction mode and segmentation depth, whether reference sample filtering is performed, whether the conditions for performing reference sample filtering according to the intra-prediction mode are met, and pixel components. Here, the filter coefficients can represent the number of filter sets.
[0324] As described above, reference sample filtering for intra-frame prediction has been described in detail. In addition to reference sample filtering for intra-frame prediction, the methods described above for determining whether to perform filtering, determining the filter type, determining the filter length, and determining the filter coefficients can be applied in the same way. Figure 1 encoder or Figure 2 The following steps are required for the decoder.
[0325] - Interpolation filtering in intra-prediction units and boundary region filtering for intra-prediction blocks
[0326] - Interpolation filtering for generating prediction blocks in motion compensation units and boundary region filtering for inter-frame prediction blocks.
[0327] - Interpolation filtering used to generate prediction blocks in the motion prediction unit
[0328] - Deblocking filtering, SAO (Sample Adaptive Offset) filtering, and ALF (Adaptive Loop Filtering) in the filter unit.
[0329] - At least one of the following filtering techniques performed in the encoder or decoder to correct (or refine or fine-tune) motion information: OBMC (Overlap Block Motion Compensation), FRUC (Frame Rate Boost), and BIO (Bidirectional Optical Flow).
[0330] Therefore, in the following description, filtering can refer to at least one of the following filters: reference sample filtering / interpolation filtering / boundary region filtering in the intra-prediction unit, boundary region filtering of the generated prediction block for interpolation filtering in the motion prediction unit and motion compensation unit, intra-loop filtering in the filter unit, and OBMC, FRUC and BIO in the encoder and decoder for correcting motion information.
[0331] According to embodiments of the present invention, at least one of the following can be determined based on block size, block shape, prediction mode, intra-frame prediction mode, inter-frame prediction mode, local features of the image, global features of the image, whether reference sample filtering is performed, whether the conditions for performing reference sample filtering according to the intra-frame prediction mode are met, pixel components, and other coding parameters: whether filtering is performed, filter type, filter length, and filter coefficients.
[0332] In the example, the type of filter used in the interpolation filtering in the intra-prediction unit can be determined based on at least one of the following: image features, block size, block shape, intra-prediction mode, segmentation depth, whether reference sample filtering is performed, whether the conditions for performing reference sample filtering according to the intra-prediction mode are met, and pixel components.
[0333] Furthermore, the filter coefficients used in the interpolation filtering in the intra-prediction unit can be determined based on at least one of the following: image features, block size, block shape, intra-prediction mode, segmentation depth, whether reference sample filtering is performed, whether the conditions for performing reference sample filtering according to the intra-prediction mode are met, and pixel components.
[0334] For example, a first set of filter coefficients may be used during interpolation filtering when reference sample filtering is performed or when the conditions for performing reference sample filtering based on the intra-prediction mode are met. A second set of filter coefficients may be used during interpolation filtering when reference sample filtering is not performed or when the conditions for performing reference sample filtering based on the intra-prediction mode are not met.
[0335] In other words, filter coefficients can be determined from a set of filter coefficients that includes at least one filter coefficient, depending on whether reference sample filtering is performed or whether the conditions for performing reference sample filtering according to the intra-frame prediction mode are met.
[0336] The set of filter coefficients used in the interpolation filtering in the intra-prediction unit can be the same set of filter coefficients used when generating the luma prediction block or chroma prediction block in the motion compensation unit.
[0337] Furthermore, when filtering is not performed on reference samples that are the target of interpolation filtering in the intra-prediction unit, different interpolation filter coefficients may be used depending on whether the conditions for performing reference sample filtering according to the intra-prediction mode are met.
[0338] In another example, when filtering is performed on reference samples that are the target of interpolation filtering in the intra-prediction unit, different interpolation filter coefficients may be used depending on whether the conditions for performing reference sample filtering according to the intra-prediction mode are met.
[0339] Here, the filter coefficient set can represent a set of K filter coefficient configurations that are distinct from each other. Furthermore, K can be a positive integer. Additionally, the filter coefficient set or filter coefficients can represent an interpolation filter coefficient set or interpolation filter coefficients.
[0340] The filtering according to embodiments of the invention can be applied to a pixel component including at least one of a luminance component and a chrominance component (e.g., Cb, Cr) by using one of the following methods.
[0341] In the example, the filter can be applied to the luminance component but not the chrominance component. Conversely, the filter can be applied to the chrominance component but not the luminance component. Alternatively, the filter can be applied to both the luminance and chrominance components.
[0342] In the example, the same filter can be applied to both the luminance and chrominance components. Alternatively, different filters can be applied to both the luminance and chrominance components.
[0343] In the example, the same filter can be applied to the Cb and Cr chromaticity components. Alternatively, different filters can be applied to the Cb and Cr chromaticity components.
[0344] The following will describe the application direction of filtering, the pixel region used for filtering, and the pixel unit used for filtering according to embodiments of the present invention.
[0345] According to embodiments of the present invention, the filtering application direction can be any one of the horizontal direction, the vertical direction, and a direction with any angle.
[0346] Figure 11 This is a diagram illustrating the direction of applied filtering according to an embodiment of the present invention.
[0347] Reference Figure 11 (a) represents the horizontal direction, (b) represents the vertical direction, and (c) represents the direction with any angle (θ). Here, θ can be an integer or a real number.
[0348] Additionally, for the target pixel being filtered, it can be achieved by... Figure 11 At least one of the directions (a), (b) and (c) is combined to apply the filter repeatedly or recursively.
[0349] According to embodiments of the present invention, the pixel region for filtering can be any one of the following: a pixel located in the horizontal direction of the target pixel, a pixel located in the vertical direction of the target pixel, a pixel located in multiple horizontal lines including the target pixel, a pixel located in multiple vertical lines including the target pixel, a pixel in a cross-shaped region including the target pixel, and a pixel in a geometric region including the target pixel.
[0350] Figure 12 This is a diagram illustrating a pixel region used for filtering according to an embodiment of the present invention.
[0351] In the example, such as Figure 12 As shown in (a), filtering can be performed using the target pixel and the pixel located in the horizontal direction.
[0352] Optionally, such as Figure 12 As shown in (b), filtering can be performed using the target pixel and the pixel located in the vertical direction.
[0353] Optionally, such as Figure 12 As shown in (c), filtering can be performed using pixels located along N horizontal lines that include the target pixel. Here, the horizontal lines can be either above or below the target pixel, or both. N can be a positive integer greater than 1. Alternatively, when N is odd, the horizontal lines can be located in equal numbers above and below the target pixel.
[0354] Optionally, such as Figure 12 As shown in (d), filtering can be performed using pixels located along N vertical lines including the target pixel. Here, the vertical lines can be the left or right side of the target pixel, or both. N can be a positive integer greater than 1. Additionally, when N is odd, the vertical lines can be located in equal numbers on the left and right sides of the target pixel.
[0355] Optionally, such as Figure 12 As shown in (e), filtering can be performed using pixels within a cross-shaped region that includes the target pixel. Here, M and N can be positive integers greater than 2, where the vertical length of the cross-shaped region is M and the horizontal length is N.
[0356] Optionally, such as Figure 12 As shown in (f), filtering can be performed using pixels within a geometric region that includes the target pixel. Here, the geometric region can be at least one of a square, a non-square, a triangle, a trapezoid, and a circle.
[0357] In addition, Figure 12 (a) to Figure 12 All pixels within the shaded pixel region in (f) can be used for filtering, or filtering can be performed using a subset of pixels within that region.
[0358] For example, filtering can be performed using the target pixel (X), and... Figure 12 (a) to Figure 13In (f), consecutive pixels of the target pixel (X) within a pixel region are represented in shaded form. Alternatively, filtering can be performed using pixels spaced at a predetermined distance K (K is a positive integer) from the target pixel (X).
[0359] According to embodiments of the present invention, the pixel unit for applied filtering can be an integer pixel unit (integer pel), a fractional pixel unit (fractional pel), or both an integer pixel unit (integer pel) and a fractional pixel unit (fractional pel). Here, the fractional unit can be 1 / 2 (half pel), 1 / 4 (quarter pel), 1 / 8 pel, 1 / 16 pel, 1 / 32 pel, 1 / 64 pel, ..., and 1 / N pel. Here, N is a positive integer.
[0360] Figure 13 This is a diagram illustrating an example of an embodiment where filtering is applied in 1 / 4 (or quarter pel) units.
[0361] exist Figure 13 In this context, pixels represented by uppercase letters in shaded form represent pixels at integer positions, while other pixels, including those with lowercase letters, represent pixels at fractional units. Furthermore, Figure 13 The "X" i,j In the symbol “i”, i can represent the horizontal index and j can represent the vertical index.
[0362] exist Figure 11 In this context, the pixel unit used for applying filtering can be at least one of the following.
[0363] - Apply the filter to pixels A in integer units. i,j
[0364] - Apply filtering unit pixels b i,j and h i,j
[0365] - Apply filtering unit pixels a i,j c i,j d i,j and n i,j
[0366] - Apply the filter to pixels e in fractional units i,j f i,j g i,j i i,j j i,j k i,j p i,j q i,j and r i,jThe pixel is located within four adjacent integer pixels forming a square.
[0367] Additionally, the pixels used to filter the target pixels in each pixel unit that may be filtered can be Figure 12 At least one direction of the applied filter and Figure 14 Any combination of at least one region used for filtering.
[0368] The following sections will describe in detail n-tap filtering, smoothing filtering, edge-preserving filtering, 1D filtering, 2D filtering, and sequential statistical filtering according to embodiments of the present invention. Here, n can be a positive integer.
[0369] The filtering using an n-tap filter according to an embodiment of the invention can be performed using the following formula 8. Here, the target pixel to be filtered is X, and the pixels used for filtering are {b1, b2, ..., b}. n The filter coefficients are {c1, c2, ..., c}. n The value of the target pixel after filtering is X', and n is a positive integer.
[0370] [Formula 8]
[0371]
[0372] In the example of performing filtering, when the target pixel being filtered is "b" 0,0 When an 8-tap filter with a target pixel length of 8 and filter coefficients of {-1, 4, -11, 40, 40, -11, 4, -1} is applied, the filter value can be calculated according to Formula 9.
[0373] [Formula 9]
[0374] b (0,0) ={-1×A (-3,0) +4×A (-2,0) -11×A (-1,0) +40×A (0,0) +40×A (1,0) -11×A (2,0) +4×A (3,0) -1×A (4,0) +32} / 64
[0375] Additionally, for the target pixel X to be filtered, such as Figure 13 As shown in the example, when a portion of the region to be filtered is outside the frame boundary, block boundary, or sub-block boundary, filtering for the target pixel X can be performed using any of the following methods.
[0376] - Do not perform filtering on the target pixel X
[0377] - Filtering for target pixel X is performed by using a region that exists within the boundaries of a frame, block, or sub-block for performing the filtering.
[0378] The length of the smoothing filter according to embodiments of the present invention can be any positive integer. Furthermore, the coefficients (or filter coefficients) of the smoothing filter can be determined using any of the following methods.
[0379] In the example, the coefficients of the smoothing filter can be derived using a Gaussian function. The 1D and 2D Gaussian functions can be expressed as Equation 10 below.
[0380] [Formula 10]
[0381]
[0382] σ: Standard deviation
[0383] The filter coefficients can be the pixel range (0–2) derived from Equation 10. BitDepth Quantization value within )
[0384] In the example, when filtering is performed in 1 / 32 pel units using a 1D Gaussian function with a tap length of 4, the filter coefficients applied to the target pixel at each integer / fractional unit position are shown in Table 2 below. Here, in Table 2, 0 can represent an integer unit pixel, and the filter coefficients from 17 / 32 to 31 / 32 can be derived by performing symmetry with respect to the filter coefficients from 16 / 32 to 1 / 32.
[0385] [Table 2]
[0386]
[0387] In another example, when filtering is performed in 1 / 32 pel units using a 1D Gaussian function with a length of 4, the filter coefficients applied to the target pixel at each integer / fraction unit position are shown in Table 3 below.
[0388] [Table 3]
[0389]
[0390] In Table 3, the sum of filter coefficients can be represented using M bits. Here, the sum of filter coefficients can not exceed 2 << M. For example, M can be a positive integer including 6. When M is 6, the sum of filter coefficients can exceed 64, that is, 2 << 6.
[0391] At least one of the filter coefficients in Table 3 can represent at least one filter coefficient within the first set of filter coefficients.
[0392] However, the filter coefficients that may be derived from the Gaussian function are not specified as the coefficient values described in Tables 2 and 3, and can be determined based on block size, block shape, intra / inter-frame prediction mode, local features of the image, global features of the image, whether reference sample filtering is performed, whether the conditions for performing reference sample filtering according to the intra-frame prediction mode are met, pixel components, and coding parameters.
[0393] The filter coefficients in Tables 2 and 3 can be examples of M-tap filter coefficients, and M can be a positive integer including 4.
[0394] The filter coefficients derived from the Gaussian function can be used for at least one of the following filters: reference sample filtering / interpolation filtering / boundary region filtering in the intra-prediction unit, boundary region filtering of the generated prediction block in the motion prediction unit and motion compensation unit, in-loop filtering in the filter unit, and OBMC, FRUC and BIO in the encoder and decoder for correcting motion information.
[0395] In another example, the coefficients of the smoothing filter can be derived using a DCT-based function. The forward and inverse (including fractional units) DCT transforms can be represented as shown in Equation 11 below.
[0396] [Formula 11]
[0397]
[0398] The filter coefficients can be the pixel range (0–2) derived from the above formula. BitDepth The quantized value within )
[0399] In the example, Figure 15 In this context, the filter coefficients applied to fractional units such as 1 / 4, 1 / 2, and 3 / 4 can be as follows.
[0400]
[0401] In another example, when filtering is performed using a DCT-based function at 1 / 32 pel units with a tap length of 4, the filter coefficients applied to the target pixel at each integer / fraction unit position are shown in Table 4 below.
[0402] [Table 4]
[0403]
[0404] In Table 4, the sum of filter coefficients can be represented using M bits. Here, the sum of filter coefficients can not exceed 2 << M. For example, M can be a positive integer including 6. When M is 6, the sum of filter coefficients can exceed 64, that is, 2 << 6.
[0405] At least one of the filter coefficients in Table 4 can represent at least one filter coefficient within the first set of filter coefficients.
[0406] However, the filter coefficients that may be derived from the DCT-based function are not specified as the coefficient values in Table 4, and can be determined based on block size, block shape, intra / inter-frame prediction mode, local features of the image, global features of the image, whether reference sample filtering is performed, whether the conditions for performing reference sample filtering according to the intra-frame prediction mode are met, pixel components, and coding parameters.
[0407] The coefficient values and filter coefficients in Table 4 can be examples of M-tap filter coefficients, and M can be a positive integer including 4.
[0408] The filter coefficients derived from the DCT-based function can be used for at least one of the following filters: reference sample filtering / interpolation filtering / boundary region filtering in the intra-prediction unit, boundary region filtering of the generated prediction block in the motion prediction unit and motion compensation unit, intra-loop filtering in the filter unit, and OBMC, FRUC and BIO in the encoder and decoder for correcting motion information.
[0409] In another example, the coefficients of a smoothing filter can be derived from a median filter. Here, a median filter can be a filter that uses the median value among the pixel values in the region used to filter the target pixel as the filter value.
[0410] The length of the edge-preserving filter according to embodiments of the present invention can be any positive integer. Furthermore, the coefficients (or filter coefficients) of the edge-preserving filter can be determined using at least one of the following methods.
[0411] For example, the coefficients of the edge-preserving filter can be derived using a two-sided function. The two-sided function can be expressed as Equation 12 below.
[0412] [Formula 12]
[0413] in,
[0414] In Formula 12, σ d It can be a parameter (or spatial parameter) that adjusts the weights by considering the distance between two pixels x and y, and σr It can be a parameter (or range parameter) that adjusts the weights by considering the difference between the pixel values I(x) and I(y) of two pixels x and y respectively. N can represent the number of pixels (y) in the region used to filter the target pixel (x), and can be a positive integer.
[0415] For the space parameter σ d or range parameter σ r Or spatial parameter σ d and range parameter σ r Both can use fixed values for all pixels. However, they are not limited to this and can use variable values based on at least one of the following: block size, block shape, intra / inter-frame prediction mode, local features of the image, global features of the image, and coding parameters.
[0416] σ can be d or σ r Or σ d and σ r Both are determined by the value of BitDepth. In the example, "σ" d or σ r =1<<(BitDepth-K)”, where K can be a positive integer equal to or less than the bit depth, or it can be 0.
[0417] In addition, σ can be determined based on the block size. d or σ r Or σ d and σ r Both. In the example, when the block size is N, N can be defined as at least one of the block's horizontal size, vertical size, the sum of its horizontal and vertical dimensions, and the product of its horizontal and vertical dimensions. Here, as N increases, σ can be used. d or σ r For larger values in N, or when N becomes smaller, σ can be used. d or σ r The larger value in the range.
[0418] In addition, σ can be determined based on intra-frame prediction mode or inter-frame prediction mode. d or σ r Or σ d and σ r Both.
[0419] Additionally, when the horizontal and vertical lengths of the current block are different, different values of σ can be used respectively. d or σ r It is applied to both horizontal and vertical filtering.
[0420] Furthermore, when the local features of pixels in the filtered region, the global features of the filtered region, or both the local features of pixels in the filtered region and the global features of the filtered region are defined as uniformity, σ can be used when the uniformity increases. d or σ r The larger value in the range. Alternatively, when the uniformity decreases, σ can be used. d or σ r The larger value in the range. Here, image features can be determined according to one of the following units: frame unit, block unit, line unit, and pixel unit.
[0421] The 1D filter according to embodiments of the present invention can be any one of a nearest neighbor filter, a linear filter, and a cubic filter. The length of the 1D filter can be any positive integer, and can be as follows: Figure 15 The derivation of the coefficients (or filter coefficients) of the 1D filter is shown in the figure.
[0422] exist Figure 16 In the diagram, pixels represented by dashed lines can represent target pixels to be filtered, and pixels represented by solid lines can be pixels within the region used to perform filtering on the target pixels. The length of each line in the solid and dashed lines can represent the magnitude of the filter coefficients.
[0423] The 2D filter according to embodiments of the present invention can be any one of a 2D nearest neighbor filter, a bilinear filter, and a bicubic filter. The length of the 2D filter can be any positive integer, and can be as follows: Figure 16 The derivation of the coefficients (or filter coefficients) of the 2D filter is shown in the figure.
[0424] exist Figure 17 In the diagram, pixels represented by dashed lines can represent target pixels to be filtered, and pixels represented by solid lines can be pixels within the region used to perform filtering on the target pixels. The length of each line in the solid and dashed lines can represent the magnitude of the filter coefficients.
[0425] The target pixel subjected to sequential statistical filtering according to embodiments of the present invention and the N pixels in the region used to perform filtering on the target pixel can be sorted in ascending or descending order, and the Kth value can be used as the filtered pixel value. Here, N and K can be positive integers satisfying K<=N.
[0426] Furthermore, for the target pixel to be filtered, at least one filter can be applied repeatedly or recursively. Additionally, filter coefficients or filter length, or both, can be set to predefined values in the encoder / decoder. Furthermore, filter coefficients or filter length, or both, can be determined in the encoder and sent to the decoder as signals.
[0427] Figure 17 This is a flowchart illustrating an image decoding method according to an embodiment of the present invention.
[0428] Reference Figure 18 In S1701, the decoder can determine the reference sample point of the current block.
[0429] Here, the reference sample point of the current block can be at least one of at least one reconstructed sample point line located to the left of the current block and at least one reconstructed sample point line located above the current block.
[0430] Furthermore, in S1702, the decoder can perform filtering on the reference sample based on the features of the region including the reference sample.
[0431] Here, the features of the region including the reference sample point can be any one of a uniform region, an edge region, and a pseudo-edge region.
[0432] In addition, the step of performing filtering on the reference sample in S1702 includes: when the region including the reference sample is characterized as a uniform region, filtering is performed by using a smoothing filter; when the region including the reference sample is characterized as an edge region, filtering is performed by using an edge-preserving filter; and when the region including the reference sample is characterized as a pseudo-edge region, filtering is performed by excluding pixels determined to be noise.
[0433] Here, the characteristics of the region including the reference sample points can be determined based on the uniformity of the region.
[0434] Additionally, the step of performing filtering on the reference sample in S1702 may include: determining a filter length based on at least one of the current block size, the current block shape, the current block intra-prediction mode, the current block segmentation depth, and the current block pixel components; and filtering the reference sample based on the determined filter length.
[0435] Furthermore, in S1703, the decoder can perform intra-frame prediction by using reference samples from which the filtering is performed.
[0436] Additionally, the decoder can generate prediction blocks by using interpolation filters when performing intra-frame prediction.
[0437] Here, the type of interpolation filter used in intra-prediction can be determined based on whether reference sample filtering is performed or whether the conditions for performing reference sample filtering according to the intra-prediction mode are met.
[0438] Figure 18 This is a flowchart illustrating an image decoding method according to another embodiment of the present invention.
[0439] Reference Figure 17 In S1801, the decoder can determine the reference sample point of the current block.
[0440] Subsequently, in S1802, the decoder can determine whether to perform reference sample filtering based on at least one of the current block size, current block shape, current block intra-prediction mode, current block partitioning depth, and current block pixel components.
[0441] Subsequently, in S1802-, when it is determined in S1802 that reference sample filtering will be performed, in S1803, the decoder may perform reference sample filtering based on the features of the region including the reference sample.
[0442] Subsequently, in S1804, the decoder can perform intra-frame prediction by using the reference samples from which the filtering was performed.
[0443] In S1802-No, if it is determined in S1802 that reference sample filtering is not performed, in S1805, the decoder can perform intra-frame prediction by using the reference samples determined in S1801.
[0444] The same exploit can be performed in the encoder. Figure 18 and The image decoding method is described.
[0445] Additionally, the recording medium according to the invention may include a bitstream generated by an image coding method, wherein the image coding method includes: determining reference samples of a current block; performing filtering on the reference samples based on features of a region including the reference samples; and performing intra-frame prediction by using the filtered reference samples.
[0446] The above embodiments can be performed in the same way in both the encoder and decoder.
[0447] At least one or a combination of the above embodiments can be used to encode / decode images.
[0448] The order in which the above embodiments are applied may differ between the encoder and the decoder, or the order in which the above embodiments are applied may be the same in the encoder and the decoder.
[0449] The above embodiments can be performed on each luminance signal and chrominance signal, or the above embodiments can be performed on both luminance and chrominance signals in the same way.
[0450] The block shape of the above embodiments of the present invention can be square or non-square.
[0451] The above embodiments of the present invention can be applied based on the size of at least one of the coding block, prediction block, transform block, block, current block, coding unit, prediction unit, transform unit, unit, and current unit. Here, size can be defined as the minimum size or maximum size, or both, for which the above embodiments are applied, or it can be defined as a fixed size to which the above embodiments are applied. Furthermore, in the above embodiments, a first embodiment can be applied to a first size, and a second embodiment can be applied to a second size. In other words, the above embodiments can be applied in combination based on sizes. Furthermore, the above embodiments can be applied when the size is equal to or greater than the minimum size and equal to or less than the maximum size. In other words, the above embodiments can be applied when the block size is included within a specific range.
[0452] For example, the above embodiments can be applied when the size of the current block is 8×8 or larger. For example, the above embodiments can be applied when the size of the current block is 4×4 or larger. For example, the above embodiments can be applied when the size of the current block is 16×16 or larger. For example, the above embodiments can be applied when the size of the current block is equal to or greater than 16×16 and equal to or less than 64×64.
[0453] The above embodiments of the present invention can be applied according to time layers. To identify the time layer to which the above embodiments can be applied, additional identifiers can be signaled, and the above embodiments can be applied to a specified time layer identified by the corresponding identifier. Here, the identifier can be defined as the lowest or highest layer to which the above embodiments can be applied, or both, or it can be defined as indicating a specific layer to which the embodiment is applied. Furthermore, a fixed time layer to which the embodiments are applied can be defined.
[0454] For example, the above embodiments can be applied when the time layer of the current image is the lowest layer. For example, the above embodiments can be applied when the time layer identifier of the current image is 1. For example, the above embodiments can be applied when the time layer of the current image is the highest layer.
[0455] The strip types to which the above embodiments of the present invention are applied can be defined, and the above embodiments can be applied according to the corresponding strip types.
[0456] In the above embodiments, the method is described based on a flowchart having a series of steps or units. However, the present invention is not limited to the order of these steps; rather, some steps may be performed simultaneously with other steps or in a different order. Furthermore, those skilled in the art should understand that the steps in the flowchart are not mutually exclusive, and other steps may be added to the flowchart or some steps may be removed from the flowchart without affecting the scope of the present invention.
[0457] The embodiments include various aspects of the examples. Not all possible combinations of the aspects may be described, but those skilled in the art will recognize different combinations. Therefore, the invention may include all substitutions, modifications, and alterations within the scope of the claims.
[0458] Embodiments of the present invention can be implemented in the form of program instructions, which can be executed by various computer components and recorded in a computer-readable recording medium. The computer-readable recording medium may individually include program instructions, data files, data structures, etc., or may include a combination of program instructions, data files, data structures, etc. The program instructions recorded in the computer-readable recording medium may be specifically designed and constructed for the present invention or are well known to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic recording media (such as hard disks, floppy disks, and magnetic tapes), optical data storage media (such as CD-ROMs or DVD-ROMs), magneto-optical media (such as floppy disks), and hardware devices specifically configured to store and implement program instructions (such as read-only memory (ROM), random access memory (RAM), flash memory, etc.). Examples of program instructions include not only machine language code formatted by a compiler but also high-level language code that can be implemented by a computer using an interpreter. The hardware device may be configured to be operated by one or more software modules or vice versa to implement the processing according to the present invention.
[0459] Although the invention has been described with respect to specific items (such as detailed elements) and limited embodiments and drawings, these are provided only to aid in a more complete understanding of the invention, and the invention is not limited to the embodiments described above. Those skilled in the art will understand that various modifications and changes can be made to the above description.
[0460] Therefore, the spirit of the present invention should not be limited to the above embodiments, and the entire scope of the appended claims and their equivalents shall fall within the scope and spirit of the present invention.
[0461] Industrial applicability
[0462] This invention can be used to encode / decode images.
Claims
1. A method for decoding an image, the method comprising: Select the reference sample line of the current block from among multiple reference sample lines adjacent to the current block; Determine the reference points included in the selected reference point line for the current block; The reference samples are filtered. as well as Intra-frame prediction is performed using filtered reference samples, and The filtering includes applying different filters to each reference sample line. The intra-frame prediction is performed by an interpolation filter on the filtered reference samples, and The interpolation filter is determined based on whether the conditions for performing filtering on reference samples according to the intra-prediction mode of the current block are met.
2. The method of claim 1, further comprising: Based on at least one of the following: the size of the current block, the intra-prediction mode of the current block, whether the current block is divided, and the pixel components of the current block, it is determined whether to perform filtering on the reference sample.
3. The method as described in claim 2, wherein, The filtering of the reference samples is performed based on the determined results.
4. The method of claim 2, wherein, The size of the current block is defined as the number of pixels in the current block.
5. A method for encoding an image, the method comprising: Select the reference sample line of the current block from among multiple reference sample lines adjacent to the current block; Determine the reference points included in the selected reference point line for the current block; The reference samples are filtered. as well as Intra-frame prediction is performed using filtered reference samples, and The filtering includes applying different filters to each reference sample line. The intra-frame prediction is performed by an interpolation filter on the filtered reference samples, and The interpolation filter is determined based on whether the conditions for performing filtering on reference samples according to the intra-prediction mode of the current block are met.
6. A method for transmitting a bit stream, the method comprising: Transmitting a bitstream generated by a video encoding method; The video encoding method includes: Select the reference sample line of the current block from among multiple reference sample lines adjacent to the current block; Determine the reference points included in the selected reference point line for the current block; Filter the reference samples; and Intra-frame prediction is performed using filtered reference samples, and The filtering includes applying different filters to each reference sample line. The intra-frame prediction is performed by an interpolation filter on the filtered reference samples, and The interpolation filter is determined based on whether the conditions for performing filtering on reference samples according to the intra-prediction mode of the current block are met.
Citation Information
Patent Citations
Image encoding / decoding method and apparatus for same
CN103404151A
Method and apparatus for pre-prediction filtering for use in block-prediction techniques
CN106134201A