Image encoding / decoding method and apparatus using model-based prediction method, and bitstream storage medium
By configuring the sample set and deriving the mapping model, entropy encoding/decoding is performed using the information redundancy between samples, which solves the efficiency limitation caused by information redundancy in image encoding/decoding and achieves more efficient encoding/decoding.
Patent Information
- Application Number
- CN202480025198.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2024-04-09
- Publication Date
- 2025-11-11
AI Technical Summary
In existing image encoding/decoding methods, there is information redundancy between luminance component samples and chrominance component samples, as well as between neighboring samples of the target block, which limits encoding efficiency.
By configuring a sample set and deriving a mapping model, the information redundancy between samples is used to reduce signaling bits, and entropy coding/decoding techniques are applied.
It improves the efficiency of image encoding/decoding and reduces the use of signaling bits.
Smart Images

Figure CN120937349A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a method and apparatus for encoding / decoding images, and more specifically, to a method and apparatus for encoding / decoding images including a weighted sum of prediction blocks. Background Technology
[0002] Recently, the demand for high-resolution and high-quality images, such as HD (High Definition) and UHD (Ultra High Definition) images, has increased across various application areas. As image data becomes higher resolution and higher quality, the relative data volume increases compared to existing image data, thus increasing transmission and storage costs when transmitting image data using media such as existing wired and wireless broadband circuits or storing image data using existing storage media. Efficient image encoding / decoding techniques may be needed for images with higher resolution and image quality to address these issues arising from the increasing resolution and quality of image data.
[0003] There are various techniques, such as inter-frame prediction techniques that use image imaging technology to predict pixel values included in the current frame from previous or subsequent frames of the current frame, intra-frame prediction techniques that use pixel information in the current frame to predict pixel values included in the current frame, transformation and quantization techniques for compressing the energy of the remaining signal, and entropy coding techniques that assign short symbols to values with high occurrence frequency and long symbols to values with low occurrence frequency. Image data can be effectively compressed, transmitted, or stored by using these image compression techniques. Summary of the Invention
[0004] Technical issues In traditional image encoding / decoding methods and devices, there is a limitation on improving encoding efficiency because of the redundancy of information between luminance component samples and chrominance component samples, as well as between samples in the vicinity of the target block and samples within the target block.
[0005] This disclosure provides a method and apparatus for improving image encoding / decoding efficiency by reducing signaling bits through model-based prediction and utilizing information redundancy between samples.
[0006] Technical solution An embodiment of this disclosure is an image encoding method and apparatus, comprising: configuring a sample set and deriving a mapping model; applying the mapping model to a target block; and entropy encoding / decoding the encoded information.
[0007] Beneficial effects This disclosure provides a method and apparatus for improving image encoding / decoding efficiency by utilizing information redundancy between samples through model-based prediction to reduce signaling bits. Attached Figure Description
[0008] Figure 1 This is a block diagram illustrating the configuration of an encoding apparatus according to an embodiment of the present disclosure.
[0009] Figure 2 This is a block diagram illustrating the configuration of an embodiment of the decoding apparatus according to which the present disclosure is applied.
[0010] Figure 3 It is a diagram that roughly illustrates the segmentation structure of an image during encoding and decoding.
[0011] Figure 4 This is a diagram illustrating an embodiment of intra-frame prediction processing.
[0012] Figure 5 This is a diagram illustrating an embodiment of inter-frame prediction processing.
[0013] Figure 6 It is a diagram used to describe the transformation and quantization processes.
[0014] Figure 7 An example illustrating the partitioning boundaries in a geometric partitioning pattern is shown.
[0015] Figure 8 It is a diagram showing the partition boundaries, partition offsets, and partition angles in the geometric partitioning pattern.
[0016] Figure 9 An example is shown that represents a weighted graph for each prediction block based on the geometric partition boundaries.
[0017] Figure 10 This is an illustration of an example of template matching.
[0018] Figure 11 This illustrates the subsampling method for the search range of template matching.
[0019] Figure 12 A subsampling method for templates used in template matching is shown.
[0020] Figures 13 to 18 The search methods in template matching according to the embodiments are shown respectively.
[0021] Figure 19 A method for configuring a first template in affine mode according to an embodiment is shown.
[0022] Figure 20 The example illustrates a method for configuring a second template in affine mode.
[0023] Figure 21 Various examples of subsampling methods in bilateral matching are shown.
[0024] Figure 22 A bilateral matching according to an embodiment is shown.
[0025] Figure 23 This is a flowchart of an encoding method for a target block according to an embodiment.
[0026] Figure 24 This is a flowchart of a decoding method for a target block according to an embodiment.
[0027] Figure 25 This is a flowchart of a decoding method for a target block according to an embodiment.
[0028] Figure 26 Examples of the first input structure and the second input structure are shown.
[0029] Figure 27 It is a diagram used to describe a specific location.
[0030] Figure 28 The reference region used to derive the mapping model is shown.
[0031] Figure 29 This shows a sample of the computational output of the mapping model.
[0032] Figure 30 It is a diagram used to describe the target block and the neighboring samples within BOUND_LINE_NUM rows / columns, as well as the samples within BOUND_LINE_NUM rows / columns of the target block.
[0033] Figure 31 An example is shown that determines whether to use top neighbor samples when configuring various sample sets based on at least one of the neighbor samples in the two top rows / columns of the target block or the samples in the two rows / columns within the target block.
[0034] Figure 32 This is a diagram showing the non-adjacent blocks of the current block.
[0035] Figure 33 A filter based on the position of the luminance component is shown.
[0036] Figure 34 An example of a 3x3 low-frequency band filter used for predictive sample correction is shown.
[0037] Figure 35 This is a reference diagram for easy description of the mapping model function.
[0038] Figure 36 The first luminance component sample and neighboring samples are shown.
[0039] Figure 37A first embodiment for generating a reference block when a motion vector (or block vector) indicates the position of a sub-pixel is shown.
[0040] Figure 38 A second embodiment for generating a reference block when a motion vector (or block vector) indicates the position of a sub-pixel is shown. Detailed Implementation
[0041] Because this disclosure can be modified in various ways and has multiple embodiments, specific embodiments are illustrated in the accompanying drawings, and are described in detail in the specific embodiments. However, this is not intended to limit this disclosure to the specific embodiments and should be understood to include all modifications, equivalents, and substitutions included within the concept and scope of this disclosure. Similar reference numerals in the drawings refer to the same or similar functions across multiple aspects. For clarity of description, the shapes and sizes of elements in the drawings may be exaggerated. The detailed description of exemplary embodiments described below refers to the accompanying drawings illustrating specific embodiments as examples. These embodiments are described in detail so that those skilled in the art can implement the embodiments. It should be understood that the various embodiments differ from one another, but they need not be mutually exclusive. For example, the particular shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the scope and spirit of this disclosure relating to the embodiments. Additionally, it should be understood that the position or arrangement of various elements in each disclosed embodiment may be changed without departing from the scope and spirit of the embodiments. Therefore, the detailed description described below is not to be considered limiting, and the scope of the exemplary embodiments is limited only by the appended claims and any scope equivalent to the scope claimed by those claims, if appropriately described.
[0042] In this disclosure, terms such as first, second, etc., may be used to describe various elements, but these elements should not be limited by these terms. These terms are used only to distinguish one element from other elements. For example, without departing from the scope of this disclosure, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. The term "and / or" includes a combination of multiple related descriptive items or any one of multiple related descriptive items.
[0043] When an element in this disclosure is referred to as "connected" or "linked" to another element, it should be understood that it may be directly connected or linked to that other element, but there may be another element between them. Similarly, when an element is referred to as "directly connected" or "directly linked" to another element, it should be understood that there is no other element between them.
[0044] Although the constituent units shown in the embodiments of this disclosure are illustrated independently to represent different functional features, this does not mean that each constituent unit is included in a separate hardware or software constituent unit. In other words, since each constituent unit is listed as a constituent unit for ease of description, at least two constituent units of each constituent unit can be combined to form a constituent unit, or a constituent unit can be divided into multiple constituent units to perform functions, and integrated embodiments and individual embodiments of each constituent unit are also included within the scope of this disclosure, unless they exceed the nature of this disclosure.
[0045] The terminology used in this disclosure is for describing particular embodiments only and is not intended to limit the disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this disclosure, it should be understood that terms such as “comprising” or “having” are intended only to specify the presence of features, numbers, steps, operations, elements, components, or combinations thereof described in this specification, and do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, elements, components, or combinations thereof. In other words, the description of “comprising” a particular configuration in this disclosure does not exclude configurations other than the corresponding configuration, and this means that additional configurations may be included within the scope of the technical ideas or embodiments of this disclosure.
[0046] Some elements of this disclosure are not essential for performing the basic functions of this disclosure and may be optional elements used only to improve performance. This disclosure may be implemented by including only the constituent units necessary to realize the essence of this disclosure, excluding elements used only for performance improvement, and a structure that includes only essential elements and excludes optional elements used only for performance improvement is also included within the scope of the claims of this disclosure.
[0047] In the following, embodiments of the present disclosure are described in detail with reference to the accompanying drawings. When describing embodiments of this specification, detailed descriptions of configurations or functions of the relevant disclosures are omitted where it is determined that such detailed descriptions may obscure the gist of the specification, and the same reference numerals are used for the same elements in the drawings, and repeated descriptions of the same elements are omitted.
[0048] In the following text, an image can refer to a frame of a video, and can also refer to the video itself. For example, "encoding and / or decoding of an image" can refer to "encoding and / or decoding of a video," and can refer to "encoding and / or decoding of one of the images in a video."
[0049] In the following text, the terms "moving image" and "video" may be used with the same meaning and may be used interchangeably.
[0050] In the following text, the target image can be an image to be encoded (the target of encoding) and / or an image to be decoded (the target of decoding). Furthermore, the target image can be an input image to the encoding device or an input image to the decoding device. Here, the target image can have the same meaning as the current image.
[0051] In the following text, the terms “image,” “picture,” “frame,” and “screen” may be used with the same meaning and may be used interchangeably.
[0052] In the following text, a target block can be a block to be encoded (the target of encoding) and / or a block to be decoded (the target of decoding). Alternatively, a target block can be the current block that is currently the target of encoding and / or decoding. For example, the terms "target block" and "current block" can be used with the same meaning and can be used interchangeably.
[0053] In the following text, the terms “block” and “unit” may be used with the same meaning and may be used interchangeably. Optionally, “block” may refer to a specific unit.
[0054] In the following text, the terms “region” and “segment” are used interchangeably.
[0055] In the following text, a specific signal can be a signal representing a specific block. For example, the original signal can be a signal representing the target block. The prediction signal can be a signal representing the prediction block. The residual signal can be a signal representing the residual block.
[0056] In this embodiment, each of the specified information, data, flags, indexes, elements, attributes, etc., may have a value. The value "0" for information, data, flags, indexes, elements, attributes, etc., may represent logical false or a first predefined value. In other words, the values "0", false, logical false, and the first predefined value are interchangeable. The value "1" for information, data, flags, indexes, elements, attributes, etc., may represent logical true or a second predefined value. In other words, the values "1", true, logical true, and the second predefined value are interchangeable.
[0057] When using variables such as i or j to represent rows, columns, or indices, the value of i can be an integer greater than or equal to 0, or an integer greater than or equal to 1. In other words, in the implementation, rows, columns, indices, etc., can be counted starting from 0, or they can be counted starting from 1.
[0058] Encoder: It refers to a device that performs encoding. In other words, it can refer to an encoding device.
[0059] Decoder: This refers to the device that performs decoding. In other words, it can refer to the decoding device itself.
[0060] A block is an M×N array of samples. Here, M and N can refer to positive integer values, and a block typically refers to a two-dimensional array of samples. A block can refer to a cell. The current block can refer to the block to be encoded (the encoding target in encoding) or the block to be decoded (the decoding target in decoding). Furthermore, the current block can be at least one of an encoded block, a prediction block, a residual block, or a transform block. In particular, the shape of the block can include not only squares but also geometric shapes that can be represented in two dimensions, such as rectangles, trapezoids, triangles, pentagons, etc. In addition, block information can include at least one of the following: block type, block size, block depth, or block encoding or decoding order, indicating the type of encoded block, prediction block, residual block, transform block, etc.
[0061] Sample: It is the basic unit of the configuration block. Depending on the bit depth (Bd), it can be represented as a value from 0 to 2Bd-1. In this disclosure, a sample can be used with the same meaning as a pixel or pixel point. In other words, a sample, pixel, and pixel point can have the same meaning.
[0062] Unit: This can refer to a unit in image encoding and decoding. In image encoding and decoding, a unit can be a region of an image that has been divided. Furthermore, a unit can refer to the partitioning unit when an image is divided into segmented units and encoded or decoded. In other words, an image can be divided into multiple units. In image encoding and decoding, predefined processing can be performed on each unit. A unit can be further divided into sub-units with smaller sizes compared to a single unit. Depending on the function, a unit can refer to a block, macroblock, coding tree unit, coding tree block, coding unit, coding block, prediction unit, prediction block, residual unit, residual block, transform unit, transform block, etc. Additionally, to distinguish a unit from a block, a unit can mean that it includes a luma component block, a corresponding chroma component block, and syntax elements for each block. Units can have various sizes and shapes, and in particular, the shape of a unit can include not only squares but also geometric shapes that can be represented in two dimensions, such as rectangles, trapezoids, triangles, pentagons, etc. In addition, the cell information may include at least one of the cell type, cell size, cell depth, or cell encoding or decoding order of the coding cell, prediction cell, residual cell, transform cell, etc.
[0063] A coding tree unit (CRU) consists of a luminance component (Y) coding tree block and two associated chrominance component (Cb, Cr) coding tree blocks. Alternatively, it can refer to the blocks it comprises and the syntax elements for each block. Each CRU can be partitioned using at least one partitioning method such as a quadtree, binary tree, etc., to configure sub-units such as coding units, prediction units, transform units, etc. It can be used as a term to refer to sample blocks that become processing units in image decoding / encoding processing, such as the partitioning of an input image.
[0064] Encoding Tree Block: Encoding tree block can be used as a term to refer to any one of the Y encoding tree block, Cb encoding tree block, or Cr encoding tree block.
[0065] Neighboring block: This can refer to a block adjacent to the current block. A block adjacent to the current block can be a block directly next to the current block or a block located within a predetermined distance from the current block. A neighboring block can also refer to a block adjacent to the vertex of the current block. Here, a block adjacent to the vertex of the current block can be a block vertically adjacent to a block horizontally adjacent to the current block, or a block horizontally adjacent to a block vertically adjacent to the current block. A neighboring block can also refer to a reconstructed neighboring block.
[0066] Reconstructing neighboring blocks: This can refer to neighboring blocks around the current block that have already been spatially / temporally encoded or decoded. In this case, reconstructing neighboring blocks can refer to reconstructing neighboring units. Reconstructing spatial neighboring blocks can be blocks in the current frame, or blocks that have already been reconstructed through encoding and / or decoding. Reconstructing temporally neighboring blocks can be reconstructed blocks or their neighboring blocks at the location corresponding to the current block in the current frame within a reference image.
[0067] Element Depth: This refers to the degree to which elements are divided. In a tree structure, the highest node (root node) corresponds to the initial undivided elements. The highest node can be called the root node. Additionally, the highest node can have the minimum depth value. In this case, the highest node can have a depth of level 0. A node with a depth of level 1 represents elements generated when the initial elements are divided once. A node with a depth of level 2 represents elements generated when the initial elements are divided twice. A node with a depth of level n represents elements generated when the initial elements are divided n times. Leaf nodes can be the lowest nodes and can be nodes that do not need to be further divided. The depth of leaf nodes can be at the maximum level. For example, the predefined value for the maximum level can be 3. It can be said that the root node has the shallowest depth, and the leaf nodes have the deepest depth. Furthermore, when elements are represented in a tree structure, the level at which the elements exist can refer to the element depth.
[0068] Bitstream: It can refer to a series of bits that includes encoded image information.
[0069] Parameter set: This corresponds to the header information in the structure of the bitstream. At least one of a video parameter set, sequence parameter set, frame parameter set, or adaptive parameter set may be included in the parameter set. Additionally, the parameter set may contain strip header information and parallel block header information.
[0070] Explanation: It can refer to the bitstream being entropy-decoded to determine the values of syntax elements, or it can refer to entropy decoding itself.
[0071] Symbol: It can refer to at least one of the syntax elements, encoding parameters, transform coefficient values, etc., of the unit to be encoded / decoded. In addition, the symbol can refer to the target of entropy encoding or the result of entropy decoding.
[0072] Prediction mode: Prediction mode can be information indicating the mode of intra-frame prediction encoding / decoding or the mode of inter-frame prediction encoding / decoding.
[0073] Prediction unit: This refers to the basic unit used in performing predictions such as inter-frame prediction, intra-frame prediction, inter-frame compensation, intra-frame compensation, and motion compensation. A prediction unit can be divided into multiple partitions or smaller sub-prediction units. Multiple partitions can also be the basic unit for performing prediction or compensation. Partitions generated by dividing prediction units can also be prediction units.
[0074] Prediction unit partitioning: This can refer to the way prediction units are divided.
[0075] Reference image list: This can refer to a list that includes at least one reference image for inter-frame prediction or motion compensation. The type of reference image list can include list combination (LC), list 0 (L0), list 1 (L1), list 2 (L2), list 3 (L3), etc., and at least one reference image list can be used for inter-frame prediction.
[0076] Inter-frame prediction indicator: This can indicate the inter-frame prediction direction of the current block (unidirectional prediction, bidirectional prediction, etc.). Optionally, it can indicate the number of reference images used when generating the prediction block for the current block. Optionally, it can indicate the number of prediction blocks used when performing inter-frame prediction or motion compensation on the current block.
[0077] Prediction list utilization flag: Indicates whether a prediction block is generated by using at least one reference image from a specific reference image list. The inter-frame prediction indicator can be derived using the prediction list utilization flag, and conversely, the prediction list utilization flag can be derived using the inter-frame prediction indicator. For example, if the prediction list utilization flag indicates a first value of 0, it indicates that a prediction block is not generated by using a reference image from the corresponding reference image list; if the prediction list utilization flag indicates a second value of 1, it indicates that a prediction block can be generated by using the corresponding reference image list.
[0078] Reference image index: It can refer to the index of a specific reference image in the list of reference images.
[0079] Reference frame: This can refer to an image referenced by a specific block for inter-frame prediction or motion compensation. Optionally, a reference image can be an image that includes a reference block referenced by the current block for inter-frame prediction or motion compensation. In the following text, the terms "reference frame" and "reference image" may be used with the same meaning and are interchangeable.
[0080] Motion vector: This can be a two-dimensional vector used for inter-frame prediction or motion compensation. A motion vector can refer to the offset between the coded / decoded target block and the reference block. For example, (mvX, mvY) can represent a motion vector. mvX can represent the horizontal component, and mvY can represent the vertical component.
[0081] Search Range: The search range can be a two-dimensional region used to search for motion vectors during inter-frame prediction. For example, the size of the search range can be M×N. M and N can both be positive integers. Furthermore, the shape of the search range can include not only squares, but also geometric shapes that can be represented in two dimensions, such as rectangles, trapezoids, triangles, pentagons, etc.
[0082] Motion vector candidates: Motion vector candidates can refer to blocks or motion vectors of blocks that become prediction candidates when predicting motion vectors. Furthermore, motion vector candidates can be included in a motion vector candidate list.
[0083] Motion vector candidate list: It can refer to a list configured using at least one motion vector candidate.
[0084] Motion vector candidate index: This can be an indicator pointing to a motion vector candidate in the motion vector candidate list. It can also be an index of a motion vector predictor.
[0085] Motion information: Motion information may refer to information including at least one of motion vectors, reference image indexes, inter-frame prediction indicators, prediction list utilization flags, reference image list information, reference images, motion vector candidates, motion vector candidate indices, merge candidates, or merge indices.
[0086] Merge candidate list: It can refer to a list configured by using at least one merge candidate.
[0087] Merging candidates: This can refer to spatial merging candidates, temporal merging candidates, combined merging candidates, combined bidirectional prediction merging candidates, zero merging candidates, etc. Merging candidates can include motion information such as inter-frame prediction indicators, reference image indices for each list, motion vectors, prediction list utilization flags, inter-frame prediction indicators, etc.
[0088] Merge Index: This can refer to an indicator that points to a merge candidate in the merge candidate list. Furthermore, the merge index can indicate a block that is reconstructed as a deduced merge candidate among blocks spatially / temporally adjacent to the current block. Additionally, the merge index can indicate at least one of the motion information of the merge candidates.
[0089] Transform unit: This can refer to the basic unit used in performing residual signal encoding / decoding (such as transform, inverse transform, quantization, dequantization, and transform coefficient encoding / decoding). A transform unit can be divided into multiple sub-transform units with smaller dimensions. Here, the transform / inverse transform may include at least one of the primary transform / inverse transform or the secondary transform / inverse transform.
[0090] Scaling: This refers to the process of multiplying the levels of transform coefficients by a factor. Transform coefficients can be generated as a result of scaling the levels of transform coefficients. Scaling can also be called inverse quantization.
[0091] Quantization parameter: This can refer to the value used when generating the transform coefficient levels during quantization. Optionally, it can also refer to the value used when generating the transform coefficients by scaling the transform coefficient levels during dequantization. The quantization parameter can be a value mapped to the quantization step size.
[0092] Residual quantization parameter (Delta quantization parameter): The residual quantization parameter refers to the difference between the predicted quantization parameter and the quantization parameter of the encoding / decoding target unit.
[0093] Scan: This can refer to a method used to sort the coefficients in cells, blocks, or matrices. For example, sorting a two-dimensional array into a one-dimensional array is called a scan. Similarly, sorting a one-dimensional array into a two-dimensional array can also be called a scan or inverse scan.
[0094] Transform coefficients: These can refer to the coefficient values generated after a transform is performed in the encoder. Alternatively, they can refer to the coefficient values generated after at least one of entropy decoding or dequantization is performed in the decoder. The level of quantization or the level of quantized transform coefficients obtained by applying quantization to the transform coefficients or residual signal can also be included in the meaning of transform coefficients.
[0095] Quantization level: The quantization level can refer to the value generated by performing quantization on the transform coefficients or residual signal in the encoder. Optionally, it can refer to the value to be dequantized before performing dequantization in the decoder. Similarly, the level of the quantized transform coefficients (the result of transform and quantization) can also be included in the meaning of quantization level.
[0096] Non-zero transform coefficients: These can refer to transform coefficients whose values are not zero or transform coefficient levels whose values are not zero.
[0097] Quantization matrix: This refers to a matrix used during quantization or dequantization to improve the subjective or objective image quality. The quantization matrix can also be called a scaling list.
[0098] Quantization matrix coefficients: These can refer to each element in the quantization matrix. Quantization matrix coefficients can also be called matrix coefficients.
[0099] Default matrix: It can refer to a predefined quantization matrix in the encoder and decoder.
[0100] Non-default matrix: This can refer to a quantization matrix that is not predefined in the encoder and decoder and is sent by the user via a signal.
[0101] Statistical value: A statistical value of a variable, coding parameter or constant that has a specific value that can be calculated can be at least one of the mean, weighted average, weighted sum, minimum, maximum, mode, median or interpolation of the corresponding specific value.
[0102] Neighboring blocks: Neighboring blocks refer to blocks adjacent to the target block. Neighboring blocks can include spatially neighboring blocks and temporally neighboring blocks. Neighboring blocks can also refer to reconstructed neighboring blocks in the reference image. A neighboring block is not necessarily adjacent to the target block.
[0103] Spatial neighbor blocks: Spatial neighbor blocks can be blocks that are spatially adjacent to the target block. The target image may include both the target block and spatial neighbor blocks. Spatial neighbor blocks may include blocks whose boundaries at least a portion of their boundaries are in contact with at least a portion of the boundary of the target block. Optionally, spatial neighbor blocks may include blocks whose distance from the target block is less than or equal to a reference value. Spatial neighbor blocks may include blocks that are diagonally adjacent to the vertices of the target block. Spatial neighbor blocks may include the top-left block adjacent to the top-left of the target block, the top block adjacent to the top of the target block, the top-right block adjacent to the top-right of the target block, the left-side block adjacent to the left side of the target block, the right-side block adjacent to the right side of the target block, the bottom-left block adjacent to the bottom-left of the target block, the bottom block adjacent to the bottom of the target block, and the bottom-right block adjacent to the bottom-right of the target block.
[0104] Temporally Proximity Blocks: Temporally proximity blocks can be blocks that are temporally adjacent to the target block. Temporally proximity blocks can include co-location blocks (col blocks). Co-location blocks can be blocks in the reconstructed image within the reference image buffer. Co-location frames (col frames) can refer to images that include co-location blocks. Co-location frames can be images included in the reference image list. Co-location blocks can be determined based on the position of the target block in the target image. When two blocks are "temporally adjacent," this can mean that the positions of the two blocks meet a specific condition. The position of a co-location block in a co-location image can be the same as the position of the target block in the target image. Optionally, the position of a co-location block in a co-location image can correspond to the position of the target block in the target image. Here, when the positions of the blocks correspond, it can mean that the blocks have the same area, the area of one block is included in the area of another block, or one block occupies a specific position in another block. For example, the position of a co-location block in a co-location image can be the same as the position of the target block in the target image. Optionally, a co-location block can be a block that includes co-location pixels in the co-location image. A co-location pixel can be a pixel with the same coordinates as a specific pixel in the target block. A temporally neighboring block can be a block that is spatially adjacent to the target block in time.
[0105] Neighboring samples: Neighboring samples can refer to samples within neighboring blocks. Neighboring samples can include predicted samples, reconstructed samples, residual samples, and decoded samples.
[0106] The terms listed below may be used with the same meaning in the embodiments and may be used interchangeably in the embodiments.
[0107] "Motion Information", "Motion Vector", "Block Vector" "Bidirectional prediction", "Bidirectional prediction", "Inter-frame dual prediction", "Inter-frame bidirectional prediction" "IBC mode", "Intra-block copy mode", "IBC", "Intra-block copy" Predefined value: A predefined value can refer to a value commonly used in both the encoding and decoding devices. For example, a predefined value can be interpreted as being restricted to a fixed value. Optionally, a predefined value can be a value shared between the encoding and decoding devices via signaling. Optionally, a predefined value can be a value derived through the same process in both the encoding and decoding devices, such that the encoding and decoding devices have a common value. Optionally, a predefined value can be a common value maintained by both the encoding and decoding devices. Values derived through the same process in both the encoding and decoding devices can include values derived through the same process for the same value and / or the same information in both the encoding and decoding devices. Values derived through the same process in both the encoding and decoding devices can include values derived by using the same conditional statements for the same value and / or the same information in both the encoding and decoding devices. The description of predefined values can also apply to predefined information. In the description, "value" can be replaced by "information".
[0108] Motion information: Motion information may refer to information including at least one of the following: reference frame list information, reference image, motion vector candidate, motion vector candidate index, merge candidate and merge index, CPMV, block vector, block vector candidate or block vector candidate index, and motion vector, reference frame index and inter-frame prediction indicator.
[0109] In this disclosure, "when the indicator indicating whether a specific method is used is true" can mean whether the following items are true: prediction mode; motion information; encoding parameters and / or the position indicated by the corresponding indicator. For example, the indicator indicating whether a specific mode is executed can have values from 0 to 3, and the specific mode can only be executed if the indicator has a value of 1 or 3. In this case, when the indicator indicating whether a specific mode is executed is true, it means that the indicator indicating whether a specific mode is executed has a value of 1 or 3.
[0110] In this disclosure, "when the indicator indicating whether a particular method is used is false" means that the indicator indicating whether a particular method is used is not true.
[0111] In an embodiment, rearrangement for a specific target may refer to sorting the specific target or the elements within the specific target.
[0112] In this disclosure, "when the indicator indicating whether a specific method (or mode) is performed on the target block is true" and "when a specific method is performed on the target block" can be used with the same meaning and can be used interchangeably.
[0113] In this disclosure, "when the indicator indicating whether a particular method (or pattern) is performed on the target block is false" and "when a particular method is not performed on the target block" can be used with the same meaning and can be used interchangeably.
[0114] In this disclosure, the matching cost may include template matching cost and bilateral matching cost.
[0115] In this disclosure, “sample (of a block or reference area or template)” may refer to all samples in a block or reference area or template, or may refer to at least one sample in a block or reference area or template.
[0116] Figure 1 This is a block diagram illustrating the configuration of an encoding apparatus according to an embodiment of the present disclosure.
[0117] The encoding device 100 may be an encoder, a video encoding device, or an image encoding device. The video may include at least one image. The encoding device 100 may encode at least one image sequentially.
[0118] refer to Figure 1 The encoding device 100 may include a motion prediction unit 111, a motion compensation unit 112, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.
[0119] The encoding device 100 can encode the input image according to intra-frame mode and / or inter-frame mode. Furthermore, the encoding device 100 can generate a bitstream including encoding information by encoding the input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or streamed via a wired / wireless transmission medium. When intra-frame mode is used as the prediction mode, switch 115 can be switched to intra-frame mode, and when inter-frame mode is used as the prediction mode, switch 115 can be switched to inter-frame mode. Here, intra-frame mode can refer to intra-frame prediction mode, and inter-frame mode can refer to inter-frame prediction mode. The encoding device 100 can generate a prediction block of the input block of the input image. Furthermore, after generating the prediction block, the encoding device 100 can encode the residual block using the residual between the input block and the prediction block. The input image can be referred to as the current image currently used as the encoding target. The input block can be referred to as the current block currently used as the encoding target or the encoding target block.
[0120] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use samples from blocks that have already been encoded / decoded around the current block as reference samples. The intra-frame prediction unit 120 can perform spatial prediction on the current block using the reference samples, and generate prediction samples for the input block through spatial prediction. Here, intra-frame prediction may refer to intra-frame prediction.
[0121] When the prediction mode is inter-frame mode, the motion prediction unit 111 can search for the region that best matches the input block from the reference image during motion prediction processing, and derive the motion vector using the searched region. In this case, the search range can be used as the region. The reference image can be stored in the reference frame buffer 190. Here, the reference image can be stored in the reference frame buffer 190 when encoding / decoding the reference image.
[0122] The motion compensation unit 112 can generate a prediction block for the current block by performing motion compensation using motion vectors. Here, inter-frame prediction can refer to inter-frame prediction or motion compensation.
[0123] When the values of the motion vectors are not integers, the motion prediction unit 111 and the motion compensation unit 112 can generate prediction blocks by applying interpolation filters to some regions in the reference image. To perform inter-frame prediction or motion compensation, the motion prediction and motion compensation method of the prediction units included in the respective coding units can be determined based on the coding unit: skip mode, merge mode, advanced motion vector prediction (AMVP) mode, or current frame reference mode, and inter-frame prediction or motion compensation can be performed according to each mode.
[0124] Subtractor 125 generates residual blocks by using the difference between the input block and the prediction block. The residual block can also be referred to as the residual signal. The residual signal can refer to the difference between the original signal and the prediction signal. Optionally, the residual signal can be transformed or quantized, or it can be a signal generated by transforming or quantizing the difference between the original signal and the prediction signal. The residual block can be a residual signal on a block-by-block basis.
[0125] Transform unit 130 can generate transform coefficients by performing a transform on the residual block and output the generated transform coefficients. Here, the transform coefficients can be coefficient values generated by performing a transform on the residual block. When a transform skip mode is applied, transform unit 130 can skip the transform on the residual block.
[0126] Quantization levels can be generated by applying quantization to the transform coefficients or the residual signal. In the following examples, the quantization level may also be referred to as the transform coefficient.
[0127] The quantization unit 140 can generate a quantization level by quantizing the transform coefficients or residual signal according to the quantization parameters, and output the generated quantization level. In this case, the quantization unit 140 can quantize the transform coefficients using a quantization matrix.
[0128] The entropy coding unit 150 can generate and output a bitstream by performing entropy coding on values calculated in the quantization unit 140 or coding parameter values calculated in the coding process according to a probability distribution. The entropy coding unit 150 can perform entropy coding on information about image samples and information used for decoding the image. For example, the information used for decoding the image may include syntax elements, etc.
[0129] When applying entropy coding, a small number of bits can be allocated to symbols with high occurrence probabilities, while a large number of bits can be allocated to symbols with low occurrence probabilities, thereby reducing the size of the bitstream of symbols to be encoded. The entropy coding unit 150 can perform entropy coding using coding methods such as Exponential Golomb, CAVLC (Context Adaptive Variable Length Coding), and CABAC (Context Adaptive Binary Arithmetic Coding). For example, the entropy coding unit 150 can perform entropy coding by using a variable-length code (VLC) table. Alternatively, the entropy coding unit 150 can derive a binarization method for the target symbol and a probability model for the target symbol / bits, and then use the derived binarization method, probability model, and context model to perform arithmetic coding.
[0130] The entropy coding unit 150 can transform the two-dimensional block form coefficients into a one-dimensional vector form through the transform coefficient scanning method to encode the transform coefficient level.
[0131] Encoding parameters can include information derived from the encoding or decoding process, as well as information encoded in the encoder and signaled to the decoder like syntax elements (flags, indexes, etc.), and can refer to information required when encoding or decoding an image. For example, encoding parameters can include at least one value or combination of the following: cell / block size, cell / block depth, cell / block partitioning information, cell / block partitioning structure, quadtree partitioning, binary tree partitioning, binary tree partitioning direction (horizontal or vertical), binary tree partitioning form (symmetric or asymmetric), prediction mode (intra-frame prediction or inter-frame prediction), intra-frame prediction mode / direction, reference sample filtering method, reference sample filter taps, reference sample filter coefficients, prediction block filtering method, prediction block filter taps, prediction block filter coefficients, prediction... Block boundary filtering method, predicted block boundary filter taps, predicted block boundary filter coefficients, inter-frame prediction mode, motion information, motion vectors, reference image index, inter-frame prediction direction, inter-frame prediction indicator, prediction list utilization flags, reference image list, reference image, motion vector prediction candidates, motion vector candidate list, whether to use merging mode, merge candidates, merge candidate list, whether to use skip mode, interpolation filter type, interpolation filter taps, interpolation filter coefficients, motion vector magnitude, motion vector representation accuracy, transform type, transform magnitude, information on whether to use primary transform, and information on whether to use primary transform. Information on whether to use secondary transform, primary transform index, secondary transform index, information on the presence of residual signal, code block style, code block flag, quantization parameters, quantization matrix, whether to apply in-loop filter, in-loop filter coefficients, in-loop filter taps, in-loop filter shape / form, whether to apply deblocking filter, deblocking filter coefficients, deblocking filter taps, deblocking filter strength, deblocking filter shape / form, whether to apply adaptive sample offset, adaptive sample offset value, adaptive sample offset category, adaptive sample offset type, whether to apply adaptive loop filter, adaptive loop... Filter coefficients, adaptive loop filter taps, adaptive loop filter shape / form, binarization / debinarization method, context model determination method, context model update method, whether to execute normal mode, whether to execute bypass mode, context binary bits, bypass binary bits, transform coefficients, transform coefficient level, quantization level, transform coefficient level scanning method, image display / output order, stripe identification information, stripe type, stripe division information, parallel block identification information, parallel block type, parallel block division information, frame type, bit depth, information about luminance signal or information about chrominance signal.
[0132] Here, sending a flag or index with a signal can indicate that the corresponding flag or index in the encoder is entropy encoded and included in the bit stream, and can also indicate that the corresponding flag or index in the decoder is entropy decoded from the bit stream.
[0133] When the encoding device 100 performs encoding via inter-frame prediction, the encoded current image can be used as a reference image for another image to be processed later. Therefore, the encoding device 100 can reconstruct or decode the current encoded image and store the reconstructed or decoded image as a reference image in the reference frame buffer 190.
[0134] The quantization level can be dequantized in the dequantization unit 160 and inverse transformed in the inverse transform unit 170. The coefficients of the dequantization and / or inverse transform can be combined with the prediction block using adder 175. A reconstruction block can be generated by combining the coefficients of the dequantization and / or inverse transform with the prediction block. Here, the coefficients of the dequantization and / or inverse transform refer to the coefficients of performing at least one of dequantization or inverse transform, and may refer to the reconstruction residual block.
[0135] The reconstructed block can be passed through filter unit 180. Filter unit 180 can apply at least one of deblocking filter, sample adaptive offset (SAO), adaptive loop filter (ALF), etc., to the reconstructed sample, reconstructed block, or reconstructed image. Filter unit 180 may also be referred to as an in-loop filter.
[0136] Deblocking filters remove block distortion generated at the boundaries between blocks. To determine whether to perform a deblocking filter, samples from several columns or rows included in the block can be used to determine whether to apply the deblocking filter to the current block. When a deblocking filter is applied to a block, different filters can be applied depending on the desired deblocking filtering intensity.
[0137] Sample-adaptive offset can be used to add an appropriate offset value to the sample value to compensate for coding errors. Sample-adaptive offset corrects the offset from the original image on a sample-by-sample basis for the image undergoing deblocking. Methods can be used to divide the samples included in the image into a specific number of regions, determine the regions to be offset, and apply the offset to the corresponding regions, or to apply the offset by considering the edge information of each sample.
[0138] An adaptive loop filter can perform filtering based on values obtained by comparing the reconstructed image with the original image. After dividing the samples included in the image into predetermined groups, filtering can be performed differently for each group by determining the filter to be applied to the corresponding group. Information related to whether an adaptive loop filter is applied can be sent to each coding unit (CU) via a signal, and the shape and filter coefficients of the adaptive loop filter to be applied can vary according to each block.
[0139] The reconstructed blocks or reconstructed image obtained through filter unit 180 can be stored in reference frame buffer 190. The reconstructed blocks obtained through filter unit 180 can be a portion of the reference image. In other words, the reference image can be a reconstructed image composed of the reconstructed blocks obtained through filter unit 180. Subsequently, the stored reference image can be used for inter-frame prediction or motion compensation.
[0140] Figure 2 This is a block diagram illustrating the configuration of an embodiment of the decoding apparatus according to the present disclosure.
[0141] The decoding device 200 can be a decoder, a video decoding device, or an image decoding device.
[0142] refer to Figure 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260, and a reference frame buffer 270.
[0143] The decoding device 200 can receive a bitstream output from the encoding device 100. The decoding device 200 can receive a bitstream stored in a computer-readable recording medium or a bitstream transmitted via a wired / wireless transmission medium. The decoding device 200 can perform decoding on the bitstream in intra-frame mode or inter-frame mode. Furthermore, the decoding device 200 can generate a reconstructed image or a decoded image through decoding and output the reconstructed image or the decoded image.
[0144] If the prediction mode used for decoding is intra-frame mode, the switch can be toggled to intra-frame mode. If the prediction mode used for decoding is inter-frame mode, the switch can be toggled to inter-frame mode.
[0145] The decoding device 200 can decode the input bitstream to obtain a reconstructed residual block and generate a prediction block. When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block to be decoded by adding the reconstructed residual block and the prediction block. The block to be decoded can be referred to as the current block.
[0146] The entropy decoding unit 210 can generate symbols by performing entropy decoding according to the probability distribution of the bitstream. The generated symbols may include symbols in the form of quantization levels. Here, the entropy decoding method can be the inverse process of the entropy encoding method described above.
[0147] The entropy decoding unit 210 can transform the one-dimensional vector form coefficients into two-dimensional block form through the transform coefficient scanning method, so as to decode the transform coefficient level.
[0148] The quantization level can be dequantized in the dequantization unit 220 and inverse transformed in the inverse transform unit 230. The quantization level is the result of performing dequantization and / or inverse transform and can be generated as a reconstruction residual block. In this case, the dequantization unit 220 can apply the quantization matrix to the quantization level.
[0149] When using intra-frame mode, intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction on the current block using sample values of already decoded blocks surrounding the block to be decoded.
[0150] When inter-frame mode is used, motion compensation unit 250 can generate a prediction block by performing motion compensation on the current block using motion vectors and a reference image stored in reference frame buffer 270. When the value of the motion vector is not an integer, motion compensation unit 250 can generate a prediction block by applying an interpolation filter to a portion of the reference image. To perform motion compensation, the motion compensation method for the prediction unit included in the corresponding coding unit can be determined based on the coding unit: skip mode, merge mode, AMVP mode, or current frame reference mode, and motion compensation can be performed according to each mode.
[0151] Adder 255 generates a reconstructed block by adding the reconstructed residual block and the prediction block. Filter unit 260 can apply at least one of a deblocking filter, a sample adaptive offset, or an adaptive loop filter to the reconstructed block or reconstructed image. Filter unit 260 can output a reconstructed image. The reconstructed block or reconstructed image can be stored in a reference frame buffer 270 and used for inter-frame prediction. The reconstructed block passed through filter unit 260 can be a portion of the reference image. In other words, the reference image can be a reconstructed image composed of the reconstructed blocks passed through filter unit 260. Subsequently, the stored reference image can be used for inter-frame prediction or motion compensation.
[0152] Figure 3 It is a diagram that roughly illustrates the partitioning structure of an image during encoding and decoding. Figure 3 An embodiment of dividing a unit into multiple sub-units is shown in outline.
[0153] To efficiently partition an image, coding units (CUs) can be used in encoding and decoding. A coding unit can serve as the basic unit for image encoding / decoding. Furthermore, during image encoding / decoding, a coding unit can be used to distinguish between intra-frame prediction modes and inter-frame prediction modes. A coding unit can also be the basic unit for processing transform coefficients, including prediction, transform, quantization, inverse transform, inverse quantization, or encoding / decoding.
[0154] Reference Figure 3The image 300 is divided sequentially in units of Maximum Coding Units (LCUs), and the partitioning structure is determined in units of LCUs. Here, LCU can be used in the same sense as Coding Tree Unit (CTU). Unit partitioning can refer to the partitioning of blocks corresponding to the unit. Block partitioning information can include information about the depth of the unit. Depth information can represent the number and / or extent of unit partitioning. A unit can be hierarchically partitioned into multiple sub-units with depth information based on a tree structure. In other words, a unit and the sub-units generated by partitioning that unit can correspond to a node and its child nodes, respectively. Each partitioned sub-unit can have depth information. Depth information can be information indicating the size of the CU, and depth information can be stored for each CU. Since unit depth represents the number and / or extent of unit partitioning, the sub-unit partitioning information can also include information about the size of the sub-units.
[0155] The partitioning structure refers to the distribution of coding units (CUs) within the LCU 310. This distribution can be determined by whether a CU is divided into multiple CUs (positive integers equal to or greater than 2, including 2, 4, 8, 16, etc.). The horizontal and vertical dimensions of the CUs generated by partitioning can be half the horizontal and vertical dimensions of the CUs before partitioning, respectively, or they can have dimensions smaller than the horizontal and vertical dimensions of the CUs before partitioning, depending on the number of partitions. CUs can be recursively partitioned into multiple CUs. Through recursive partitioning, at least one of the horizontal or vertical dimensions of the partitioned CUs can be reduced compared to at least one of the horizontal or vertical dimensions of the CUs before partitioning. CU partitioning can be performed recursively until a predefined depth or predefined size is reached. For example, the depth of the LCU can be 0, and the depth of the minimum coding unit (SCU) can be a predefined maximum depth. Here, the LCU can be a coding unit with the maximum coding unit size as described above, and the SCU can be a coding unit with the minimum coding unit size. Partitioning starts from the LCU 310, and the depth of the CU increases by 1 each time the horizontal and / or vertical dimensions of a CU decrease due to partitioning. For example, for each depth, an undivided CU can have a size of 2N×2N. Alternatively, for a divided CU, a CU of size 2N×2N can be divided into four CUs of size N×N. The size of N can be halved each time the depth increases by 1.
[0156] Additionally, the partitioning information of a CU can be used to indicate whether a CU has been partitioned. This partitioning information can be 1 bit. All CUs except the SCU can include partitioning information. For example, if the value of the partitioning information is a first value, the CU may not be partitioned, and if the value is a second value, the CU may be partitioned.
[0157] refer to Figure 3 An LCU with a depth of 0 can be a 64×64 block. 0 can be the minimum depth. An SCU with a depth of 3 can be an 8×8 block. 3 can be the maximum depth. CUs with depths of 32×32 and 16×16 can be represented as depth 1 and depth 2, respectively.
[0158] For example, when a coding unit is divided into four coding units, the horizontal and vertical dimensions of the four divided coding units can be half the size of the original coding unit. For example, when a 32×32 coding unit is divided into four coding units, each of the four divided coding units can have a size of 16×16. When a coding unit is divided into four coding units, the coding unit can be said to be divided in the form of a quadtree.
[0159] For example, when a coding unit is divided into two coding units, the horizontal or vertical dimensions of the two divided coding units can be half the size of the original coding unit. As an example, when a 32×32 coding unit is vertically divided into two coding units, each of the two divided coding units can have a size of 16×32. When a coding unit is divided into two coding units, the coding unit can be said to be divided in the form of a binary tree. Figure 3 LCU 320 is an example of an LCU that applies both quadtree-based partitioning and binary tree-based partitioning.
[0160] Figure 4 This is a diagram illustrating an embodiment of intra-frame prediction processing.
[0161] from Figure 4 The arrow from the center to the outside can indicate the prediction direction of the intra-frame prediction mode.
[0162] Intra-frame coding and / or decoding can be performed using reference samples from neighboring blocks of the current block. A neighboring block can be a reconstructed neighboring block. For example, intra-frame coding and / or decoding can be performed using values or coding parameters of reference samples included in the reconstructed neighboring block.
[0163] A prediction block can refer to a block generated as a result of performing intra-frame prediction. A prediction block can correspond to at least one of a CU, PU, or TU. The unit of a prediction block can be the size of at least one of a CU, PU, or TU. A prediction block can be a square block with dimensions of 2x2, 4x4, 16x16, 32x32, 64x64, etc., or a rectangular block with dimensions of 2x8, 4x8, 2x16, 4x16, 8x16, etc.
[0164] Intra-prediction can be performed based on the intra-prediction mode of the current block. The number of intra-prediction modes that the current block can have can be a predefined fixed value, or it can be a value that varies depending on the attributes of the predicted block. For example, the attributes of the predicted block can include the size of the predicted block, the form of the predicted block, etc.
[0165] Regardless of the block size, the number of intra-prediction modes can be fixed at N. Optionally, for example, the number of intra-prediction modes can be 3, 5, 9, 17, 34, 35, 36, 65, or 67, etc. Optionally, the number of intra-prediction modes can vary depending on the block size and / or the type of color components. For example, the number of intra-prediction modes can vary depending on whether the color component is a luma signal or a chroma signal. For example, as the block size increases, the number of intra-prediction modes can increase. Optionally, the number of intra-prediction modes for the luma component block can be greater than the number of intra-prediction modes for the chroma component block.
[0166] Intra-frame prediction modes can be non-directional or directional. Non-directional modes can be DC or planar modes, and directional modes (angle modes) can be prediction modes with a specific direction or angle. Intra-frame prediction modes can be represented as at least one of mode number, mode value, mode number, mode angle, or mode direction. The number of intra-frame prediction modes can be M or greater than 1, including non-directional and directional modes.
[0167] For intra-frame prediction of the current block, it can be checked whether samples included in reconstructed neighboring blocks can be used as reference samples for the current block. If there are samples that cannot be used as reference samples for the current block, they can be used as reference samples for the current block after replacing the sample values of at least one sample contained in the reconstructed neighboring block with the sample values of the samples that cannot be used as reference samples using duplication and / or interpolation.
[0168] During intra-frame prediction, filters can be applied to at least one of the reference samples or the prediction samples based on at least one of the intra-frame prediction mode or the size of the current block.
[0169] For planar mode, when generating the prediction block for the current block, the sample value of the predicted target sample can be generated by using a weighted sum of the top and left reference samples of the current sample, and the top-right and bottom-left reference samples of the current block, based on the position of the predicted target sample within the prediction block. Furthermore, for DC mode, when generating the prediction block for the current block, the average of the top and left reference samples of the current block can be used. Additionally, for directional mode, the prediction block can be generated using the top, left, top-right, and / or bottom-left reference samples of the current block. Interpolation in cells or real numbers can also be performed to generate predicted sample values.
[0170] The intra-prediction mode of the current block can be predicted from the intra-prediction modes of blocks existing around it, and entropy encoding / decoding of the intra-prediction mode of the current block can be performed. If the intra-prediction mode of the current block is the same as that of neighboring blocks, information indicating that the intra-prediction mode of the current block is the same as that of neighboring blocks can be signaled using predetermined flag information. Furthermore, indicator information of intra-prediction modes among multiple neighboring blocks that are the same as the intra-prediction mode of the current block can be signaled. If the intra-prediction mode of the current block is different from that of neighboring blocks, entropy encoding / decoding of the intra-prediction mode information of the current block can be performed based on the intra-prediction modes of neighboring blocks.
[0171] Figure 5 This is a diagram illustrating an embodiment of inter-frame prediction processing.
[0172] Figure 5 The square shown can represent an image. Additionally, in Figure 5 In this context, arrows can also be used to indicate the prediction direction. Each image can be classified into I-frames (intra-frame images), P-frames (predictive images), B-frames (bidirectional prediction images), etc., based on its encoding type.
[0173] I-frames can be encoded / decoded via intra-frame prediction without inter-frame prediction. P-frames can be encoded / decoded using inter-frame prediction with reference images present only in a unilateral direction (e.g., forward or reverse direction). B-frames can be encoded / decoded using inter-frame prediction with reference images present in both unilateral directions (e.g., forward and reverse directions). Furthermore, for B-frames, encoding / decoding can be performed using inter-frame prediction with reference images present in both unilateral directions or with reference images present in one of the forward and reverse directions. Here, the unilateral direction can be forward and reverse. When using inter-frame prediction, the encoder can perform inter-frame prediction or motion compensation, and the decoder can perform the corresponding motion compensation.
[0174] The following describes in detail the inter-frame prediction according to the embodiments.
[0175] Inter-frame prediction or motion compensation can be performed using reference images and motion information.
[0176] Motion information of the current block can be derived by each of the encoding device 100 and the decoding device 200 during inter-frame prediction. Motion information can be derived using motion information of reconstructed neighboring blocks, collocation blocks (col blocks), and / or motion information of blocks adjacent to the collocation block. A collocation block can be a block within a reconstructed collocation frame (col frame) corresponding to the spatial location of the current block. Here, a collocation frame can be one of at least one reference image included in a list of reference images.
[0177] The method used to derive motion information can vary depending on the prediction mode of the current block. For example, prediction modes applied to inter-frame prediction may include AMVP mode, merge mode, skip mode, current frame reference mode, etc. Here, the merge mode can be referred to as motion merge mode.
[0178] For example, when AMVP is applied as a prediction mode, a motion vector candidate list can be generated by identifying at least one of the motion vectors of reconstructed neighboring blocks, motion vectors of co-located blocks, motion vectors of blocks adjacent to co-located blocks, or (0,0) motion vectors as motion vector candidates. Motion vector candidates can be derived using the generated list of motion vector candidates. Motion information for the current block can be determined based on the derived motion vector candidates. Here, the motion vectors of co-located blocks or blocks adjacent to co-located blocks can be referred to as temporal motion vector candidates, and the motion vectors of reconstructed neighboring blocks can be referred to as spatial motion vector candidates.
[0179] Encoding device 100 can calculate the motion vector difference (MVD) between the motion vector of the current block and the motion vector candidates, and entropy encode the MVD. Furthermore, encoding device 100 can generate a bitstream by entropy encoding the motion vector candidate index. The motion vector candidate index indicates the optimal motion vector candidate selected from the motion vector candidates included in the motion vector candidate list. Decoding device 200 can entropy decode the motion vector candidate index from the bitstream, and select a motion vector candidate for the target block from the motion vector candidates included in the motion vector candidate list by using the entropy-decoded motion vector candidate index. Furthermore, decoding device 200 can deduce the motion vector of the target block by summing the entropy-decoded MVD with the motion vector candidate.
[0180] The bitstream may include a reference image index, etc., indicating a reference image. The reference image index may be entropy encoded and transmitted from the encoding device 100 to the decoding device 200 via a signal through the bitstream. The decoding device 200 may generate a predicted block for the decoded target block based on the derived motion vector and reference image index information.
[0181] Another example of a method for deriving motion information is a merging pattern. A merging pattern can refer to merging the motion of multiple blocks. It can also refer to a pattern for deriving the motion information of the current block from the motion information of neighboring blocks. When applying a merging pattern, a list of merging candidates can be generated by using the reconstructed motion information of neighboring blocks and / or the motion information of co-located blocks. Motion information may include at least one of 1) motion vectors, 2) reference image indices, or 3) inter-frame prediction indicators. Prediction indicators can be unidirectional (L0 prediction, L1 prediction) or bidirectional.
[0182] The merge candidate list can represent a list storing motion information. The motion information stored in the merge candidate list can be at least one of the following: motion information of neighboring blocks adjacent to the current block (spatial merge candidate), motion information of blocks in the reference image that are in the same position as the current block (temporal merge candidate), new motion information generated by combining motion information already existing in the merge candidate list, or zero merge candidate.
[0183] Encoding device 100 can generate a bitstream by entropy encoding at least one of a merge flag or a merge index, and then send it as a signal to decoding device 200. The merge flag can be information indicating whether a merge mode is performed for each block, and the merge index can be information about which of the neighboring blocks adjacent to the current block will be merged. For example, the neighboring blocks of the current block can include at least one of the left-hand neighboring block, the top-hand neighboring block, or the time-hand neighboring block.
[0184] Skip mode can be a mode in which motion information of neighboring blocks is applied to the current block as is. When skip mode is used, encoding device 100 can entropy encode information about the motion information of the block that will be used as motion information for the current block and send it as a signal to decoding device 200 via a bitstream. In this case, encoding device 100 may not send syntax elements about at least one of motion vector difference information, encoded block flags, or transform coefficient levels to decoding device 200.
[0185] The current frame reference mode can refer to a prediction mode that uses the pre-reconstructed region within the current frame to which the current block belongs. In this case, a vector can be defined to specify the pre-reconstructed region. Whether the current block is encoded in the current frame reference mode can be determined by using the reference image index of the current block. A flag or index indicating whether the current block is encoded in the current frame reference mode can be sent by a signal, or it can be inferred from the reference image index of the current block. If the current block is encoded in the current frame reference mode, the current frame can be added to a fixed or random position in the reference image list of the current block. For example, a fixed position could be the position with reference image index 0 or the last position. When the current frame is added to a random position in the reference image list, a separate reference image index showing the random position can be sent by a signal.
[0186] Figure 6 It is a diagram used to describe the transformation and quantization processes.
[0187] like Figure 6 As shown, quantization levels can be generated by performing transform and / or quantization processing on the residual signal. The residual signal can be generated as the difference between the original block and the predicted block (intra-frame prediction block or inter-frame prediction block). Here, the predicted block can be a block generated by intra-frame prediction or inter-frame prediction. Here, the transform can include at least one of a primary transform or a secondary transform. Transform coefficients can be generated by performing a primary transform on the residual signal, and secondary transform coefficients can be generated by performing a secondary transform on the transform coefficients.
[0188] A primary transformation can be performed using at least one of a plurality of predefined transformation methods. As an example, the plurality of predefined transformation methods may include a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), or a Caronanello Transform (KLT)-based transformation, etc. A secondary transformation can be performed on the transform coefficients generated after the primary transformation. The transformation method applied during the primary and / or secondary transformations can be determined based on at least one of the coding parameters of the current block and / or neighboring blocks. Optionally, transformation information indicating the transformation method can be transmitted via a signal.
[0189] Quantization levels can be generated by performing quantization on the residual signal or the result of performing primary and / or secondary transforms. Quantization levels can be scanned based on at least one of the intra-frame prediction mode or block size / shape, according to at least one of upper-right diagonal scan, vertical scan, or horizontal scan. For example, the coefficients of a block can be scanned using an upper-right diagonal scan to transform them into a one-dimensional vector form. Depending on the transform block size and / or intra-frame prediction mode, a vertical scan scanning two-dimensional block coefficients along the column direction or a horizontal scan scanning two-dimensional block coefficients along the row direction can be used instead of an upper-right diagonal scan. The scanned quantization levels can be entropy-encoded and included in the bitstream.
[0190] The decoder generates quantization levels by entropy decoding the bitstream. The quantization levels can be inversely scanned and arranged in two-dimensional blocks. In this case, at least one of the following inverse scanning methods can be performed: upper right diagonal scan, vertical scan, or horizontal scan.
[0191] It can perform dequantization on the quantization level, perform secondary inverse transform depending on whether to perform secondary inverse transform, and perform primary inverse transform depending on whether to perform primary inverse transform on the result of secondary inverse transform, thereby generating the reconstructed residual signal.
[0192] The following sections will describe Adaptive Motion Vector Resolution (AMVR) and Adaptive Block Vector Resolution (ABVR).
[0193] In the embodiments, the term "resolution" may refer to the term "motion vector resolution".
[0194] In adaptive motion vector resolution, the resolution of the motion vector difference can be adjusted on a block-by-block basis.
[0195] Adaptive motion vector resolution information can represent the resolution of motion vector differences. The resolution of the motion vector difference of a target block can be determined by transmitting / encoding / decoding the adaptive motion vector resolution information.
[0196] Information related to adaptive motion vector resolution may include at least one of the following: an indicator indicating whether adaptive motion vector resolution is used, an index indicating the motion vector resolution, the number of resolution candidates in the adaptive motion vector resolution, or the type of resolution candidates in the adaptive motion vector resolution. As an example, if the indicator indicating whether adaptive motion vector resolution is used in the target block is true, the index indicating the motion vector resolution in the target block can be signaled / encoded / decoded.
[0197] The motion vector resolution that can be applied to blocks can be the same or different.
[0198] For example, the resolution of the motion vector that can be applied to the target block can be determined based on at least one of the target block's encoding parameters, motion information, or pattern information.
[0199] Adaptive motion vector resolution can improve coding efficiency by adjusting the resolution of the motion vector difference.
[0200] For example, the adjusted resolution can be one of 16 pixels, 8 pixels, 4 pixels, full pixels, half pixels, and quarter pixels, and it is not limited to the pixels listed above.
[0201] When the adjusted resolution is n pixels, if the value of one component of the motion vector difference is changed by 1, the position indicated by the motion vector difference can be changed by n pixels. In other words, if the adjusted resolution of the target block is n pixels, each component of the motion vector difference can indicate the reference block in units of n pixels.
[0202] For example, if the motion vector difference that should actually be applied to the target block is (a, b) and the adjusted resolution is p pixels, then (a / p, b / p) can be encoded instead of (a, b). In other words, the motion vector difference transmitted / encoded in the encoding device can be (a / p, b / p). The decoding device can derive the original motion vector difference (a, b) again by multiplying the transmitted motion vector difference (a / p, b / p) by p.
[0203] Implementations of the inter-frame prediction modes in this disclosure can be extended.
[0204] First, because the target block is predicted with reference to a specific reconstructed block, inter-frame prediction mode, intra-frame block copy (IBC) mode, and intra-frame template matching prediction mode can share common characteristics.
[0205] Therefore, in embodiments, the inter-frame prediction mode can be replaced by the IBC mode or the intra-frame template matching mode. When the IBC mode or the intra-frame template matching mode is used for the target block, the description of the case where the inter-frame prediction mode is used for the target block can also be applied. The description of the inter-frame prediction mode can be applied to the IBC mode or the intra-frame template matching mode, and the inter-frame prediction mode can be replaced by the IBC mode or the intra-frame template matching mode. Furthermore, information related to the inter-frame prediction mode can be considered as information related to the IBC mode or the intra-frame template matching mode. The description of information related to the inter-frame prediction mode can be applied to information related to the IBC mode or the intra-frame template matching mode. For example, when the IBC mode or the intra-frame template matching mode is used for the target block, the value of the inter-frame prediction mode indicator can be 0 (or false). However, in this case, the value of the IBC mode indicator or the intra-frame template matching mode indicator can be 1 (or true).
[0206] Furthermore, in the embodiments, the inter-frame predicted motion vector (MV) can be replaced by the block vector (BV) of IBC. When BV is used for the target block, the description of the case where MV is used for the target block can also be applied. The description of MV can be applied to BV, and MV can be replaced by BV. In addition, information related to MV can be considered as information related to BV. The description of information related to MV can be applied to information related to BV. However, BV can be information indicating a specific reconstructed block in the target image that includes the target block, rather than in the reference image.
[0207] Therefore, by replacing the motion vector with a block vector, the description of adaptive motion vector resolution (AMVR) can also be applied to adaptive block vector resolution (ABVR).
[0208] In embodiments related to inter-frame prediction modes, a reference block and a reference template with template matching existing in the reference image are described. On the other hand, when IBC mode or intra-frame template matching mode is used for the target block, the reference block and reference template may exist only in the target image. Therefore, in embodiments, the reference image described in relation to the inter-frame prediction mode can be considered the target image in both IBC mode and intra-frame template matching mode. Optionally, in embodiments, the reference image described in relation to the inter-frame prediction mode may be limited to the target image in both IBC mode and intra-frame template matching mode, and images other than the target image may not be referenced in both IBC mode and intra-frame template matching mode.
[0209] Since the reference block of the target block is determined based on the block vector of the target block, and the predicted block of the target block is determined based on the determined reference block of the target block, IBC and intra-frame template matching methods can have common characteristics.
[0210] Therefore, in this embodiment, IBC can be replaced by intra template matching. When intra template matching mode is used for a target block, the description of the case where IBC mode is used for the target block can also be applied. The description of IBC mode can be applied to intra template matching mode, and IBC mode can be replaced by intra template matching mode. Furthermore, information related to IBC mode can be considered as information related to intra template matching mode. The description of information related to IBC mode can be applied to information related to intra template matching mode.
[0211] IBC modes can include at least one of IBC merge mode, IBC-AMVP mode, IBC mode based on sub-blocks, or IBC-MBVD mode.
[0212] In this embodiment, intra-template matching can be replaced by IBC. When IBC mode is used for the target block, the description of the case where intra-template matching mode is used for the target block can also be applied. The description of intra-template matching mode can be applied to IBC mode, and intra-template matching mode can be replaced by IBC mode. Furthermore, information related to intra-template matching mode can be considered as information related to IBC mode. The description of information related to intra-template matching mode can be applied to information related to IBC mode.
[0213] Intra-frame template matching modes may include at least one of the following: intra-frame template matching mode on the basis of all blocks (i.e., intra-frame template matching mode not on the basis of sub-blocks), template matching mode on the basis of sub-blocks, or IBC template matching mode.
[0214] The subsampling method will be described below. When performing subsampling, this means selecting only a subset of samples from a specific region.
[0215] Subsampling methods can be applied in various ways to the prediction stage, transformation stage, reconstruction stage, etc., and can be applied to various filters (downsampling filters, upsampling filters, reference sample filters, filters at cell boundaries, interpolation filters, low-pass filters, high-pass filters, etc.) and the techniques disclosed herein for selecting some samples from samples in a specific region, such as motion vector adjustment, etc.
[0216] When performing subsampling, this can mean: 1) selecting the SUBSAMPLE_START_HOR sample in the horizontal direction and the SUBSAMPLE_START_VER sample in the vertical direction based on specific samples; and 2) when selecting only some samples in a specific region, selecting samples whose sample intervals are multiples of SUBSAMPLE_STEP_HOR in the horizontal direction and SUBSAMPLE_STEP_VER in the vertical direction based on the samples in 1). Optionally, in the case of performing subsampling, this can mean: selecting a portion of the samples in 1) and 2). The location of the specific sample can be the top-left, bottom-left, bottom-right, or top-right sample of the region where subsampling is performed. However, the location of the specific sample is not limited to the top-left sample of the region where subsampling is performed.
[0217] SUBSAMPLE_START_HOR and SUBSAMPLE_START_VER can be 0 or positive integers. Information about at least one of SUBSAMPLE_START_HOR or SUBSAMPLE_START_VER can be sent / encoded / decoded using signals, or the values of SUBSAMPLE_START_HOR and / or SUBSAMPLE_START_VER can be determined to predefined values without sending / encoding / decoding information using signals.
[0218] SUBSAMPLE_STEP_VER and SUBSAMPLE_STEP_HOR can be 0, 1, 2, or positive integers, respectively. Information about at least one of SUBSAMPLE_STEP_VER or SUBSAMPLE_STEP_HOR can be sent / encoded / decoded using signals, or the values of SUBSAMPLE_STEP_VER and / or SUBSAMPLE_STEP_HOR can be determined to predefined values without sending / encoding / decoding information using signals.
[0219] Subsampling methods can be classified by the region where subsampling is performed, the location of the region where subsampling is performed, the size of the region where subsampling is performed, SUBSAMPLE_START_HOR, SUBSAMPLE_START_VER, SUBSAMPLE_STEP_HOR, and SUBSAMPLE_STEP_VER. However, the criteria used to classify subsampling methods are not limited to the values mentioned above.
[0220] Information about the subsampling method can be sent / encoded / decoded using signals, or a predefined subsampling method can be used without sending / encoding / decoding using signals.
[0221] Information about the subsampling method can be used to determine the subsampling method.
[0222] For example, information about the subsampling method could be information about the region where subsampling is performed, the location of the region where subsampling is performed, the size of the region where subsampling is performed, and at least one of SUBSAMPLE_START_HOR, SUBSAMPLE_START_VER, SUBSAMPLE_STEP_HOR, or SUBSAMPLE_STEP_VER.
[0223] For example, the subsampling method can be determined based on at least one of the target block's motion information, encoding parameters, size, or prediction mode.
[0224] The subsampling method may be determined as at least one of sequence level, frame level, sub-frame level, parallel block level, parallel block group level, stripe level, coding tree unit (CTU) level, coding unit (CU) level, or prediction unit (PU) level, but the unit of determination is not limited to this.
[0225] The Geometric Partitioning Model (GPM) will be described below.
[0226] Figure 7 An example illustrating the partitioning boundaries in a geometric partitioning pattern is shown.
[0227] Figure 8 It is a diagram showing the partition boundaries, partition offsets, and partition angles in the geometric partitioning pattern.
[0228] Figure 9 An example is shown that represents a weighted graph for each prediction block based on the geometric partition boundaries.
[0229] Reference Figure 7 The geometric partitioning pattern disclosed herein can have various types of division forms. Geometric partitioning forms may include at least one of vertical partitioning, horizontal partitioning, or diagonal partitioning. (See reference...) Figure 8Each partition type can be determined based on at least one of partition offset or partition angle. The partition offset and partition angle can be predefined values or determined from information encoded / decoded from the bitstream. Figure 7 Of the 20 geometric partitioning patterns, 4 are vertical or horizontal partitions, and 16 are diagonal partitions. These 20 geometric partitioning patterns can be specified by indexes and can be defined as separate tables in the encoder / decoder.
[0230] The prediction method using geometric partitioning patterns determines the partition boundaries that divide the target block into two parts in each direction in the form of straight lines, and may include a method for performing a weighted sum on the two prediction blocks (or reference blocks) using a weight map determined based on the partition boundaries. The subsampling method described above can be performed when performing the weighted sum.
[0231] At least one prediction block (or at least one reference block) in a geometric partitioning pattern may refer to a prediction block of unidirectional and / or bidirectional prediction, or a reference block in at least one direction of unidirectional and / or bidirectional prediction.
[0232] Optionally, for example, at least one prediction block in the geometric partitioning mode may refer to a prediction block predicted intra-frame.
[0233] For example, in geometric partitioning mode, one prediction block can be predicted inter-frame, and another prediction block can be predicted intra-frame. Here, as mentioned above, inter-frame prediction can be replaced by prediction using IBC mode or prediction using intra-template matching mode.
[0234] Template matching is described below.
[0235] Figure 10 This is an illustration of an example of template matching.
[0236] In template matching, the motion information of the target block can be determined and / or changed based on the calculation results of the cost function between the target template and the reference template.
[0237] The reference block may include at least one of the following: 1) a block indicated by initial motion information, 2) a block indicated by motion information derived in the template matching search process, 3) a block indicated by motion information finally improved by template matching, 4) a block in which a sample (or position) belonging to the search range of template matching is selected as one of the top left, bottom left, top right, bottom right and center, or 5) a block finally determined by template matching.
[0238] The size of the reference block can be the same as the size of the target block.
[0239] Motion information improved by template matching can be motion information with the lowest matching cost derived from template matching search processing. However, the methods used to derive motion information are not limited to the above criteria.
[0240] Template matching cost can refer to the result of calculating the templates of the target block and the reference block used in template matching using a cost function.
[0241] Each of the reference block, reference template, and reference region may include at least one of a predicted sample, a reconstructed sample, a residual sample, or a decoded sample of the reference image. Optionally, each of the reference block, reference template, and reference region may include at least one of a predicted sample, a reconstructed sample, a residual sample, or a decoded sample of the target image.
[0242] The target template may include neighboring samples of the target block. The reference region of the target block may include neighboring samples of the target block.
[0243] For example, the reference region of the target block may include at least one of the samples located in the lower left, left side, upper left, top, or upper right regions surrounding the target block.
[0244] For example, in template matching, the target template can be the same as the reference region of the target block.
[0245] For example, a sample based on a target template of a target block can be a sample corresponding to a sample based on a reference template of a reference block.
[0246] For example, when configuring a target template in template matching, you can select some samples within the reference region of the target block. The target template can be configured using the selected samples.
[0247] For example, the sample selected for configuring the target template based on the target block can be a sample corresponding to the sample selected for configuring the reference block template based on the reference block.
[0248] For example, the reference region of a target block based on a target block can be a region corresponding to the reference region of a reference block based on a reference block.
[0249] The reference template may include neighboring samples of the reference block. The reference region of the reference block may include neighboring samples of the reference block.
[0250] For example, the reference region of a reference block may include at least one of the samples located in the lower left, left side, upper left, bottom, or upper right regions surrounding the reference block.
[0251] For example, in template matching, the reference template can be the same as the reference region of the reference block.
[0252] For example, a sample based on a reference template of a reference block can be a sample corresponding to a sample based on a target template of a target block.
[0253] For example, when configuring a reference template in template matching, you can select some samples within the reference region of the reference block. The reference template can be configured using the selected samples.
[0254] For example, the sample selected for configuring the reference template based on the reference block can be a sample corresponding to the sample selected for configuring the target block template based on the target block.
[0255] For example, the reference region of a reference block based on a reference block can be a region corresponding to the reference region of a target block based on a target block.
[0256] The template matching method may include at least one of intra-frame template matching mode or inter-frame template matching mode.
[0257] Intra-frame template matching mode can refer to a template matching method for each of the reference block, reference template, and reference region, including at least one of the predicted sample, reconstructed sample, residual sample, or decoded sample of the target image.
[0258] Inter-frame template matching mode can refer to a template matching method for each of the reference block, reference template, and reference region, including at least one of the predicted sample, reconstructed sample, residual sample, or decoded sample of the target image.
[0259] The target / reference template for template matching may include at least one of the following: 1) at least one sample in the TMSIZE_LEFT line adjacent to the left of the target / reference block; or 2) at least one sample in the TMSIZE_ABOVE line adjacent to the top of the target / reference block. However, the method for configuring the template and / or the positional relationship between each sample in the template and the target / reference block is not limited to the above relationships or methods.
[0260] Each of TMSIZE_LEFT and TMSIZE_ABOVE can be 0, 1, 2, 3, 4, or a positive integer greater than or equal to 4. TMSIZE_LEFT and TMSIZE_ABOVE can be equal to each other. Optionally, TMSIZE_LEFT and TMSIZE_ABOVE can be different. Each of TMSIZE_LEFT and TMSIZE_ABOVE can be a predefined value or a value determined based on information transmitted / encoded / decoded using signals. Each of TMSIZE_LEFT and TMSIZE_ABOVE can be determined based on at least one of the target block's motion information, encoding parameters, size, or prediction mode.
[0261] When configuring the template for template matching, all samples in the reference region can be used, or only some samples in the reference region can be used. When determining these samples, the subsampling method described above can be applied.
[0262] The template used for template matching can refer to at least one of the templates of the target block or the template of the reference block.
[0263] The reference region can refer to at least one of the reference regions of the target block used for template matching or the reference regions of the reference block.
[0264] When configuring a template for template matching by using only some samples located in the reference region, subsampling can be performed on all or part of the reference region.
[0265] When configuring a template for template matching using only some samples located within a reference region, the reference region can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, 2) a region where subsampling is not performed and is used for template configuration, and 3) a region not used for template configuration. The template for template matching can be configured using samples selected through subsampling in 1) and samples from the regions in 2).
[0266] For example, the region corresponding to 1) could be the region within the reference region that belongs to the left and / or upper left of the block.
[0267] For example, the area corresponding to 1) could be the top and / or upper left area of the block within the reference area.
[0268] Optionally, when configuring the template for template matching by using only some samples located in the reference region, the reference region can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, and 2) a region not used for template configuration. The template for template matching can be configured using samples selected through subsampling in 1).
[0269] When calculating the cost function between a target template and a reference template in template matching, all samples from each template can be used, or only some samples from each template can be used. In other words, the cost function can be calculated only for some samples.
[0270] When calculating the cost function between templates by using only at least a portion of the samples within the template, subsampling can be performed on all or part of the template region.
[0271] When calculating the cost function between templates by using only at least a portion of the samples within the template, the region of the template used for template matching can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, 2) a region where subsampling is not performed but used to calculate the cost function, and 3) a region not used to calculate the cost function. The cost function between templates in template matching can be calculated using samples selected by subsampling in 1) and samples from the regions in 2).
[0272] Optionally, when calculating the cost function between templates by using only at least a portion of the samples within the template, the region of the template used for template matching can be divided into at least two regions. The divided regions can be one of the following: 1) a region for which subsampling is performed, and 2) a region not used for calculating the cost function. The cost function between templates in template matching can be calculated using the samples selected by subsampling in 1).
[0273] When performing template matching search processing, all samples / locations within the search area can be used, or only some samples / locations within the search area can be selected. Search and / or matching costs can be calculated only for the selected samples / locations. Optionally, search and / or matching costs can be calculated only for motion information indicating the selected samples / locations.
[0274] When performing template matching search processing using only some samples / locations in the search region, subsampling can be performed on all or part of the search region.
[0275] When performing template matching search processing using only some samples / locations within the search area, the search area can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, 2) a region where search processing is performed without subsampling, and 3) a region where search processing is not performed. Template matching search processing can be performed on pixels and / or locations selected by subsampling in 1) and samples / locations within the regions in 2). Optionally, template matching search processing can be performed on motion information indicating samples / locations selected by subsampling in 1) and samples / locations within the regions in 2).
[0276] Optionally, when performing template matching search processing by using only some samples / locations within the search area, the search area can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, and 2) a region where no search processing is performed. Template matching search processing can be performed using the samples / locations selected through subsampling in 1). Optionally, template matching search processing can be performed based on motion information indicating the samples / locations selected through subsampling in 1).
[0277] Figure 11 This illustrates the subsampling method for the search range of template matching.
[0278] Figure 12 A subsampling method for templates used in template matching is shown.
[0279] Figure 11 and Figure 12 The shaded sample (or location) in the text can refer to the sample (or location) selected through subsampling.
[0280] For all or part of the reference area, it can be as follows Figure 12 The subsampling shown can be performed, and the template can be configured by using only the selected sample (or location).
[0281] For all or part of the template region matched by the template, it can be like this: Figure 12 The diagram shows the performance of subsampling, and the cost function can be computed only on selected samples (or locations).
[0282] For all or part of the search region for template matching, it can be like this: Figure 11 The diagram shows the performance of subsampling, and the search and / or matching costs can be calculated only for selected pixels and / or locations.
[0283] For all or part of the search region for template matching, as shown below Figure 11 The subsampling shown can be performed, and the search and / or matching costs can be calculated only on motion information indicating the selected pixels and / or locations.
[0284] By performing a first search step using first motion information as initial motion information, second motion information can be derived as a result of correcting the first motion information. In a second search step performed after the first search step, the second motion information can be used as initial motion information.
[0285] If the initial motion information (e.g., the initial motion vector or the initial block vector) is not motion information in integer pixels (i.e., in fractional pixels), then the result of rounding (or rounding down or rounding up) the corresponding initial motion information can be used as the initial motion information instead.
[0286] For example, in order to generate a reference template at the location indicated by motion information obtained by adding a specific offset to initial motion information in fractional pixels during the search process, samples at fractional pixel locations must be generated by applying an interpolation filter to samples at integer pixel locations. However, if the initial motion information is constrained to integer pixels, the complexity can be reduced because interpolation at fractional pixel locations is not required during the search process.
[0287] The search can be performed by using a cost function to determine the similarity between NUM_TEMPLATE_COMPARE templates.
[0288] The search may include the process of determining at least one piece of motion information that meets specific conditions within a specific search range. The motion information of the target block may be determined and / or modified based on the at least one piece of motion information determined through the search.
[0289] Motion information that meets specific conditions can refer to the motion information with the lowest matching cost among the motion information within the search range, but is not limited to this.
[0290] Optionally, the search may include the process of identifying at least one block within a specific search range that satisfies specific conditions. Motion information of the blocks identified through the search can be used as motion information for the target block.
[0291] A block that meets specific conditions can refer to the reference block with the lowest matching cost among the reference blocks in the search range, but is not limited to this.
[0292] The cost function can refer to a function that determines the similarity between at least one sample in the target template and at least one sample in the reference template.
[0293] The similarity between the first and second values can be determined by using at least one of the following: 1) comparing a specific value to the difference between the two values, 2) the ratio between the two values, or 3) the difference between the two values.
[0294] The cost function can be a function that determines the similarity between at least one sample in the target template and a sample in its corresponding reference template.
[0295] The cost function can be at least one of the following: Sum of Absolute Differences (SAD), Sum of Absolute Transformed Differences (SATD), Sum of Absolute Differences with Mean Removed (MR-SAD), Mean Squared Error (MSE), and Sum of Squared Errors (SSE). However, the cost function is not limited to the items listed above.
[0296] The cost function used in template matching can be predefined or determined based on information sent / encoded / decoded by signals.
[0297] For example, MR-SAD can be used as a cost function in template matching if the target block satisfies the activation conditions of bilateral matching and / or a portion thereof, or if bilateral matching is performed in the target block.
[0298] For example, if the target block does not satisfy the activation conditions of bilateral matching and / or a portion thereof, or if bilateral matching is not performed in the target block, SAD can be used as the cost function in template matching.
[0299] For example, the type of cost function in a bilateral match can be determined based on whether specific conditions are met in the bilateral match. In this case, the type of cost function in the template match can be determined based on whether the activation conditions of the bilateral match and the specific conditions used to determine the type of cost function in the bilateral match are met.
[0300] For example, if the target block satisfies the activation condition of bilateral matching and the specific condition, then MR-SAD can be used as the cost function in bilateral matching; otherwise, SAD can be used as the cost function in bilateral matching.
[0301] For example, if the target block satisfies the activation condition of bilateral matching, and bidirectional prediction using weights is performed or the number of samples in the target block is greater than a certain value, then MR-SAD can be used as the cost function in template matching; otherwise, SAD can be used as the cost function in template matching.
[0302] The search area can be a specific range centered on the location indicated by the initial motion information. In other words, the center of the search area can be the location indicated by the initial motion information.
[0303] Optionally, the search range can be a specific range indicated by the initial motion information to the upper left. In other words, the upper left of the search range can be the position indicated by the initial motion information.
[0304] Optionally, the search range may consist of a pre-reconstructed region surrounding the target block. As an example, the search range may include at least one of the samples (or the locations of samples) located in the lower left, left side, upper left, top, or upper right regions surrounding the target block.
[0305] At least one of the dimensions or shapes of the target block search range can be predefined in the encoder and decoder.
[0306] Optionally, at least one of the size or shape of the search range in the target block can be determined based on at least one of the target block size, the target block encoding parameters, the target block motion information, or the target block prediction mode.
[0307] Optionally, information representing one of the dimensions and shape of the search range of the target block can be encoded and transmitted using a signal.
[0308] The search area can be a rectangle with a horizontal length of SR_X and a vertical length of SR_Y. Alternatively, the search area can be a rhombus with a horizontal length of SR_X and a vertical length of SR_Y. Alternatively, the search area can be a cross shape with a horizontal length of SR_X and a vertical length of SR_Y. However, the shape and size of the search area are not limited to the above embodiments.
[0309] Each of SR_X and SR_Y can be a positive integer. Each of SR_X and SR_Y can be a predefined value or a value determined based on information sent / encoded / decoded by signals.
[0310] Initial motion information can be determined based on at least one of the following: motion information of the target block, encoding parameters of the target block, motion vector of the target block, reference image of the target block, block vector of the target block, motion vector predictor of the target block, block vector predictor of the target block, motion information of at least one neighboring block of the target block, merging candidate of the target block, motion vector difference of the target block, or block vector difference of the target block.
[0311] Search methods can be classified based on at least one of the following: search mode, search resolution, search range, initial motion information, or units from which motion information is derived.
[0312] The search method can be determined based on the motion information of the target block, the encoding parameters of the target block, the size of the target block, the prediction mode of the target block, the reference image of the target block, at least one sample value in the target block, the target template, at least one sample value in the target template, or at least one sample value in the region of the target template.
[0313] The search pattern can be one of the diamond pattern, cross pattern, or full search pattern. However, the search pattern is not limited to the patterns listed above.
[0314] When (0, 0) represents the position indicated by the initial motion information, the search using the diamond pattern can mean searching for at least one of the following positions: (0, 2×RR), (RR, RR), (2×RR, 0), (RR, -RR), (0, -RR), (-RR, -RR), (-RR, 0), (-RR, RR), and (0, 0).
[0315] When (0, 0) represents the position indicated by the initial motion information, the search using the cross pattern can mean searching for at least one of the positions (0, RR), (RR, 0), (0, -RR), (-RR, 0), and (0, 0).
[0316] RR can be the search resolution or a value determined based on the search resolution, or it can be a predefined positive number.
[0317] Using the full search mode means searching all locations within a predefined search range.
[0318] For example, when FS_i has values from -FS_X to FS_X and FS_j has values from -FS_Y to FS_Y, a search using full search mode can mean searching for the position of (FS_i × RR, FS_j × RR). Here, (0, 0) can be the position indicated by the initial motion information. However, the search range is not limited to the positions mentioned above. Each of FS_X and FS_Y can be a predefined positive number.
[0319] The search resolution can be one of 4 pixels, full pixels, half pixels, or a quarter pixels. However, the search resolution is not limited to the above pixels.
[0320] The search resolution can be predefined, determined based on at least one piece of information about the adaptive motion vector resolution, or determined based on values sent / encoded / decoded by signals.
[0321] The units from which motion information is derived can include units of the entire block and units of sub-blocks.
[0322] Figures 13 to 18 The search methods in template matching according to the embodiments are shown respectively.
[0323] Based on the motion information of the target block, encoding parameters, prediction mode, and adaptive motion vector resolution, it is possible to determine Figures 13 to 18 The table shown contains specific columns. Searches can be performed from top to bottom using the search mode and search resolution corresponding to the rows marked "v".
[0324] For example, if the target block is in AMVP mode and the resolution is determined by adaptive motion vector resolution is Figure 13 If the value is 4 pixels, then a crosshair search with a 4-pixel search resolution can be performed after a diamond-shaped search with a 4-pixel search resolution.
[0325] ALT_IF can refer to the index of an adaptive interpolation filter. An interpolation filter can be applied to compute the pixel value at a specific resolution sampling location. The adaptive interpolation filter can be one of several interpolation filters selected by its index. In other words, when applying an adaptive interpolation filter, different interpolation filters can be used depending on the index to compute the pixel value at a specific resolution sample location.
[0326] For example, a specific resolution can be half a pixel. However, a specific resolution is not limited to half a pixel.
[0327] For example, the interpolation filter determined by the index can be one of a 6-tap interpolation filter and an 8-tap interpolation filter. However, the method for determining the interpolation filter is not limited to the methods described above.
[0328] Figure 19 A method for configuring a first template in affine mode according to an embodiment is shown.
[0329] CPMV refers to Affine Control Point Motion Vector (CPMV). The MV of each sub-block within a target block can be derived using CPMV. CPMV can be achieved using either a 4-parameter affine motion model employing two CPMVs or a 6-parameter affine motion model employing three CPMVs.
[0330] When the target block is in affine mode, it can be divided into sub-blocks with a width of N and a height of M. The motion information of each sub-block can be determined based on at least one of the target block's motion information, encoding parameters, or dimensions. The template matching cost of the target block can be determined based on at least one of the template matching costs of the divided sub-blocks. For example, the template matching cost of the target block can be the sum of the template matching costs of the divided sub-blocks, or the average of the template matching costs of each sub-block.
[0331] N and M can be 2, 4, 8, 16, or positive integers. Each of N and M can be a predefined value or a value determined based on information transmitted / encoded / decoded using signals.
[0332] Figure 20 The example illustrates a method for configuring a second template in affine mode.
[0333] refer to Figure 20A, the reference templates (A0~A3 to L0~L4) of the target block can be configured by the target block-based CPMV including at least one sample from the samples at the positions indicated by the target template positions corresponding to each reference template, based on the motion vector (or block vector) determined for each reference template. The motion vector (or block vector) determined for each reference template position can refer to the motion vector (or block vector) determined relative to at least one position in the corresponding reference template (e.g., the center position of the corresponding reference template).
[0334] refer to Figure 20 B may include configuring the reference templates of the target block (A0~A3 to L0~L4) by including at least one sample from the samples at the location indicated by the target template position corresponding to each reference template, based on the motion vector (or block vector) of the nearest sub-block within the target block.
[0335] For example, the motion information of a target block can be determined based on the motion information of neighboring blocks. This can be expressed as "the target block inherits motion information from its neighboring blocks".
[0336] For example, when the target block is in merge mode, a merge candidate can be specified from the merge candidate list based on the merge index, and the motion information of the specified merge candidate can be used as the motion information of the target block.
[0337] For example, when the target block is in AMVP mode, an MV candidate can be specified from the MV candidate list based on the MV candidate index, and the motion information of the specified MV candidate can be used as the motion information of the target block.
[0338] When the motion information inherited by the target block from neighboring blocks indicates bidirectional prediction, an embodiment of template matching in the target block can be performed as follows. Each of the following steps can be performed sequentially.
[0339] In the first step, template matching can be performed for each of the L0 and L1 directions, and the template matching cost (C0, C1) for the motion information determined for the L0 and L1 directions can be calculated.
[0340] In this scenario, when performing template matching for each direction, it can be performed in the same way as template matching in unidirectional prediction for the corresponding direction, without considering motion information in other directions. For example, in this case, if the target block satisfies predefined conditions, MR-SAD can be used as the cost function; otherwise, SAD can be used as the cost function.
[0341] The predefined condition may be a condition based on at least one of the following: whether a model-based prediction method is performed in the target block; an indicator indicating whether a model-based prediction method is performed in the target block; whether bilateral matching is performed in the target block; an indicator indicating whether bilateral matching is performed in the target block; motion information of the target block; the size of the target block; coding parameters of the target block; motion information of neighboring blocks of the target block; coding parameters of neighboring blocks of the target block; or the type of cost function in template matching in neighboring blocks of the target block.
[0342] For example, if whether a model-based prediction method is performed in the target block is true, or if the indicator indicating whether a model-based prediction method is performed in the target block is true, then MR-SAD can be used as the cost function; otherwise, SAD can be used as the cost function.
[0343] For example, if whether bilateral matching is performed in the target block is true, or if the indicator indicating whether bilateral matching is performed in the target block is true; and the number of samples in the target block is equal to or greater than a specific value, then MR-SAD can be used as the cost function; otherwise, SAD can be used as the cost function.
[0344] The cost function may refer to the cost function used when searching for template matching; and / or the cost function for calculating at least one of C0, C1, or C'. The cost function used when searching for template matching and the cost function for calculating at least one of C0, C1, or C' may be the same or different.
[0345] For example, when searching for template matching, MR-SAD can be used as the cost function, and C0, C1, and C' can be calculated by using SAD.
[0346] In the second step, for C0 < C1, a new target template T' can be generated by using the target template and the template in the L0 direction.
[0347] For example, when the target template is T, the reference template in the L0 direction is T0, and the reference template in the L1 direction is T1, it can be determined according to Here, and can be predefined values respectively. In addition, and can be values determined based on whether weighted bidirectional prediction is performed in the target block; and / or the weights in weighted bidirectional prediction. For example, can be 2, and can be -1.
[0348] For C0 > C1, a new target template T' can be generated using the target template and the template in the L1 direction. If the value of C0 is the same as the value of C1, it can be considered as one of C0 < C1 or C1 > C0, and the process in the second step can be executed.
[0349] In the third step, for C0 < C1, template matching can be performed by treating the target template as T' for the L1 direction, and the template matching cost C' of the motion information determined for the L1 direction can be calculated.
[0350] For C0 > C1, template matching can be performed by treating the target template as T' for the L0 direction, and the template matching cost C' of the motion information determined for the L0 direction can be calculated.
[0351] If the value of C0 is the same as the value of C1, it can be considered as one of C0 < C1 or C1 > C0, and the process in the third step can be executed.
[0352] In the fourth step, for , the motion information of the target block can be changed to the motion information indicating unidirectional prediction in the L0 direction or the L1 direction based on the values of C0 and C1.
[0353] For C0 < C1, the motion information of the target block can be changed to the motion information indicating unidirectional prediction in the L0 direction. Optionally, the motion information in the L1 direction can be considered unavailable in the target block.
[0354] For C0 > C1, the motion information of the target block can be changed to the motion information indicating unidirectional prediction in the L0 direction. Optionally, the motion information in the L1 direction can be considered unavailable in the target block.
[0355] If the value of C0 is the same as the value of C1, it can be considered as one of C0 < C1 or C1 > C0, and the process in the fourth step can be executed. Here, and can be predefined values respectively. Here, and can be values determined based on whether weighted bidirectional prediction is performed in the target block; and / or the weights in the weighted bidirectional prediction. For example, can be 1, and can be 1 / 8.
[0356] For example, the second step to the fourth step can be executed only when the target block meets the predefined conditions.
[0357] For example, steps two through four can only be executed if the target block is bidirectional prediction and bidirectional matching is not performed in the target block or the target block does not meet the activation conditions for bidirectional matching.
[0358] In bilateral matching, the reference block in the L0 direction and the reference block in the L1 direction can be used as templates, and the motion information of the target block can be determined and / or changed based on the calculation results of the cost function between the two templates.
[0359] A reference block may include at least one of the following: 1) a reference block indicated by initial motion information; 2) a reference block indicated by motion information derived in the bilateral matching search process; or 3) a reference block indicated by motion information finally improved by bilateral matching.
[0360] For example, when configuring a template for bilateral matching, a reference block in the L0 direction and a reference block in the L1 direction can be used as templates.
[0361] The bilateral matching cost can refer to the result value calculated using the cost function on the templates of the reference blocks in the L0 direction and the L1 direction used in bilateral matching.
[0362] If the target block is in Intra-Block Copy (IBC) mode and is predicted using at least two reference blocks, bilateral matching can also be performed using two different reference blocks from the target block's reference blocks as templates.
[0363] When configuring a template for bilateral matching, you can select only some pixels and / or positions within the reference blocks in the L0 direction and the L1 direction. The template can be configured using only the selected pixels and / or selected positions.
[0364] The template used for bilateral matching may refer to at least one of the templates in the L0 direction or the templates in the L1 direction.
[0365] For example, when configuring a template for bilateral matching, subsampling can be used for the reference block in the L0 direction and the reference block in the L1 direction.
[0366] For example, when configuring a template for bilateral matching, subsampling can be used for a portion of the reference block in the L0 direction and a portion of the reference block in the L1 direction.
[0367] Optionally, for example, when configuring a template for bilateral matching, the reference block in the L0 direction and the reference block in the L1 direction can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, 2) a region where subsampling is not performed and is used for template configuration, or 3) a region not used for template configuration. The template for bilateral matching can be configured using pixels and / or positions selected by subsampling in 1) and pixels and / or positions in the regions in 2).
[0368] Optionally, for example, when configuring a template for bilateral matching, the reference block in the L0 direction and the reference block in the L1 direction can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, and 2) a region not used for template configuration. The template for bilateral matching can be configured using pixels and / or positions selected by subsampling in 1).
[0369] For example, when configuring a template for bilateral matching, the area used to configure the template can be a portion of the reference block in the L0 direction and a portion of the reference block in the L1 direction.
[0370] For example, when configuring a bilateral matching template, pixels (or positions) for configuring the template can be selected from only a portion of the reference block in the L0 direction and a portion of the reference block in the L1 direction.
[0371] The size of a portion of the reference block in the L0 direction can be smaller than the size of the reference block in the L0 direction. For example, the height (vertical dimension) of a portion of the reference block in the L0 direction can be smaller than the height (vertical dimension) of the reference block in the L0 direction. For example, the width (horizontal dimension) of a portion of the reference block in the L0 direction can be smaller than the width (horizontal dimension) of the reference block in the L0 direction.
[0372] Optionally, the size of a portion of the reference block in the L1 direction can be smaller than the size of the reference block in the L1 direction. For example, the height (vertical dimension) of a portion of the reference block in the L1 direction can be smaller than the height (vertical dimension) of the reference block in the L1 direction. For example, the width (horizontal dimension) of a portion of the reference block in the L1 direction can be smaller than the width (horizontal dimension) of the reference block in the L1 direction.
[0373] When calculating the cost function between templates in bilateral matching, you can select only some pixels and / or locations within the template region. The cost function can be calculated only for the selected pixels and / or selected locations.
[0374] For example, when calculating the cost function between templates in a bilateral matching process, subsampling can be performed on the template regions in the L0 direction and the template regions in the L1 direction.
[0375] For example, when calculating the cost function between templates in a bilateral matching, subsampling can be performed on a portion of the template region in the L0 direction and a portion of the template region in the L1 direction.
[0376] Optionally, for example, when calculating the cost function between templates in bilateral matching, each of the template regions in the L0 direction and the template regions in the L1 direction can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, 2) a region where subsampling is not performed but used to calculate the cost function, and 3) a region not used to calculate the cost function. The calculation of the cost function between templates in bilateral matching can be performed using pixels and / or positions selected by subsampling in 1) and pixels and / or positions within the regions in 2).
[0377] Optionally, for example, when calculating the cost function between templates in bilateral matching, each of the template regions in the L0 direction and the L1 direction can be divided into at least two regions. The divided regions can be one of the following: 1) regions where subsampling is performed, and 2) regions not used for calculating the cost function. The calculation of the cost function between templates in bilateral matching can be performed using pixels and / or positions selected by subsampling in 1).
[0378] For example, when calculating the cost function between templates in a bilateral matching process, the region used to calculate the cost function can be a portion of the template in the L0 direction and a portion of the template in the L1 direction.
[0379] The dimensions of a portion of the template in the L0 direction can be smaller than the dimensions of the template in the L0 direction. For example, the height (vertical dimension) of a portion of the template in the L0 direction can be smaller than the height (vertical dimension) of the template in the L0 direction. For example, the width (horizontal dimension) of a portion of the template in the L0 direction can be smaller than the width (horizontal dimension) of the template in the L0 direction.
[0380] Optionally, the size of a portion of the template in the L1 direction can be smaller than the size of the template area in the L1 direction. For example, the height (vertical dimension) of a portion of the template in the L1 direction can be smaller than the height (vertical dimension) of the template in the L1 direction. For example, the width (horizontal dimension) of a portion of the template in the L1 direction can be smaller than the width (horizontal dimension) of the template in the L1 direction.
[0381] When performing bilateral matching search processing, only some of the pixels and / or locations within the search area may be selected. Search and / or matching costs may be calculated only for the selected pixels and / or selected locations. Alternatively, search and / or matching costs may be calculated only for motion information indicating the selected pixels and / or locations.
[0382] For example, when performing a bilateral matching search, subsampling can be performed on all or part of the search region.
[0383] Optionally, for example, when performing bilateral matching search processing, the search area can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, 2) a region where subsampling is not performed but search processing is performed, and 3) a region where search processing is not performed. Bilateral matching search processing can be performed on pixels and / or positions selected by subsampling in 1) and pixels and / or positions in the region in 2). Optionally, bilateral matching search processing can be performed on motion information indicating pixels and / or positions selected by subsampling in 1) and pixels and / or positions in the region in 2).
[0384] Optionally, for example, when performing bilateral matching search processing, the search area can be divided into at least two regions. The divided regions can be one of the following: 1) a region where subsampling is performed, and 2) a region where search processing is not performed. Bilateral matching search processing can be performed using the pixels and / or positions selected by subsampling in 1). Optionally, bilateral matching search processing can be performed based on motion information indicating the pixels and / or positions selected by subsampling in 1).
[0385] Figure 21 Various examples of subsampling methods in bilateral matching are shown.
[0386] Figure 21 The shaded sample (or location) in the text can refer to the sample (or location) selected through subsampling.
[0387] For the reference block region in the L0 direction and the reference block region in the L1 direction, it can be as follows: Figure 21 The subsampling shown can be performed, and the template can be configured by using only the selected samples (or locations).
[0388] For a portion of the reference block region in the L0 direction and a portion of the reference block region in the L1 direction, it can be as follows: Figure 21 The subsampling shown can be performed, and the template can be configured by using only the selected samples (or locations).
[0389] For all or part of the template region of bilateral matching, it can be as follows: Figure 21The example shows the performance of subsampling, and the cost function can be computed only for selected samples (or locations).
[0390] For all or part of the search region in a bilateral matching, it can be like... Figure 21 The subsampling shown can be performed, and the search and / or matching costs can be calculated only for selected pixels and / or locations.
[0391] For all or part of the search region in a bilateral matching, it can be like... Figure 21 The subsampling shown can be performed, and the search and / or matching cost can be calculated only for motion information indicating the selected pixel and / or location.
[0392] Two-sided matching can always be performed.
[0393] In contrast, bilateral matching can operate only when predefined enabling conditions are met.
[0394] For example, bilateral matching can be performed when inter-frame prediction mode is used for the target block and at least two reference blocks are used.
[0395] For example, bilateral matching can only be performed if the first direction and the second direction are different from each other and the first POC interval and the second POC interval are the same. The first direction can be the direction from the target image to the reference image in the L0 direction. The second direction can be the direction from the target image to the reference image in the L1 direction. The first POC interval can be the difference between the POC of the target image and the POC of the reference image in the L0 direction. The second POC interval can be the difference between the POC of the target image and the POC of the reference image in the L1 direction.
[0396] For example, bilateral matching can only be performed if the first direction and the second direction are different from each other. The first direction can be the direction from the target image to the reference image in the L0 direction. The second direction can be the direction from the target image to the reference image in the L1 direction.
[0397] Here, when the first direction and the second direction are different from each other, it can mean that the following equation is satisfied.
[0398]
[0399] Here, when the first direction and the second direction are the same, it can mean that the following equation is satisfied.
[0400]
[0401] It can be the POC of the target image. POC0 can be the POC of the reference image in the L0 direction. POC1 can be the POC of the reference image in the L1 direction.
[0402] The search can be performed by using a cost function to determine the similarity between two templates.
[0403] The search may include the process of determining at least one piece of motion information that meets specific conditions within a specific search range. The motion information of the target block may be determined and / or modified based on the at least one piece of motion information determined through the search.
[0404] Motion information that meets specific conditions can refer to the motion information with the lowest matching cost among the motion information within the search range, but is not limited to this.
[0405] Optionally, the search may include the process of identifying at least one block within a specific search range that satisfies specific conditions. Motion information of the blocks identified through the search can be used as motion information for the target block.
[0406] A block that meets specific conditions may refer to motion information with the lowest matching cost among reference blocks within the search range, but is not limited to this.
[0407] The cost function can refer to a function that determines the similarity between at least one sample in the first template and at least one sample in the second template of two templates in a bilateral matching.
[0408] The similarity between a first value and a second value can be determined by using at least one of the following: 1) comparing a specific value to the difference between two values, 2) the ratio between the two values, or 3) the difference between the two values.
[0409] It can be a function that determines the similarity between at least one sample in a first template and a sample in a corresponding second template.
[0410] The cost function can be at least one of the following: Sum of Absolute Differences (SAD), Sum of Absolute Transformed Differences (SATD), Sum of Absolute Differences with Mean Removed (MR-SAD), Mean Squared Error (MSE), and Sum of Squared Errors (SSE). However, the cost function is not limited to the items listed above.
[0411] The cost function used in bilateral matching can be predefined or determined based on information sent / encoded / decoded by signals.
[0412] For example, if the size of the target block is smaller than a certain value, SAD can be used as the cost function in bilateral matching; otherwise, MR-SAD can be used as the cost function in bilateral matching. The size of the block can include at least one of the following: the width of the block, the height of the block, the sum of the width and height of the block, or the product of the width and height of the block.
[0413] For example, if BCW is not performed in the target block or the same weights are used for the reference blocks in the L0 and L1 directions in BCW, then SAD can be used as the cost function in bilateral matching; otherwise, MR-SAD can be used as the cost function in bilateral matching.
[0414] For example, if a local illumination compensation mode is not performed in the target block, then SAD can be used as the cost function in two-sided matching; otherwise, MR-SAD can be used as the cost function in two-sided matching. A local illumination compensation mode can refer to a mode that calculates the correlation between the template of the target block and the template of the reference block to derive at least one of weights or offsets and multiplies or adds it to all or part of the target block (or the reference block of the target block), or multiplies and adds them together.
[0415] If the initial motion information (e.g., the initial motion vector or the initial block vector) is not motion information in integer pixels (i.e., in fractional pixels), then rounding the result of the corresponding initial motion information (or rounding down or rounding up) can be used instead as the initial motion information.
[0416] For example, in order to generate a template at the location indicated by motion information obtained by adding a specific offset to initial motion information in fractional pixels during the search process, it is necessary to generate samples at fractional pixel locations by applying an interpolation filter to samples at integer pixel locations. However, if the initial motion information is constrained to integer pixels, the complexity can be reduced because interpolation at fractional pixel locations is not required during the search process.
[0417] Two-sided matching may include at least one search step.
[0418] For example, bilateral matching can be configured to sequentially include the steps of 1) deriving motion information for the entire block and 2) deriving motion information for the sub-blocks of the block. However, the methods for deriving motion information executed in each step and the order of the steps are not limited to the configuration described above.
[0419] In each search step of bilateral matching, the motion information in BM_NUM directions can be improved between the motion information in the L0 direction and the motion information in the L1 direction. BM_NUM can be 0, 1, 2, or a positive integer. The BM_NUM used in the steps of bilateral matching can be the same or different.
[0420] For example, if BM_NUM is 1 in a particular search step, motion information improvement can be performed only on motion information in the LXBM direction in that particular search step.
[0421] For example, if BM_NUM is 1 and XBM is 0 in a specific search step, then in that specific search step, only the search in the L0 direction can be performed, while the template and motion information in the L1 direction are fixed.
[0422] XBM can be 0, 1, or a positive integer.
[0423] XBM can be predefined.
[0424] For example, the XBM direction can be a direction with a large POC interval between the L0 and L1 directions. The POC can be the difference between the POC of the reference image (in a specific direction) and the POC of the target image.
[0425] For example, the XBM direction can be a direction with a higher template matching cost between the L0 and L1 directions. Here, the template matching cost in a specific direction can be the template matching cost of motion information in that specific direction.
[0426] For example, information about XBM can be sent / encoded / decoded using signals.
[0427] For example, if the first POC difference is greater than the second POC difference, then XBM can be 0; otherwise, XBM can be 1. Optionally, if the first POC difference is greater than the second POC difference, then XBM can be 1; otherwise, XBM can be 0. The first POC difference can be the difference between the POC of the target image and the POC of the reference image in the L0 direction. The second POC difference can be the difference between the POC of the target image and the POC of the reference image in the L1 direction.
[0428] For example, XBMs can be determined based on contextual and / or probabilistic models used for entropy encoding and decoding of motion information and encoding parameters of the target block.
[0429] For example, the XBM can be determined based on at least one of the context model and / or probability model used when entropy coding and entropy decoding of inter-frame prediction indicators in the target block.
[0430] For example, based on the context model and / or probabilistic model used in entropy coding and decoding of inter-frame prediction indicators, a more potential direction between unidirectional predictions in the L0 direction and unidirectional predictions in the L1 direction within the target block can be selected as the LXBM direction.
[0431] Optionally, for example, based on the context model and / or probabilistic model used when entropy coding and entropy decoding of inter-frame prediction indicators, a more potential direction between unidirectional predictions in the L0 direction and unidirectional predictions in the L1 direction in the target block can be selected as the L(1-XBM) direction.
[0432] A more potential direction may refer to a direction that uses fewer bits when entropy encoding is performed using a context model and / or a probability model. Alternatively, a more potential direction may refer to a direction with a higher probability indicated by the context model and / or the probability model for that direction.
[0433] For example, the LXBM can be determined based on the inter-frame bidirectional prediction weights of the target block. For instance, the LXBM could be a direction in the target block that has a higher inter-frame bidirectional prediction weight between the L0 and L1 directions. Alternatively, for example, the LX could be a direction in the target block that has a lower inter-frame bidirectional prediction weight between the L0 and L1 directions.
[0434] For example, an XBM can be determined based on at least one of the motion information of neighboring blocks or encoding parameters.
[0435] For example, it can be based on and Figure 19 The motion information or encoding parameters of at least one of the corresponding neighboring blocks of A0, A1, B0, B1 or B2 in the target block are used to determine X.
[0436] For example, the XBM in the target block can be determined based on at least one of the inter-frame prediction indicator or the inter-frame bidirectional prediction weight in the neighboring block.
[0437] For example, one or more context models and / or probability models can be used for entropy encoding and entropy decoding of XBM.
[0438] Among multiple context models and / or probabilistic models, a context model and / or probabilistic model for entropy encoding and entropy decoding of XBMs in a target block can be determined based on at least one of the motion information or encoding information of neighboring blocks.
[0439] For example, the context model and / or probability model for entropy encoding and entropy decoding of XBMs within a block can be the same. Alternatively, the context model and / or probability model for entropy encoding and entropy decoding of XBMs within a block can be different based on at least one of the inter-frame prediction direction or inter-frame bidirectional prediction weights of neighboring blocks.
[0440] The same XBM can be used in the search steps of bilateral matching. Alternatively, different XBMs can be used in the search steps of bilateral matching.
[0441] For example, when performing bilateral matching, the cost function used to calculate the cost of bilateral matching can be determined based on the weights of the target block's bidirectional prediction and / or the weight index of the bidirectional prediction.
[0442] In an embodiment, bidirectional prediction using weights can refer to bidirectional prediction using CU-level weights (BCW).
[0443] For example, if the first weight and the second weight of the target block are the same, the bilateral matching cost can be calculated using SAD or SATD. If the first weight and the second weight of the target block are different, the bilateral matching cost can be calculated using MRSAD or MRSATD. Here, the first weight can be the weight for the L0 direction in bidirectional prediction using weights. The second weight can be the weight for the L1 direction in bidirectional prediction using weights.
[0444] BM_NUM can be 0, 1, 2 or a positive integer.
[0445] BM_NUM can be predefined.
[0446] BM_NUM in each search step of bilateral matching can be determined based on encoding parameters. Alternatively, BM_NUM can be determined based on at least one of motion information, the search step of bilateral matching, the matching cost in the previous search step, the matching cost of the initial motion information in the current search step, or BM_NUM in the previous search step.
[0447] For example, in the first search step of a bilateral match, BM_NUM can be 1 or 2.
[0448] For example, BM_NUM in the current search step can be determined based on the matching cost in the previous search step.
[0449] For example, if the difference between the matching cost of the initial motion information in the previous search step and the matching cost of the improved motion information in the previous search step is less than COSTDIFF_FORBMNUM, then BM_NUM in the current search step can be 0.
[0450] COSTDIFF_FORBMNUM can be 0, 1, 2, 4, 8, 16 or a positive integer.
[0451] For example, COSTDIFF_FORBMNUM can be determined based on the size of the target block. COSTDIFF_FORBMNUM can be the product of the number of pixels in the target block and a specific value. The specific value can be 0, 1, 2, 4, 8, or a positive integer.
[0452] For example, if BM_NUM was 0 in a previous search step, then BM_NUM can be 0 in the current search step.
[0453] For example, if the matching cost of the initial motion information in the current search step is less than COSTDIFF_FORBMNUM_INIT, then the BM_NUM of the target block can be 0.
[0454] COSTDIFF_FORBMNUM can be 0, 1, 2, 4, 8, 16 or a positive integer.
[0455] For example, COSTDIFF_FORBMNUM_INIT can be determined based on the size of the target block. COSTDIFF_FORBMNUM_INIT can be the product of the number of pixels in the target block and a specific value. The characteristic value can be 0, 1, 2, 4, 8, or a positive integer.
[0456] For example, when motion improvement is performed in the search step of bilateral matching, motion information improvement for motion information in the L0 direction can only be performed if the matching cost of motion information in the L0 direction of the initial motion information in the current search step is greater than COSTDIFF_FORBMNUM_INIT.
[0457] For example, when motion improvement is performed in the search step of bilateral matching, motion information improvement for motion information in the L1 direction can only be performed if the matching cost of motion information in the L1 direction of the initial motion information in the current search step is greater than COSTDIFF_FORBMNUM_INIT.
[0458] COSTDIFF_FORBMNUM can be 0, 1, 2, 4, 8, 16 or a positive integer.
[0459] For example, COSTDIFF_FORBMNUM_INIT can be determined based on the size of the target block. COSTDIFF_FORBMNUM_INIT can be the product of the number of pixels in the target block and a specific value. The specific value can be 0, 1, 2, 4, 8, or a positive integer.
[0460] When BM_NUM is 0 in a specific search step of a bilateral matching, it can indicate that motion information improvement is not performed in that specific search step. Alternatively, when BM_NUM is 0 in a specific search step of a bilateral matching, it can indicate that that specific search step is not performed.
[0461] For example, when performing bilateral matching, BM_NUM in the motion information derivation step for the entire block can be 1, and BM_NUM in the motion information derivation step for the sub-block can be 2. In this case, in the motion information derivation step for the entire block, only the motion information in the LXBM direction can be improved, and in the motion information derivation step for the sub-block, both the motion information in the L0 direction and the motion information in the L1 direction can be improved.
[0462] Figure 22 A bilateral matching according to an embodiment is shown.
[0463] Figure 22 This illustrates the case where BM_NUM is 2 in the motion information derivation step for the entire block during bilateral matching.
[0464] MV0 can be the initial motion information in the L0 direction.
[0465] MV1 can be the initial motion information in the L1 direction.
[0466] MVdiff can refer to the improved value of motion information derived through bilateral matching.
[0467] MV0' and MV1' can be motion information derived through bilateral matching.
[0468] In bilateral matching, the magnitude of the motion information improvement value in the L0 direction can be the same as the magnitude of the motion information improvement value in the L1 direction. The directions of the motion information improvement values in the L0 and L1 directions can be opposite to each other. In other words, the following equation can be established.
[0469] MV0' = MV0 + MVdiff MV1'=MV1-MvDIFF Subsampling in the embodiments may be performed based on at least one of the following: whether template matching is performed; an indicator indicating whether template matching is performed; whether bilateral matching is performed; an indicator indicating whether bilateral matching is performed; motion information of the target block; encoding parameters of the target block; or a search step in template matching and / or bilateral matching.
[0470] For example, the subsampling method can be determined based on at least one of the following: whether template matching is performed; an indicator indicating whether template matching is performed; whether bilateral matching is performed; an indicator indicating whether bilateral matching is performed; motion information of the target block; encoding parameters of the target block; or search steps in template matching and / or bilateral matching.
[0471] For example, when performing template matching and / or bilateral matching, the subsampling method used can be the same for the search steps.
[0472] Optionally, for example, when performing template matching and / or bilateral matching, the subsampling method used can be different for each search step.
[0473] Whether to perform subsampling in the embodiments can be determined based on at least one of the following: whether template matching is performed; an indicator indicating whether template matching is performed; whether bilateral matching is performed; an indicator indicating whether bilateral matching is performed; motion information of the target block; encoding parameters of the target block; or a search step in template matching and / or bilateral matching.
[0474] For example, whether to perform subsampling can be determined based on at least one of the following: whether template matching is performed; an indicator indicating whether template matching is performed; whether bilateral matching is performed; an indicator indicating whether bilateral matching is performed; motion information of the target block; encoding parameters of the target block; or search steps in template matching and / or bilateral matching.
[0475] For example, whether or not subsampling is performed can be the same for the search step when performing template matching and / or bilateral matching.
[0476] Optionally, for example, whether to perform subsampling when performing template matching and / or bilateral matching can be different for each search step.
[0477] For example, whether to perform subsampling in the horizontal direction and whether to perform subsampling in the vertical direction can be the same.
[0478] For example, whether to perform subsampling in the horizontal direction and whether to perform subsampling in the vertical direction can be different.
[0479] The subsampling method in the search step that performs a search on the entire target block in template matching and / or bilateral matching can be the same or different from the subsampling method in the search step that performs a search on sub-blocks within the target block in template matching and / or bilateral matching.
[0480] The motion information derivation method on the decoder side may include the template matching method described above, the bilateral matching method described above, the method for improving motion information, the method for reordering the motion information list, and the method for specifying the position in the search area.
[0481] In methods for improving motion information, second motion information can be derived by modifying first motion information. Furthermore, a prediction block for the target block can be generated by performing a prediction using the second motion information. At least one of the reference blocks for the target block can be specified using the second motion information.
[0482] In this embodiment, refinement of specific information can refer to modifying, correcting, or updating specific information. In this embodiment, the terms "refinement," "modification," and "correction" are used interchangeably. Refinement of specific information can be performed to generate refined information.
[0483] Performing a correction for specific motion information may mean performing at least one of the following methods for the corresponding motion information.
[0484] As an example, specific information included in the corresponding motion information can be changed to a predetermined offset.
[0485] As an example, as a result of performing specific operations on the corresponding information and the predetermined offset, specific information included in the corresponding motion information can be changed.
[0486] Specific operations may include at least one of squaring, weighted average, weighted sum, arithmetic operations, or filtering.
[0487] The predetermined offset may include at least one of motion vector, reference frame index and inter-frame prediction indicator, reference frame list information, reference image, motion vector candidate, motion vector candidate index, merge candidate and merge index, block vector, block vector candidate or block vector candidate index.
[0488] For example, a predetermined offset can be specified from a list of predetermined offset candidates.
[0489] Information that can be sent / encoded / decoded using signals to specify a predetermined offset can be used.
[0490] Signals can be sent / encoded / decoded to specify the index of at least one offset from a predetermined list of offset candidates.
[0491] For a list of N motion information items, the method for reordering the list of motion information items based on the matching cost of each motion information item in the list can be a method for reordering at least one of the motion information items in the list.
[0492] For example, the order of motion information in a list of N motion information can be reordered in ascending order of matching cost.
[0493] For example, the order of at least one motion information in a list of N motion information can be reordered in ascending order of matching cost.
[0494] For example, the list can be reconfigured after at least one piece of motion information in a list consisting of N pieces of motion information is reordered.
[0495] When a specific list of motion information is reconfigured, this may mean at least one of the following: 1) configuring the list by using only some of the motion information that configured the corresponding list of motion information; 2) removing at least one of the motion information that constitutes the corresponding list of motion information from the list; or 3) inserting at least one new piece of motion information into the corresponding list of motion information.
[0496] For example, in 3) above, the new motion information can be motion information determined based on at least one motion information in the motion information list. As an example, the new motion information can be the average or weighted average of the two motion information with the lowest index in the list.
[0497] Optionally, for example, in 3) above, the new motion information can be the default motion information.
[0498] N can be 2, 3, 4, 6, 8, 12 or a positive integer.
[0499] At least one motion information can be specified from the reordered list.
[0500] A predicted block for the target block can be generated by performing a prediction using motion information specified from a reordered list.
[0501] At least one of the reference blocks of the target block can be specified by using motion information specified from a reordered list.
[0502] For example, a list consisting of N motion information items can refer to a candidate list for merging motion information of the target block. Alternatively, a list consisting of N motion information items can refer to a list configured using a decoder-side motion information derivation method.
[0503] The method for specifying a location within a search area can be a method for specifying at least one location within a search area based on a matching cost, or a method for specifying at least one of some locations within a search area based on a matching cost.
[0504] In this embodiment, position and motion information can be used interchangeably. In this embodiment, a specific position can be used interchangeably with motion information indicating a movement from a target block to a corresponding position. In this embodiment, a specific position can refer to motion information indicating a movement from a target block to a corresponding position. In this embodiment, specific motion information can refer to a position from the target block indicated by the corresponding motion information.
[0505] In an embodiment, a specified location may refer to a sample at a specified location.
[0506] In an embodiment, a specified location may refer to motion information indicating the movement from the target block to the corresponding location.
[0507] For example, the decoder-side motion information derivation method can be to select N positions with the lowest matching cost among the positions in the specified search area.
[0508] The matching cost at a specific location can refer to the matching cost of motion information from the target block to the corresponding location.
[0509] For example, you can specify N locations within the search area that have the lowest matching cost. Alternatively, you can specify, for example, N locations within a subset of locations in the search area that have the lowest matching cost.
[0510] N can be 1, 2, 3, 4, 5, 6, 8, 12 or a positive integer.
[0511] A list can be configured with N specified positions. At least one position can be specified in the configured list, and predictions in the target block can be performed based on that position. Information from the list can be sent / encoded / decoded using signals.
[0512] A mapping model can be a model that shows the association from a first set of variables to a second set of variables.
[0513] The first set of variables can be called the set of independent variables (or the set of independent variables). The set of independent variables can be a set of at least one variable that affects or is expected to affect other variables (dependent variables).
[0514] The second set of variables may be referred to as the set of dependent variables (or the set of dependent variables). The set of dependent variables may be a set of at least one variable, wherein the value of said at least one variable is determined or expected to be determined based on other variables (independent variables).
[0515] A mapping model can be a model that shows how dependent variables change as independent variables change.
[0516] The set of independent samples can be the set of independent variables of the mapping model. Alternatively, the set of independent variables of the mapping model can be determined based on the set of independent samples.
[0517] The dependent sample set can be the set of dependent variables of the mapping model. Alternatively, the set of dependent variables of the mapping model can be determined based on the dependent sample set.
[0518] The test sample set can be a set of variables to which the mapping model is applied to predict at least one sample of the target block and / or to change at least one sample value of the target block. Optionally, the set of variables to which the mapping model is applied to predict at least one sample of the target block and / or to change at least one sample value of the target block can be determined based on the test sample set.
[0519] As an example, variables may include sample values, and when variables are sample values, the set of independent variables and the set of dependent variables may be represented as the set of independent samples and the set of dependent samples. Here, a sample value can be any one of the sample values of a single sample, a value obtained by modifying a sample value using neighboring samples, or a representative value of multiple samples. Here, a representative value can be any one of the mean, median, maximum, and minimum values.
[0520] Prediction of at least one sample in a target block can be performed based on a set of test samples or a mapping model that computes the association between independent and dependent sample sets.
[0521] Prediction of at least one sample of the target block can be performed based on the results obtained by inputting a set of test samples into a mapping model or a set of results obtained.
[0522] At least one of the sample values of the target block can be changed based on the test sample set or a mapping model that calculates the association between the independent sample set and the dependent sample set.
[0523] At least one of the sample values of the target block can be changed based on the results obtained by inputting the set of test samples into the mapping model or the set of results obtained.
[0524] Figure 22 This is a flowchart of an encoding method for a target block according to an embodiment.
[0525] The target block can be encoded by at least one of the following operations: configuring a sample set and deriving a mapping model; applying the mapping model to the target block; or entropy encoding the encoded information.
[0526] An embodiment may be a part of a video encoding method or an encoding method for a target block.
[0527] Figure 23 This is a flowchart of an encoding method for a target block according to an embodiment.
[0528] The target block can be encoded by at least one of the following operations: deriving a mapping model; applying a mapping model; or entropy encoding the encoded information.
[0529] An embodiment may be a part of a video encoding method or an encoding method for a target block.
[0530] Figure 24 This is a flowchart of a decoding method for a target block according to an embodiment.
[0531] The target block can be decoded by at least one of the following operations: entropy decoding of the encoded information; configuring a sample set and deriving a mapping model; or applying the mapping model to the target block.
[0532] An embodiment may be a part of a video decoding method or a decoding method for a target block.
[0533] Figure 25 This is a flowchart of a decoding method for a target block according to an embodiment.
[0534] The target block can be decoded by at least one of the following operations: entropy decoding of the encoded information; deriving a mapping model; or applying a mapping model.
[0535] An embodiment may be a part of a video decoding method or a decoding method for a target block.
[0536] Model-based prediction methods may include at least one of the following operations: configuring at least one of an independent sample set, a dependent sample set, or a subject sample set; configuring a mapping model input for each sample / location within the independent sample set; configuring a mapping model target output for samples / locations within the dependent sample set corresponding to each sample / location within the independent sample set; deriving the coefficients of the mapping model to minimize the sum (or average) of the differences between the input and target output for all samples / locations within the independent sample set; configuring a mapping model input for each sample / location within the subject sample set; or, after inputting a mapping model for each sample / location within the subject sample set and outputting a result value from the mapping model, using it to perform (at least one of prediction, reconstruction, or decoding) and / or (a change in at least one value of prediction, reconstruction, or decoding) in the target block.
[0537] The input structure can specify how the input is configured when configured for a given sample (or location). The output structure can specify the form of the output for the given input.
[0538] The input structure of a mapping model may include information including at least one of the following: information about at least one input component of the mapping model (including gradient information of the input component), the relationship between at least two input components of the mapping model (e.g., the positional relationship between a first input component and a second input component; e.g., a method for deriving the second input component from the first input component), the relationship between at least one input component of the mapping model and at least one output component of the mapping model, or information about the number of input components of the mapping model (the size of the input).
[0539] The output structure of a mapping model may include information including at least one of the following: information about each output component of the mapping model, the relationship between at least two output components of the mapping model, the relationship between at least one input component of the mapping model and at least one output component of the mapping model, and information about the number of output components of the mapping model (the size of the output).
[0540] Information about a particular input (or output) component may include at least one of the following: the type of the corresponding input (or output) component, the method used to derive the corresponding input (or output) component, the method used to determine the corresponding input (or output) component, its relationship to another input (or output) component (e.g., positional relationship, derivation method), or its relationship to at least one output (or input) component.
[0541] The type of a particular input (or output) component may include at least one of the following: the value of the sample and / or a value based on the value of the sample; the position of the sample and / or a value based on the position of the sample; or a bias.
[0542] The sample can be (at least one of the target block, the neighboring block of the target block, the reference block of the target block, or the neighboring block of the reference block of the target block) (at least one of the luminance component or chrominance component) (at least one of the prediction, reconstruction sample, residual sample, or decoding sample).
[0543] The value based on a specific sample can refer to the result of performing a specific operation on the value of a specific sample or the value that includes the value of a specific sample.
[0544] The value based on the location of a specific sample can refer to the result of performing a specific operation on the location of the specific sample or the value of the location including the location of the specific sample.
[0545] As an example, the input structure can have four input components. When configuring the input for a specific sample, each of the four input components can be the corresponding sample, the left sample of the corresponding sample, the right sample of the corresponding sample, and the bias.
[0546] Figure 26 Examples of the first input structure and the second input structure are shown.
[0547] When the first input structure includes the second input structure, this may mean that when the input for the same sample (or location) is configured using the first input structure and the second input structure, the input using the first input structure includes all input components of the input using the second input structure.
[0548] Optionally, when the first input structure includes the second input structure, this may mean that when the input for the same sample (or location) is configured using the first input structure and the second input structure, the input using the first input structure includes at least one input component of the input using the second input structure.
[0549] Offsets can be used with the same meaning and can be interchanged with each other. Offsets can indicate motion information of the target block and / or the target.
[0550] Here, each of the bias and bias values can be a determined value based on at least one of the following: the encoding parameters of the block, the size of the target block, the range of values that the luminance component of the target block can have, the range of values that the chrominance component of the target block can have, the sample of the target block, the sample of the co-located block of the target block, the availability of the neighboring blocks of the target block, the encoding parameters of the neighboring blocks of the target block, or the neighboring samples of the target block.
[0551] For example, at least one of the bias values can be a value obtained by dividing the range of possible values of the luminance component (or chrominance component) of the target block by a specific value. The specific value can be 2, 4, or a positive integer.
[0552] For example, at least one of the bias values can be the median of the values that the luminance component (or chrominance component) of the target block can have.
[0553] For example, at least one of the bias values may be a value of a sample existing at a specific location based on the target block or a value of a sample existing at a specific location based on the co-location block.
[0554] Figure 27 It is a diagram used to describe a specific location.
[0555] Reference Figure 27 At least one of the bias values can exist in Figure 27 The value of a sample based on at least one of the shadow position of the target block and the shadow position of the co-position block.
[0556] The size of a block can refer to at least one of the following: the width of the block, the height of the block, or the number of samples in the block.
[0557] Here, a specific operation may include at least one of exponentiation, weighted average, weighted sum, arithmetic operations, or filtering.
[0558] For example, a particular operation may include at least one of a weighted average, a weighted sum, or arithmetic operations with a particular value.
[0559] A specific value can be determined based on at least one of the following: the prediction mode of the target block, the motion information of the target block, the encoding parameters of the target block, the size of the target block, the range of possible values for the luminance component of the target block, the range of possible values for the chrominance component of the target block, the availability of neighboring blocks of the target block, the encoding parameters of the neighboring blocks of the target block, or the neighboring samples of the target block.
[0560] For example, a specific value can exist in Figure 27 The value of at least one of the samples in the shaded position.
[0561] For example, a specific value could be the median of the range of values that the luminance (or chrominance) component of the target block can have.
[0562] The position of a sample can refer to its relative position to the top-left sample of the target block or its relative position to the top-left sample of the image. The position within the image can refer to its relative position to the top-left sample of the image.
[0563] For example, a specific value could be the x-component and / or y-component of the position of the top-left sample of the target block.
[0564] For example, the dynamic range of the values computed during the derivation of the mapping model can be limited by using a second value, which is the result of performing a specific operation (e.g., subtraction, multiplication, division) on the first value, instead of the first value. Therefore, the complexity of the operation can be reduced, and the size of the memory required for the operation can be limited.
[0565] For example, it may be possible to allow mapping patterns to express more complex relationships by using a second value instead of the first value, which is the result of performing a specific operation (e.g., exponentiation) on the first value.
[0566] Here, the value based on the first value can refer to either the first value or the result of performing a specific operation on the first value.
[0567] Configuring a sample set can include selecting at least one sample.
[0568] Optionally, configuring the sample set may include determining the region of the sample, and deriving the mapping model by using at least one of the following: the value of at least one sample within the corresponding region; the location information of at least one sample; the value of at least one sample; the motion information of at least one sample; or the value of the motion information and / or location information based on the location of at least one sample.
[0569] Figure 28 The reference region used to derive the mapping model is shown.
[0570] Configuring a sample set by using samples within a reference region may include: determining Figure 28 The reference region is determined by using the values of at least one sample within the corresponding region to derive the mapping model. The sample set can refer to at least one of the independent sample set, the dependent sample set, or the subject sample set.
[0571] Information for determining at least one of the independent sample set, dependent sample set, or subject sample set can be predefined in the encoder and decoder. The sample set used to derive the mapping model can be configured according to predefined rules in the encoder and decoder; and / or the sample set to which the mapping model will be applied can be configured.
[0572] The input to the mapping model can be configured for each sample (or location) within the (independent sample set and / or subject sample set).
[0573] If the input to the mapping model is configured for each sample (or location) within an independent sample set, then the target output for the corresponding input can be configured for each sample (or location) within a subordinate sample set corresponding to each sample (or location) within the subject sample set.
[0574] Figure 29 This shows a sample of the computational output of the mapping model.
[0575] Reference Figure 29 When the input to the mapping model is configured for the target sample and the output of the mapping model is computed using the target sample and samples around the target sample, it can be considered the same as performing a cross-shaped filter centered on the location of the target sample.
[0576] Each of the independent sample set, the dependent sample set, and the subject sample set may include at least one of the following: a sample in the reference region, a sample in the target block, or a sample in the block corresponding to the target block.
[0577] For example, the reference region may exclude samples within the reference region of the target block, or it may include at least one sample within the reference region of the target block.
[0578] For example, a reference region may exclude samples within a reference region of a reference block, or it may include at least one sample within a reference region of a reference block.
[0579] For example, the reference region can be at least one of the regions adjacent to the target block, the reference region, or the block corresponding to the target block.
[0580] In this article, "adjacent" does not necessarily mean "adjacent".
[0581] Figure 28 The reference region in the example can represent the reference region in the model-based prediction method according to the embodiment.
[0582] As an example, LINENUM_L, LINENUM_LB, LINENUM_A, and LINENUM_RA can be 0, 1, 2, or positive integers, respectively.
[0583] As an example, LINENUM_L, LINENUM_LB, LINENUM_A, and LINENUM_RA can be predefined values.
[0584] As an example, LINENUM_MORE_X can be 0, 1, 2 or a positive integer (X is A, B, L, R).
[0585] As an example, LINENUM_MORE_X can be a predefined value. Optionally, information about LINENUM_MORE_X can be sent / encoded / decoded using signals.
[0586] As an example, when configuring the input of a mapping model for specific samples within an independent set of samples, if samples located at positions up to N apart in a specific direction are used based on the position of the corresponding sample, the value of LINENUM_MORE_X for that direction can be determined to be N. N can be 0, 1, or a positive integer. The specific direction can include at least one of top, bottom, left, or right.
[0587] The size and / or location of the reference region can be determined based on at least one of the following: the prediction pattern of the target block, the motion information of the target block, the encoding parameters of the target block, the size of the target block, the range of possible values for the luminance component of the target block, the range of possible values for the chrominance component of the target block, the availability of neighboring blocks of the target block, the encoding parameters of the neighboring blocks of the target block, and the neighboring samples or mapping model of the target block.
[0588] The size of the reference area may include at least one of LINENUM_L, LINENUM_LB, LINENUM_A, LINENUM_RA, or LINENUM_MORE_X.
[0589] For example, at least one of LINENUM_MORE_A, LINENUM_MORE_B, LINENUM_MORE_L, or LINENUM_MORE_R can be determined based on the input structure in a model-based prediction model.
[0590] The configuration of the reference region may differ from the configuration shown. As an example, in Figure 28In this configuration, the reference region can be configured using only the top adjacent region, or it can be configured using only the left adjacent region. The configuration of the reference region can vary depending on the target to be predicted by the matching model.
[0591] Reference regions can also be configured to not be adjacent to the target block. Furthermore, after specifying multiple reference region candidates, one of these candidates can be selected. A reference region candidate can be at least one of the following: a region within the target image; a region within the reference image; or a region within the co-located image. As an example, the availability of a region spaced at a predetermined distance from the target block can be determined, and if the corresponding region is available, it can be inserted into the list as a reference region candidate. Here, the region spaced at a predetermined distance can be a region of a specific size, where the position obtained by spacing the horizontal and vertical distances from one of the top and left boundaries of the current block by a unit distance is considered the center / top-left position. As another example, the availability of a region spaced at a predetermined distance from a co-located block can be determined, and if the corresponding region is available, it can be inserted into the list as a reference region candidate. Here, the region spaced at a predetermined distance can be a region of a specific size, where the position obtained by spacing the horizontal and vertical distances from one of the top and left boundaries of the co-located block by a unit distance is considered the center / top-left position. Here, the unit distance can be N, 2N, 3N, 4N, etc., and N can be 4, 8, 16, etc. A region of a specific size can be 4×4, 8×8, or 16×16.
[0592] When multiple candidates are inserted into the list, information indicating one of the multiple reference region candidates can be encoded and sent with a signal.
[0593] Information representing the configuration of a reference region can be encoded and transmitted via signals. This information can indicate whether at least one of the top adjacent region, left adjacent region, upper left adjacent region, or lower right adjacent region is included in the reference region. As an example, this information can be an index indicating one of the configuration candidates for the reference region. Configuration candidates can include one of the following: a candidate consisting of the top and left adjacent regions, a candidate consisting only of the top adjacent regions, a candidate consisting only of the left adjacent regions, a candidate consisting of the top and upper right regions, and a candidate consisting of the left and lower left regions.
[0594] The block corresponding to the target block may refer to at least one of the following: a reference block of the target block; a chroma component block corresponding to the target block when the target block is a luma component block; or a luma component block corresponding to the target block when the target block is a chroma component block, but the type of block corresponding to the target block is not limited to these.
[0595] The reference region in a model-based prediction method can be the same as the reference region in other prediction methods. Alternatively, the reference region in a model-based prediction method can be different from the reference region in other prediction methods. Other prediction methods may include template matching.
[0596] In this paper, "determining specific information based on a mapping model" can mean determining specific information based on at least one of the following: MODEL_INPUT_NUM of the mapping model; MODEL_OUTPUT_NUM of the mapping model; at least one bias value in the mapping model; a method for configuring the input of the mapping model; the input structure of the mapping model; the output structure of the mapping model; a method for performing (at least one of prediction, reconstruction, and decoding) on samples within a target block based on the output of the mapping model; or a method for changing the value of (at least one of prediction, reconstruction, and decoding) on samples within a target block based on the output of the mapping model, but the meaning of "determining specific information based on a mapping model" is not limited to these.
[0597] Each of the independent sample set, dependent sample set, and test sample set may include at least one of the following: a predicted sample, a reconstructed sample, a residual sample, or a decoded sample, of at least one of the following: a target block, a neighboring block of the target block, a reference block of the target block, and a neighboring block of the reference block of the target block. Additionally, the sample may be associated with at least one of the luminance or chrominance components.
[0598] When a specific sample is included in a specific sample set, this may mean that at least one of the following is included in the specific sample set: motion information in a block containing the sample values of the corresponding specific sample; motion information of the corresponding specific sample; encoding parameters of the corresponding specific sample; prediction mode of the corresponding specific sample; x-component of the sample's position; y-component of the sample's position; and motion information of the corresponding specific sample or at the corresponding specific sample's position. Alternatively, when a specific sample is included in a specific sample set, this may mean that it is included in the specific sample set based on the value of at least one of the following: motion information in a block containing the sample values of the corresponding specific sample; motion information of the corresponding specific sample; encoding parameters of the corresponding specific sample; prediction mode of the corresponding specific sample; x-component of the sample's position; y-component of the sample's position; and motion information of the corresponding specific sample and at the corresponding specific sample's position.
[0599] "A specific sample is included in a specific sample set" and "A specific sample is used to configure a specific sample set" can be used with the same meaning and can be used interchangeably in the embodiments.
[0600] "Value based on a specific value" can refer to the result of performing a specific operation on a specific value.
[0601] Predicted samples can refer to the predicted luminance component and / or the predicted chrominance component. Reconstructed samples can refer to the reconstructed luminance component and / or the reconstructed chrominance component. Residual samples can refer to the residual luminance component and / or the residual chrominance component. Decoded samples can refer to the decoded luminance component and / or the decoded chrominance component.
[0602] The reference block may include at least one of the following: a reference block in the L0 direction of inter-frame prediction, a reference block in the L1 direction of inter-frame prediction, a reference block in intra-frame template matching, or a reference block in intra-frame block copying.
[0603] Intra-frame template matching mode can refer to a template matching method for each of the reference block, reference template, and reference region, including at least one of the predicted sample, reconstructed sample, residual sample, or decoded sample of the target image.
[0604] Information about the method for configuring at least one of the independent sample set, dependent sample set, or test sample set can be transmitted / encoded / decoded using signals.
[0605] A method for configuring a specific sample set can refer to a method that includes at least one of the following: 1) the number of samples used to configure the sample set and / or the maximum number of samples used to configure the sample set; 2) the size and / or location of the region where the samples for configuring the sample set exist; 3) the shape, size, and / or location of the region where the samples for configuring the sample set are selected; 4) the type of at least one sample in the samples used to configure the sample set; 5) whether a bias value is included; 6) a method for determining the bias value; 7) a filling method when configuring the corresponding sample set; or 8) the size and / or location of a reference region.
[0606] Methods for configuring a specific sample set can be categorized based on the following: 1) the number of samples used to configure the sample set and / or the maximum number of samples used to configure the sample set; 2) the size and / or location of the region where the samples for configuring the sample set exist; 3) the shape, size, and / or location of the region where the samples for configuring the sample set are selected; 4) the type of at least one sample in the samples used to configure the sample set; 5) whether a bias value is included; 6) the method used to determine the bias value; 7) the filling method when configuring the corresponding sample set; or 8) the size and / or location of the reference region.
[0607] Information regarding the method for configuring the sample set may include at least one of the following: 1) the number of samples used to configure the sample set and / or the maximum number of samples used to configure the sample set; 2) the size and / or location of the region where the samples for configuring the sample set exist; 3) the shape, size, and / or location of the region where the samples for configuring the sample set are selected; 4) the type of at least one sample in the samples used to configure the sample set; 5) whether a bias value is included; 6) the method used to determine the bias value; 7) the filling method when configuring the corresponding sample set; or 8) the size and / or location of the reference region.
[0608] Information about a method for configuring at least one of the following can be determined based on the target block's prediction pattern, target block motion information, target block encoding parameters, target block size, range of possible values for the target block's luminance component, range of possible values for the target block's chrominance component, availability of the target block's neighboring blocks, encoding parameters of the target block's neighboring blocks, or neighboring samples of the target block.
[0609] For example, information about a method for configuring at least one of the following can be determined based on the target block's prediction pattern, target block motion information, target block encoding parameters, target block size, range of possible values for the target block's luminance component, range of possible values for the target block's chrominance component, availability of the target block's neighboring blocks, encoding parameters of the target block's neighboring blocks, and neighboring samples or mapping models of the target block.
[0610] For example, a method for configuring at least one of the independent sample set, dependent sample set, or subject sample set can be determined based on whether neighboring blocks in a specific direction are available to the target block.
[0611] Whether it is available can include whether it exists. The specific direction can be one of the following: top left, top, top right, left, right, bottom left, bottom, or bottom right.
[0612] When configuring at least one of an independent sample set, a dependent sample set, or a subject sample set, it can be determined whether to include at least one sample from the samples in the corresponding specific direction in the corresponding sample set, based on whether at least one neighboring block located in a specific direction is available to the target block.
[0613] For example, based on the ratio of the width to the height of the target block, a method for configuring at least one of the independent sample set, the dependent sample set, or the subject sample set can be determined.
[0614] When configuring at least one of an independent sample set, a dependent sample set, or a test sample set, it can be determined whether to include at least one sample in a specific direction in the corresponding sample set based on the ratio of the width to the height of the target block.
[0615] A specific direction can be one of the following: top left, top, top right, left, right, bottom left, bottom, and bottom right.
[0616] When configuring at least one of an independent sample set, a dependent sample set, or a subject sample set, the size and / or location of the reference region can be determined based on the ratio of the width to the height of the target block.
[0617] For example, a method for configuring at least one of an independent sample set, a dependent sample set, or a subject sample set can be determined based on at least one of the prediction patterns, motion information, or coding parameters of the target block.
[0618] When configuring at least one of an independent sample set, a dependent sample set, or a subject sample set, the type of at least one sample in the sample set used for configuring the sample set can be determined based on at least one of the target block's prediction pattern, motion information, or coding parameters.
[0619] The sample type may include at least one of the following: at least one of the following: a prediction sample, a reconstruction sample, a residual sample, or a decoded sample of at least one of the following: a target block, a neighboring block of the target block, a reference block of the target block, or a neighboring block of the reference block of the target block.
[0620] For example, when the target block is in a geometric partitioning mode, a method for configuring at least one of the independent sample set, dependent sample set, or subject sample set can be determined.
[0621] For example, when the target block is in a geometric partitioning pattern, at least one of the independent sample set, dependent sample set, or subject sample set can be determined based on at least one of the information about the geometric partitioning pattern.
[0622] Information about geometric partitioning patterns may refer to at least one of the partitioning boundaries, partitioning patterns, partitioning angles, or partitioning offsets of the geometric partitioning pattern, but information about geometric partitioning patterns is not limited to these.
[0623] For example, for the boundary in a specific direction of the target block, it can be determined whether to use the neighboring samples in the specific direction to configure at least one of the independent sample set, dependent sample set, or subject sample set based on the sample values of at least one neighboring sample in BOUND_LINE_NUM rows / columns and / or at least one sample in BOUND_LINE_NUM rows / columns of the target block.
[0624] BOUND_LINE_NUM can be 1, 2, or a positive integer. BOUND_LINE_NUM can be determined based on at least one of the following: motion information of the target block, encoding parameters of the target block, size of the target block, range of possible values for the luminance component of the target block, range of possible values for the chrominance component of the target block, availability of neighboring blocks of the target block, encoding parameters of neighboring blocks of the target block, or neighboring samples of the target block.
[0625] Figure 30 It is a diagram used to describe the target block and the neighboring samples within BOUND_LINE_NUM rows / columns, as well as the samples within BOUND_LINE_NUM rows / columns of the target block.
[0626] Figure 31 An example is shown that determines whether to use top neighbor samples when configuring various sample sets based on at least one of the neighbor samples in the two top rows / columns of the target block or the samples in the two rows / columns within the target block.
[0627] Reference Figure 31 For the top-direction boundary of the target block, it can be determined whether to use the top-direction neighboring samples to configure at least one of the independent sample set, dependent sample set, or subject sample set based on the sample values of at least one neighboring sample in two adjacent rows / columns and / or at least one sample in two rows / columns within the target block.
[0628] exist Figure 31 In the middle, if If the value is equal to or greater than a certain threshold, then neighboring samples in the top direction of the target block may not be used to configure the independent sample set, dependent sample set, and subject sample set. For example, if If the value is less than a certain threshold, at least one of the independent sample set, dependent sample set, or subject sample set can be configured using at least one neighboring sample in the top direction of the target block.
[0629] In the embodiment, because the block size is 8×8, x has a value from 0 to 7, but even when the block size is 8×8, x can have a value in a different range and can have values determined at intervals greater than 1 (e.g., 0, 2, 4, 6).
[0630] A mapping model can be a regression model. For example, a mapping model can be a filter.
[0631] For example, a mapping model can be derived through regression analysis.
[0632] For example, to reduce the complexity of deriving the mapping model, at least one of LU decomposition, LDL decomposition, or Gaussian elimination can be used.
[0633] A mapping model can be a model that outputs a result value with MODEL_OUTPUT_NUM components based on an input consisting of MODEL_INPUT_NUM components.
[0634] When a specific sample is included in the input of a mapping model, this can mean that at least one of the following is included in the specific sample set: the sample value of the corresponding specific sample; the motion information of the corresponding specific sample; the encoding parameters of the corresponding specific sample; the prediction pattern of the corresponding specific sample; the x-component of the sample's position; the y-component of the sample's position; motion information in the block containing the corresponding specific sample or motion information at the position of the corresponding specific sample. Alternatively, when a specific sample is included in the specific sample set, this can mean that a value based on at least one of the following is included in the input: the motion information of the corresponding specific sample; the encoding parameters of the corresponding specific sample; the prediction pattern of the corresponding specific sample; the x-component of the sample's position; the y-component of the sample's position; motion information in the block containing the corresponding specific sample or motion information at the position of the corresponding specific sample.
[0635] "Specific samples are included in the input" and "Specific samples are used to configure the input" can be used with the same meaning and can be used interchangeably in the embodiments.
[0636] "Result of the mapping model" and "output of the mapping model" can be used with the same meaning and can be used interchangeably in the embodiments.
[0637] MODEL_INPUT_NUM and MODEL_OUTPUT_NUM can both be positive integers.
[0638] MODEL_INPUT_NUM and MODEL_OUTPUT_NUM can be predefined values. Optionally, information about at least one of MODEL_INPUT_NUM or MODEL_OUTPUT_NUM can be sent / encoded / decoded using signals.
[0639] The mapping model can be a linear model. Alternatively, the mapping model can be a nonlinear model.
[0640] Mapping model f, input and result value The relationship between them can be represented as shown in the following examples.
[0641] /
[0642]
[0643] ...
[0644]
[0645] It can refer to the coefficients of the mapping model f.
[0646] i can have values from 1 to MODEL_INPUT_NUM.
[0647] j can have values from 1 to MODEL_OUTPUT_NUM.
[0648] When deriving a mapping model f, it can mean deriving at least one of the coefficients of the mapping model f.
[0649] For a specific sample within an independent sample set, an input consisting of MODEL_INPUT_NUM components can be determined, including at least one of the following: the corresponding sample; neighboring samples adjacent to the corresponding sample; neighboring samples not adjacent to the corresponding sample; a bias value; or the result of performing a specific operation on at least one of the above elements.
[0650] For example, at least one of the neighboring samples and / or at least one of the bias values used to configure inputs for a specific sample in an independent sample set may not be included in the independent sample set.
[0651] For example, the input can include bias values that are independent of a particular sample.
[0652] For a specific sample in an independent sample set, a target output consisting of MODEL_OUTPUT_NUM components can be determined, including at least one of the following: a sample in the subordinate sample set corresponding to the corresponding sample; a neighboring sample adjacent to a sample in the subordinate sample set corresponding to the corresponding sample; a neighboring sample that is not adjacent (not contiguous) to a sample in the subordinate sample set corresponding to the corresponding sample; or a bias value.
[0653] A mapping model can be a model that calculates the association between the input and the target output for each sample in an independent set of samples.
[0654] When the input and target output are determined for each sample in an independent set of samples, a mapping model can be derived to ensure that the result value of the mapping model for the input of all samples is similar to the target output on average.
[0655] For a specific sample in an independent sample set, when using the corresponding sample's neighboring samples as input, if the specific neighboring sample is unavailable, then at least one determined sample from the available samples whose position is closest to the corresponding neighboring sample, instead of the specific neighboring sample, can be used as input. This process can be expressed as performing padding for the corresponding neighboring sample.
[0656] A sample in a subordinate sample set corresponding to a specific sample in an independent sample set can refer to at least one of the following: a sample in a subordinate sample set corresponding to the position of the corresponding sample; and / or at least one of the target outputs determined for the corresponding sample.
[0657] For example, when configuring independent sample sets and / or dependent sample sets, NUM_SAMPLE_TO_REF samples can be used.
[0658] NUM_SAMPLE_TO_REF can be a fixed constant value. Alternatively, NUM_SAMPLE_TO_REF can be a value adaptively determined based on at least one of the following: the prediction mode of the target block, the motion information of the target block, the encoding parameters of the target block, the size of the target block, the range of possible values for the luma component of the target block, the range of possible values for the chroma component of the target block, the availability of neighboring blocks of the target block, the encoding parameters of the neighboring blocks of the target block, or the neighboring samples of the target block.
[0659] For example, if the size of the target block is less than or equal to a specific value, it can have the value NUM_SAMPLE_TO_REF=(H+W); otherwise, it can have the value NUM_SAMPLE_TO_REF=(H+W) / STEP_TO_REF.
[0660] STEP_TO_REF can be 2, 4, 8 or a positive integer.
[0661] Optionally, for example, if the size of the target block is less than or equal to a specific value, it can have the value NUM_SAMPLE_TO_REF=(H+W); otherwise, it can have the value NUM_SAMPLE_TO_REF=NUM_SAMPLE_TO_REF_CONST.
[0662] NUM_SAMPLE_TO_REF_CONST can be 8, 16, 32, 64, 128, or a positive integer.
[0663] The size of the target block can refer to at least one of the following: the target block's height, width, the product of height and width, the sum of height and width, the larger of height and width, the smaller of height and width, or the diagonal length.
[0664] For a specific sample within the subordinate sample set, a target output consisting of MODEL_OUTPUT_NUM components can be determined, including at least one of the following: the corresponding sample; neighboring samples adjacent to the corresponding sample; neighboring samples not adjacent to the corresponding sample; a bias value; or the result of performing a specific operation on at least one of the above elements.
[0665] For example, at least one of the neighboring samples and / or at least one of the bias values used to configure inputs for a specific sample in the subordinate sample set may not be included in the independent sample set.
[0666] For a specific sample in the dependent sample set, an input consisting of MODEL_INPUT_NUM components can be determined, including at least one of the following: a sample in the independent sample set corresponding to the corresponding sample; a neighboring sample adjacent to a sample in the independent sample set corresponding to the corresponding sample; a neighboring sample that is not adjacent (not contiguous) to a sample in the dependent sample set corresponding to the corresponding sample; a bias value; or the result of performing a specific operation on at least one of the above elements.
[0667] A mapping model can be a model that calculates the association between the input and the target output for each sample in the subordinate sample set.
[0668] A mapping model can be derived to ensure that, when the input and target output are determined for each sample in the subordinate sample set, the result value of the mapping model for the input of all samples is similar to the target output on average.
[0669] For a specific sample in a subset of samples, when using the neighboring samples of the corresponding sample as the target output, if the specific neighboring sample is unavailable, then at least one determined sample from the available samples whose position is closest to the corresponding neighboring sample, instead of the specific neighboring sample, can be used as the target output. This process can be expressed as performing padding for the corresponding neighboring sample.
[0670] A sample in the independent sample set corresponding to a specific sample in the subordinate sample set can refer to at least one of the following: a sample in the independent sample set corresponding to the position of the corresponding sample; and / or at least one of the inputs determined for the corresponding sample.
[0671] Mapping model f, input and target output The relationship between them can be represented as shown in the following examples. It can refer to the error between the target output and the output value of the mapping model for the input.
[0672]
[0673] The mapping model f can be derived to minimize the sum of all samples in the dependent sample set or some samples in the dependent sample set. The sum or average.
[0674] For example, when deriving a mapping model, at least one coefficient of the mapping model can be determined as a value selected from a predefined candidate list of DEFAULT_COEF_NUM values. DEFAULT_COEF_NUM can be 1, 2, 3, or a positive integer. The value of DEFAULT_COEF_NUM can be predefined, or information about DEFAULT_COEF_NUM can be signaled / encoded / decoded. For example, information about which of the DEFAULT_COEF_NUM candidate values will be used to determine the coefficients of the mapping model can be signaled / encoded / decoded.
[0675] The input to the mapping model may include at least one of the following: 1) the value of at least one sample in the independent sample set and / or the value based on that value; 2) the position of at least one sample in the independent sample set and / or the value based on that position; 3) the value of at least one sample in the target sample set and / or the value based on that value; 4) the position of at least one sample in the target sample set and / or the value based on that position; or 5) a bias.
[0676] Information about model-based prediction methods stored in the target block can be stored. The model-based prediction method can be executed in a block encoded / decoded after the target block based on the information stored in the target block. Alternatively, the model-based prediction method can be executed in the target block based on information about the model-based prediction method in a block encoded / decoded before the target block.
[0677] It may be information including at least one of the following: whether a model-based prediction method is performed in the target block and / or an indicator indicating whether a model-based prediction method is performed in the target block; a method for configuring at least one of the independent sample set, dependent sample set, or subject sample set in the target block; at least one of the samples, locations, or sample values included in the independent / dependent / subject sample set in the target block; a method for determining a reference region in the target block; mapping model information in the target block; whether a multi-mapping model is used in the target block and / or an indicator indicating whether a multi-mapping model is used in the target block; or the number and / or maximum number of mapping models in the target block.
[0678] The mapping model information may include at least one of the following: the mapping model's MODEL_INPUT_NUM; the mapping model's MODEL_OUTPUT_NUM; at least one coefficient of the mapping model; the mapping model's input structure; or the mapping model's output structure.
[0679] The type of a particular variable can be related to a component of that variable. For example, the type of a particular variable can be classified based on at least one of whether an operation is performed, the type of operation performed, or the type of a component.
[0680] The type of component may include at least one of the following: a luminance component and a chrominance component of the target block, a neighboring block of the target block, a reference block of the target block, or a neighboring block of the reference block of the target block; a prediction sample, a reconstructed sample, a residual sample, or a decoded sample, or at least one of at least one bias value.
[0681] Based on the information about the model-based prediction method stored in the second block, which was encoded / decoded prior to the first block, the model-based prediction method in the first block can be executed. Alternatively, based on the information about the model-based prediction method stored in the second block, which was encoded / decoded prior to the first block, the information about the model-based prediction method in the first block can be executed.
[0682] Here, "inheriting information about the model-based prediction method from the first block in the second block" can mean executing the model-based prediction method in the first block based on the information about the model-based prediction method stored in the second block (encoded / decoded before the first block). Alternatively, here, "inheriting information about the model-based prediction method from the first block in the second block" can mean executing the information about the model-based prediction method in the first block based on the information about the model-based prediction method stored in the second block (encoded / decoded before the first block).
[0683] For example, at least one coefficient of at least one mapping model contained in the information about the model-based prediction method stored in the second block can be determined as one of the coefficients of the mapping model in the first block.
[0684] The information about the model-based prediction method inherited from the first block in the second block can be considered as the derivation of the mapping model f in the first block.
[0685] For example, a portion of the coefficients of at least one mapping model contained in the information about model-based prediction methods stored in the second block can be identified as a portion of the coefficients of the mapping model in the first block. The remaining coefficients of the mapping model in the first block can then be derived through regression analysis.
[0686] Information about the model-based prediction method for the current block can be inherited from at least one candidate block. A candidate block can include at least one of a neighboring block adjacent to the current block or a non-neighboring block not adjacent to the current block.
[0687] As an example, such as Figure 27 As shown, adjacent blocks can include at least one of the blocks that are adjacent to the top, top left, top right, left, or bottom left of the current block.
[0688] Figure 32 This is a diagram showing the non-adjacent blocks of the current block. For example... Figure 32 In the example shown, non-adjacent blocks can include blocks at a distance of N from at least one of the top or left boundaries of the current block. N can be an integer, such as 4, 8, or 16. After deriving merge candidates from adjacent blocks, merge candidates can be derived using non-adjacent blocks.
[0689] Optionally, a non-adjacent block can be a block indicated by at least one of the following: the block vector of the target block, the motion vector of the target block, the block vector of the luma component block corresponding to the target block when the target block is a chroma component block, the motion vector of the luma component block corresponding to the target block when the target block is a chroma component block, the block vector of a neighboring block, the motion vector of a neighboring block, the block vector around the target block, or the motion vector around the target block. For example, a neighboring block can be an adjacent block.
[0690] For example, for at least two blocks indicated by the block vector of the target block, the motion vector of the target block, the block vector of the luma component block corresponding to the target block when the target block is a chroma component block, the motion vector of the luma component block corresponding to the target block when the target block is a chroma component block, the block vector of the neighboring block, the motion vector of the neighboring block, the block vector around the target block, or the motion vector around the target block, a merge candidate can be derived from the N blocks among them that have the lowest template matching cost.
[0691] Merge candidates can also be derived from temporally adjacent blocks. For example, in... Figure 27 In the example shown, a time-adjacent block can be a block in the same frame that corresponds to the position of the current block (a co-position block), or it can be an adjacent block of the co-position block.
[0692] For example, deriving a merge candidate by using a specific candidate block may mean that the specific candidate block is used as a merge candidate. Alternatively, for example, deriving a merge candidate by using a specific candidate block may mean that the mapping model used in the specific candidate block, or a portion thereof, is used as a merge candidate.
[0693] Merge candidates can also be derived from temporally adjacent blocks. For example, in... Figure 27In the example shown, a time-adjacent block can be a block in the same frame that corresponds to the position of the current block (a co-position block), or it can be an adjacent block of the co-position block.
[0694] As an example, temporally adjacent blocks can be blocks located to the right, bottom, lower right, upper left, or lower left of the same block.
[0695] There may be N temporally adjacent blocks used to derive merge candidates for the target block. N can be 0, 1, 2, 3, 6, 12, or a positive integer. Additionally, there may be M merge candidates derived from the temporally adjacent blocks in the target block. M can be 0, 1, 2, 3, 6, 12, or a positive integer.
[0696] The number of temporally adjacent blocks can be adaptively determined based on the number of available spatially adjacent blocks.
[0697] Merge candidates can also be derived from non-adjacent blocks of the same position. As an example, merge candidates can be derived from time blocks that are at least one unit distance or N times the unit distance from the top-left position of the same position, either horizontally or vertically. For example, if the top-left position of the same position is (x, y), merge candidates can be derived from time blocks at the top-left positions (x + (N*horizontal), y), (x, y + ((N*vertical)), or (x + (N*horizontal), y + (N*vertical)).
[0698] Here, "horizontal" represents the unit distance in the horizontal direction, and "vertical" represents the unit distance in the vertical direction. The unit distance can be determined based on the size of the target block. As an example, the unit distance in the horizontal direction and the unit distance in the vertical direction can be set to be the same as the width and height of the target block, respectively.
[0699] N can be a natural number greater than or equal to 1. If the derived time block is unavailable when the first value is applied to N, the time block can be searched by incrementing N by 1.
[0700] Merge candidates can be derived from the block at the shifted position of the time block (hereinafter referred to as the shifted time block). Here, the shifted time block can be located at the position of the motion vector at which the time block is spaced apart. Here, the motion vector can be the motion vector of the spatially adjacent block of the target block. As an example, the shifted time block can be determined based on the motion vector of the left-hand or top-hand adjacent block of the target block. Optionally, when searching for spatially adjacent blocks of the target block according to a predefined order, the shifted time block can be determined based on the motion vector of the first available block found. If no available block is found among the spatially adjacent blocks, the time block at the original position can be used without performing a shift.
[0701] Whether a time-adjacent block is available can be determined based on at least one of the frame / strip type (e.g., whether it is an I-frame / I-strip) or the type of the matching model.
[0702] Based on information about the prediction methods for at least one candidate block, a prediction method information merging list can be configured. Information about candidate blocks or their prediction methods can be inserted as merging candidates into this list.
[0703] Default prediction method information can be inserted into a prediction method information merging list. As an example, the default prediction method information can be a predefined set of default scaling parameters in the encoder / decoder. Alternatively, the default prediction method information can be derived by adding a predefined offset to the merging candidates pre-inserted into the prediction method information merging candidate list, or subtracting a predefined offset from the merging candidates pre-inserted into the prediction method information merging candidate list. Here, the pre-inserted merging candidate can be the merging candidate with the smallest index or the merging candidate with the largest index among the merging candidates inserted into the prediction method information merging candidate list. A default prediction method can be inserted when the number of merging candidates included in the prediction method information merging candidate list is less than a threshold. The threshold can be the maximum value of the merging candidates that can be included in the prediction method information merging candidate list, or a value M less than the maximum value. M can be an integer such as 1, 2, etc.
[0704] As an example, even if merge candidates are inserted into the prediction method information merge list according to a predefined order, if the number of merge candidates is less than a threshold, the default prediction method information can be inserted into the prediction method information merge list. As an example, the predefined order could be the order of spatially adjacent blocks, spatially non-adjacent blocks, and prediction method information stored in the cumulative prediction method information table described later.
[0705] Here, a non-adjacent block can refer to a block that is at least one unit distance or N times the unit distance from a specific location of the target block or a specific location around the target block, either horizontally or vertically. As an example, as shown in Figure 40, if the top-left position of the target block is (x, y), merge candidates can be derived from spatially non-adjacent blocks at the top-left positions (x-(N*horizontal), y), (x, y-(N*vertical)), or (x-(N*horizontal), y-(N*vertical)).
[0706] Here, "horizontal" represents the unit distance in the horizontal direction, and "vertical" represents the unit distance in the vertical direction. The unit distance can be determined based on the size of the target block. As an example, the unit distance in the horizontal direction and the unit distance in the vertical direction can be set to be the same as the width and height of the target block, respectively.
[0707] N can be a natural number greater than or equal to 1. If spatially non-adjacent blocks are unavailable when the first value is applied to N, spatially non-adjacent blocks can be searched by incrementing N by 1.
[0708] Prediction method information for blocks predicted based on a mapping model can be accumulated and stored in a prediction method information table. For example, if a block is predicted based on a mapping model, the prediction method information table can be updated using the prediction method information for that block. The number of candidates to be included in the prediction method information table can be predefined in the encoder and decoder. The prediction method information table can be initialized by predefined units. Here, the predefined units can be CTUs, parallel blocks, stripes, CTU rows, or sub-pictures.
[0709] Alternatively, the prediction method information table may not store prediction method information for blocks predicted based on the mapping model, but rather blocks predicted based on the mapping model.
[0710] When configuring the prediction method information merging list, the prediction method information table can be used. For example, if the number of merging candidates included in the prediction method information table is less than a threshold, the prediction method information stored in the prediction method information table can be inserted into the prediction method information merging list as new merging candidates.
[0711] If the prediction method information merging list includes multiple merging candidates, information for one of the specified multiple merging candidates can be encoded and sent using a signal.
[0712] Alternatively, the prediction method information table itself can be used as the prediction method information merging list. When multiple candidates are included in the prediction method information merging list, information for one of the specified candidates can be encoded and sent using a signal. As an example, information indicating whether the prediction method information table itself is used as the prediction method information merging list can be sent / encoded / decoded using a signal.
[0713] In the target block, information indicating whether a prediction method information merging list is used can be encoded and sent via a signal. This information can be a 1-bit flag. If the information indicates that the prediction method information merging list is used, the prediction method information of the current block can be inherited from the merge candidate. If the information indicates that the prediction method information merging list is not used, the prediction method information of the current block can be derived from the reference area.
[0714] After reconstructing the target block, a mapping model for the target block can be derived based on samples of the reconstructed target block.
[0715] In the blocks encoded / decoded after the target block, the derived mapping model of the target block can be referenced.
[0716] In blocks encoded / decoded after the target block, information about the model-based prediction method can be inherited from the derived mapping model of the target block.
[0717] As an example, the mapping model for a target block could be a mapping model representing the relationship from an independent set of predicted samples of at least one luminance component of the target block to a dependent set of reconstructed samples of at least one luminance component of the target block. As another example, the mapping model for a target block could be a mapping model representing the relationship from an independent set of reconstructed samples of at least one luminance component of the target block to a dependent set of reconstructed samples of at least one chrominance component of the target block. As yet another example, the mapping model for a target block could be a mapping model representing the relationship from an independent set of predicted samples of at least one chrominance component of the target block to a dependent set of reconstructed samples of at least one chrominance component of the target block.
[0718] In this case, the mapping model can only be derived in the same manner as described above when the model-based prediction method is executed in the target block. Optionally, the mapping model can be derived in all blocks in the same manner as described above, regardless of whether the model-based prediction method is executed in the target block. Optionally, if the target block satisfies certain conditions, the mapping model can be derived from the target block in the same manner as described above.
[0719] Specific conditions may be associated with at least one of the following: the prediction mode of the target block, the motion information of the target block, the coding parameters of the target block, the size of the target block, the range of possible values for the luma component of the target block, the range of possible values for the chroma component of the target block, the availability of neighboring blocks of the target block, the coding parameters of the neighboring blocks of the target block, and the neighboring samples or mapping model of the target block. As an example, a specific condition may be related to whether the prediction mode of the target block is inter-frame prediction. As an example, a specific condition may be related to whether the target block is a luma component block.
[0720] Here, "deriving a mapping model for a specific block" can mean at least one of the following: 1) deriving at least one mapping model for the entire corresponding block; 2) dividing the corresponding block into specific units and deriving at least one mapping model for each divided block; or 3) dividing the corresponding block into specific units and deriving at least one mapping model for at least one of the divided blocks. In this case, the specific unit may include at least one of a TU unit and a sub-block unit of size N×M. N and M above can be 1, 2, 4, 8, 16, 32, or a positive integer.
[0721] Predictions using the first mapping model can be performed in the target block. Subsequently, a second mapping model for the target block can be derived from the independent sample set and / or dependent sample set that includes the reconstructed samples of the target block, and the second mapping model can replace (or update) the first mapping model.
[0722] As an example, the second mapping model could be a mapping model representing the relationship between an independent set of predicted samples of at least one luminance component including the target block and a subordinate set of reconstructed samples of at least one luminance component including the target block. As another example, the second mapping model could be a mapping model representing the relationship between an independent set of reconstructed samples of at least one luminance component including the target block and a subordinate set of reconstructed samples of at least one chrominance component including the target block. As yet another example, the second mapping model could be a mapping model representing the relationship between an independent set of predicted samples of at least one chrominance component including the target block and a subordinate set of reconstructed samples of at least one chrominance component including the target block.
[0723] In the blocks encoded / decoded after the target block, the second mapping model of the target block can be referenced.
[0724] In blocks encoded / decoded after the target block, information about the model-based prediction method can be inherited from the second mapping model of the target block.
[0725] A mapping model for the target block can be derived and stored from an independent set of samples and / or a set of dependent samples that include reconstructed samples of the target block.
[0726] It can store mapping models related to luminance component prediction and luminance component reconstruction.
[0727] For example, after reconstructing the luminance component samples of the target block (or reconstructing the target block as a luminance component block), a mapping model for the target block can be derived and stored from the independent sample set including the luminance component prediction samples of the target block and the dependent sample set including the luminance component reconstruction samples of the target block.
[0728] Information regarding model-based prediction methods can be inherited from the mapping model of the target block derived by this method from the first block encoded / decoded after the target block. Subsequently, in the first block, the values of at least one luminance component prediction samples in the first block can be modified based on a set of subject samples including at least one luminance component prediction sample from the first block.
[0729] In this case, there is an effect of predicting the residual signal of the first block of luminance components and adding it to the predicted luminance component of the first block, thereby improving coding efficiency.
[0730] It can store mapping models related to chromaticity component prediction and chromaticity component reconstruction.
[0731] For example, after reconstructing the chromaticity component samples of the target block (or reconstructing the target block as a chromaticity component block), a mapping model for the target block can be derived and stored from the independent sample set including the chromaticity component prediction samples of the target block and the dependent sample set including the chromaticity component reconstruction samples of the target block.
[0732] Information about model-based prediction methods can be inherited from the mapping model of the target block derived by the method from the first block encoded / decoded after the target block. Then, in the first block, the values of at least one chromaticity component prediction samples in the first block can be changed based on a set of subject samples including at least one chromaticity component prediction sample from the first block.
[0733] In this case, there is an effect of predicting the residual signal of the first block of chroma components and adding it to the predicted chroma components of the first block, thereby improving coding efficiency.
[0734] Luminance component reconstruction and chrominance component reconstruction can be stored.
[0735] For example, after reconstructing the chromaticity component samples of the target block (or reconstructing the target block as a chromaticity component block), a mapping model for the target block can be derived and stored from an independent sample set including the chromaticity component reconstructed samples of the target block (or the reconstructed samples of the luminance component blocks corresponding to the target block) and a dependent sample set including the chromaticity component reconstructed samples of the target block.
[0736] Information about model-based prediction methods can be inherited from the mapping model of the target block derived by the method from the first block encoded / decoded after the target block. Then, in the first block, the value of at least one luminance component prediction sample of the first block can be changed based on a set of subject samples including at least one luminance component reconstruction sample of the first block (or, reconstruction samples of the luminance component block corresponding to the target block).
[0737] In this case, coding efficiency can be improved by utilizing the information redundancy between the luminance and chrominance components.
[0738] When using a prediction method information merging list in a target block, signal sending / encoding / decoding can be used to specify merging candidates for inheriting the mapping model information that will be used in the target block.
[0739] For example, the information used to specify the merge candidates for inheriting the mapping model information to be used in the target block can be the index of the corresponding merge candidate in the prediction method information merge list.
[0740] When using a prediction method information merging list in a target block, merging candidates for inheriting the mapping model information to be used in the target block can be determined based on at least one of the following: target block size, whether the target block is a luma component block, whether the target block is a chroma component block, target stripe type, neighboring blocks of the target block, availability of neighboring blocks of the target block, neighboring samples of the target block, luma component samples of the target block, chroma component samples of the target block, availability of neighboring samples of the target block, or encoding parameters of the target block.
[0741] There can be one or more merge candidates that inherit mapping model information from the target block. As an example, there can be 6 merge candidates.
[0742] Furthermore, based on the prediction modes of the target block and / or the corresponding luminance / chrominance component blocks, it can be determined whether a mapping-based prediction method is permitted to be performed in the target block.
[0743] As an example, a model-based prediction method can only be performed in the target block if the target block and / or the corresponding luma / chroma component block is encoded / decoded via intra-frame prediction.
[0744] As another example, even if the target block and / or the corresponding luma / chroma component block is predicted via inter-frame prediction, IBC, or template-based intra-frame prediction (intraTMP), a model-based prediction method can be performed in the target block.
[0745] Depending on the encoding / decoding mode of the target block and / or the corresponding luminance / chrominance component blocks, at least one of the information regarding the model-based prediction method may differ.
[0746] For example, information about model-based prediction methods may refer to at least one of the following: a method for configuring at least one of the independent / dependent / subject sample sets, the coefficients of the mapping model, the input structure of the mapping model, the output structure of the mapping model, the type of the mapping model, a method for deriving the coefficients of the mapping model, or a method for deriving the mapping model.
[0747] As an example, if the target block and / or the corresponding luma / chroma component block is predicted via inter-frame prediction, IBC, or intraTMP, a nonlinear model can be used as the mapping model in the target block, but a linear model is not required.
[0748] As an example, if the target block and / or the corresponding luma / chroma component block is predicted via inter-frame prediction, IBC, or intraTMP, a nonlinear model can be used in a fixed manner.
[0749] As an example, if the target block and / or the corresponding luma / chroma component block is predicted via inter-frame prediction, IBC, or intraTMP, then the method for selecting the mapping model by using a merged candidate list is not required.
[0750] At least one of the information regarding the model-based prediction method may differ between the case where the target block and / or the corresponding luma / chroma component block is predicted via intra-frame prediction and the case where the target block and / or the corresponding luma / chroma component block is predicted via other prediction modes (i.e., inter-frame prediction, IBC, or intraTMP).
[0751] As an example, the number of coefficients used in the mapping model or the method used to derive the coefficients may differ between the case where the target block and / or the corresponding luma / chroma component block is predicted by intra-frame prediction and the case where the target block and / or the corresponding luma / chroma component block is predicted by other prediction modes (i.e., inter-frame prediction, IBC, or intraTMP).
[0752] Here, the target block and the current block can be used with the same meaning and can be used interchangeably.
[0753] When configuring the merged candidate list, the available blocks can be limited based on the prediction patterns of independent sample sets.
[0754] As an example, when encoding / decoding the independent sample set of the current block via intra-frame prediction, only the temporal / spatial blocks that are encoded / decoded via intra-frame prediction in the temporal / spatial blocks can be used to derive merging candidates.
[0755] As an example, if the independent set of samples for the current block is predicted by inter-frame prediction, IBC, or intraTMP, then merge candidates can be derived using only the temporal / spatial blocks predicted by inter-frame prediction, IBC, or intraTMP within the temporal / spatial blocks.
[0756] When configuring the merge candidate list, the available blocks can be limited based on the prediction modes of the target block and / or the corresponding luminance / chrominance component blocks.
[0757] As an example, when encoding / decoding a target block and / or the corresponding luma / chroma component block via intra-frame prediction, merge candidates can be derived using only the temporal / spatial blocks within the temporal / spatial blocks that are encoded / decoded via intra-frame prediction for the corresponding block or the corresponding luma / chroma component block.
[0758] As an example, when a target block and / or a corresponding luma / chroma component block is encoded / decoded by at least one of inter-frame prediction, IBC, or intraTMP, merge candidates can be derived using only the temporal / spatial blocks in which the corresponding block or the corresponding luma / chroma component block is encoded / decoded by at least one of inter-frame prediction, IBC, or intraTMP.
[0759] As an example, when encoding / decoding a target block and / or the corresponding luma / chroma component block by intra-frame prediction, the merging candidates can be derived using only the blocks included in the prediction method information table that are encoded / decoded by intra-frame prediction for the corresponding block or the corresponding luma / chroma component block.
[0760] As an example, when encoding / decoding a target block and / or a corresponding luma / chroma component block using at least one of inter-frame prediction, IBC, or intraTMP, merging candidates can be derived using only the blocks included in the prediction method information table that are encoded / decoded using at least one of inter-frame prediction, IBC, or intraTMP for the corresponding block or the corresponding luma / chroma component block.
[0761] As an example, when encoding / decoding a target block and / or a corresponding luma / chroma component block via intra-frame prediction, the merging candidate can be derived using only the prediction method information included in the prediction method information table, specifically the prediction method information included in the prediction method information table for encoding / decoding blocks using the corresponding prediction method information or luma / chroma component blocks corresponding to blocks using the corresponding prediction method information.
[0762] As an example, when encoding / decoding a target block and / or a corresponding luma / chroma component block using at least one of inter-frame prediction, IBC, or intraTMP, merging candidates can be derived using only the prediction method information included in the prediction method information table, specifically the prediction method information included in the prediction method information table for encoding / decoding blocks using the corresponding prediction method information or luma / chroma component blocks corresponding to blocks using the corresponding prediction method information.
[0763] Optionally, even when the target block and / or the corresponding luma / chroma component block is encoded / decoded via intra-frame prediction, not only the temporal / spatial blocks within the temporal / spatial blocks that are encoded / decoded via intra-frame prediction for the corresponding block or the corresponding luma / chroma component block, but also the temporal / spatial blocks within the temporal / spatial blocks that are predicted via at least one of inter-frame prediction, IBC, or intraTMP for the corresponding block or the corresponding luma / chroma component block, can be used to derive merging candidates.
[0764] Optionally, even when the target block and / or the corresponding luma / chroma component block is encoded / decoded by at least one of inter-frame prediction, IBC, or intraTMP, not only the temporal / space blocks in the temporal / space blocks that are predicted by at least one of inter-frame prediction, IBC, or intraTMP, but also the temporal / space blocks in the temporal / space blocks that are encoded / decoded by intra-frame prediction can be used to derive merging candidates.
[0765] Optionally, for example, even when encoding / decoding the target block and / or the corresponding luma / chroma component block via intra-frame prediction, not only the blocks included in the temporal / spatial block and prediction method information table that are encoded / decoded via intra-frame prediction for the corresponding block or the corresponding luma / chroma component block, but also the blocks included in the temporal / spatial block and prediction method information table that are predicted via at least one of inter-frame prediction, IBC, or intraTMP for the corresponding block or the corresponding luma / chroma component block, can be used to derive merging candidates.
[0766] Optionally, for example, even when the target block and / or the corresponding luma / chroma component block is encoded / decoded by at least one of inter-frame prediction, IBC, or intraTMP, not only the blocks included in the temporal / spatial block and prediction method information table that are predicted by at least one of inter-frame prediction, IBC, or intraTMP, but also the blocks included in the temporal / spatial block and prediction method information table that are encoded / decoded by intra-frame prediction for the corresponding block or the corresponding luma / chroma component block, can be used to derive merging candidates.
[0767] Optionally, for example, even when encoding / decoding the target block and / or the corresponding luma / chroma component block via intra-frame prediction, not only the temporal / spatial blocks encoded / decoded via intra-frame prediction or the luma / chroma component blocks corresponding to the blocks using the corresponding prediction method information, or the blocks included in the prediction method information table, but also the prediction method information included in the prediction method information table that predicts via at least one of inter-frame prediction, IBC, or intraTMP for blocks using the corresponding prediction method information or the luma / chroma component blocks corresponding to the blocks using the corresponding prediction method information, can be used to derive merging candidates.
[0768] Optionally, for example, even when the target block and / or the corresponding luma / chroma component block is encoded / decoded by at least one of inter-frame prediction, IBC, or intraTMP, not only the temporal / spatial blocks or blocks included in the prediction method information table that are encoded / decoded by at least one of inter-frame prediction, IBC, or intraTMP using the corresponding prediction method information or the luma / chroma component blocks corresponding to the blocks using the corresponding prediction method information, but also the prediction method information included in the prediction method information table that is predicted by at least one of inter-frame prediction, IBC, or intraTMP using the corresponding prediction method information or the luma / chroma component blocks corresponding to the blocks using the corresponding prediction method information, can be used to derive merging candidates.
[0769] Here, a time / space block may include a time / space block that is adjacent to the target block and a time / space block that is not adjacent to the target block.
[0770] When configuring a merged candidate list for model-based prediction modes, the merged candidate list can be configured to ensure that at least one of the information regarding the model-based prediction method used in all candidate blocks of the merged candidate list is the same. Alternatively, when configuring a merged candidate list for model-based prediction modes, the merged candidate list can be configured to include only blocks whose information regarding the model-based prediction method used is the same as the information regarding a specific model-based prediction method as candidate blocks.
[0771] For example, when configuring a merged candidate list for model-based prediction patterns, the merged candidate list can be configured to ensure that all input structures of the mapping model used in all candidate blocks of the merged candidate list are identical.
[0772] For example, when configuring a merged candidate list for model-based prediction patterns, the merged candidate list can be configured by using only blocks that have a specific input structure or an input structure that includes that specific input structure as candidate blocks.
[0773] The specific model-based prediction method can be a predefined model-based prediction method. Optionally, the specific model-based prediction method can be adaptively determined based on at least one of the following: the prediction pattern of the target block, the motion information of the target block, the encoding parameters of the target block, the size of the target block, the range of possible values for the luminance component of the target block, the range of possible values for the chrominance component of the target block, the availability of neighboring blocks of the target block, the encoding parameters of the neighboring blocks of the target block, or the neighboring samples of the target block.
[0774] Optionally, for all candidate blocks in the merged candidate list of model-based prediction patterns, there may be no information about a common model-based prediction method. In other words, the information about the model-based prediction method for each candidate block in the merged candidate list of model-based prediction patterns may differ from each other.
[0775] For example, at least one of the input structures used in each candidate block of a merged candidate list of model-based prediction patterns may be different from each other.
[0776] Each merge candidate in the model-based prediction pattern merge candidate list may include at least one of the mapping model information for the luminance component, Cb component, and Cr component.
[0777] Optionally, each merge candidate in the merge candidate list of model-based prediction patterns may include only the mapping model of one component among the mapping model information of the luminance component, Cb component, and Cr component.
[0778] For example, the first pooled candidate list of model-based prediction patterns can be configured using only the mapping model information about the Cb component, and the second pooled candidate list of model-based prediction patterns can be configured using only the mapping model information about the Cr component.
[0779] When specifying merge candidates for inheriting mapping model information from the target block from the merge candidate list of model-based prediction modes, the merge candidates for inheriting mapping model information about at least two of the luminance, Cb, or Cr components can always be the same.
[0780] For example, the mapping model information for the Cb component and the mapping model information for the Cr component of a specified merging candidate can be inherited from the merging candidate list of model-based prediction patterns.
[0781] Optionally, when specifying merging candidates for inheriting mapping model information from the merging candidate list of model-based prediction modes, the merging candidates for inheriting mapping model information may be the same or different from each other for at least two of the luminance component, Cb component, or Cr component.
[0782] For example, the mapping model information regarding the Cb component of a first merging candidate specified from the model-based prediction mode merging candidate list can be inherited, and the mapping model information regarding the Cr component of a second merging candidate specified from the model-based prediction mode merging candidate list can also be inherited. In this case, the first and second merging candidates can be the same as or different from each other.
[0783] Information about a mapping model for a specific component can refer to information about a mapping model used to perform at least one of the prediction, reconstruction, or alteration of the sample values of the corresponding component.
[0784] Optionally, regardless of the prediction mode of the target block and / or the corresponding luminance / chrominance component block, the temporal / spatial block to which the mapping model is applied can be set as a merge candidate.
[0785] In addition, as mentioned above, at least one of the methods used to derive the mapping model and / or mapping model coefficients may be different depending on the prediction mode of the target block and / or the corresponding luminance / chrominance component block.
[0786] As an example, when encoding a target block and / or its corresponding luma / chroma component blocks via intra-frame prediction, a mapping model based on a linear model or a first nonlinear model can be applied to the current block. On the other hand, if the luma / chroma component blocks of a candidate block are predicted via inter-frame prediction, IBC, or intraTMP, the merge candidate derived from the corresponding candidate block can be a second nonlinear model. When a merge candidate derived from a candidate block is selected from the merge candidate list, common information between the mapping model (e.g., a linear model or a first nonlinear model) applicable only to the current block and the mapping model indicated by the corresponding merge candidate (e.g., a second nonlinear model) can be inherited.
[0787] For example, if the input structure of the mapping model used in the first model-based prediction mode includes the input structure of the mapping model used in the second model-based prediction mode, then in the first model-based prediction mode, at least one of the information about the model-based prediction method used in the second model-based prediction mode can be inherited.
[0788] For example, if the input structure of the mapping model used in the first model-based prediction mode includes the input structure of the mapping model used in the second model-based prediction mode, then in the first model-based prediction mode, the block executing the second model-based prediction mode can be used as a candidate block.
[0789] For example, when a first model-based prediction mode inherits at least one piece of information about the model-based prediction method used in a second model-based prediction mode, there may be information about the model-based prediction method that is used in the first model-based prediction mode but not in the second model-based prediction mode. In this case, the corresponding information can be determined as a predetermined value.
[0790] For example, when a first model-based prediction model inherits at least one piece of information about the model-based prediction method used in a second model-based prediction model, there may be at least one input component used in the first model-based prediction model but not in the second model-based prediction model. In this case, in the first model-based prediction model, the coefficients of the mapping model for the corresponding input type can be determined and / or changed to predetermined values.
[0791] Optionally, for example, when the first model-based prediction mode inherits at least one piece of information about the model-based prediction method used in the second model-based prediction mode, there may be at least one input component used in the first model-based prediction mode but not in the second model-based prediction mode. In this case, the input structure in the mapping model in the first model-based prediction mode may be determined and / or changed to an input structure that does not use the corresponding input type.
[0792] Optionally, for example, when the first model-based prediction mode inherits at least one piece of information about the model-based prediction method used in the second model-based prediction mode, there may be at least one input component used in the first model-based prediction mode but not in the second model-based prediction mode. In this case, the value corresponding to the corresponding input component in the input of the mapping model in the first model-based prediction mode can always be determined as a predetermined value (as an example, it can be replaced by a bias).
[0793] In the above embodiments, the predetermined value can be 0, 1, 2, 256, 512, 1024 or an integer.
[0794] For example, when performing prediction on a target chroma component block, each of the first model-based prediction mode and the second model-based prediction mode can be a model-based prediction mode that generates a predicted value of CHROMA_PIX by using at least one of the following as input: component 1) position or gradient information of a specific sample CHROMA_PIX within the target chroma component; component 2) information of the luminance component sample LUMA_PIX corresponding to CHROMA_PIX; component 3) neighboring luminance samples of LUMA_PIX; component 4) the result value of applying a predetermined filter centered on the position of the luminance component corresponding to CHROMA_PIX; component 5) a bias value; or component 6) the result value of performing a specific operation on at least one of component 1, component 2, component 3, component 4, or component 5.
[0795] In components 1 through 4, LUMA_PIX can be a subsampled luminance component sample corresponding to CHROMA_PIX. Alternatively, in components 1 through 4, LUMA_PIX can be a non-subsampled luminance component sample corresponding to CHROMA_PIX.
[0796] Figure 33 A filter based on the position of the luminance component is shown.
[0797] For example, the predetermined filter in component 4 above could be a filter used to calculate the gradient. Figure 33 One of the filters in [the system]. In [the context of the filter]. Figure 33 In this context, C can refer to the position of the luminance component corresponding to CHROMA_PIX.
[0798] The embodiments can be applied when the input structure in the first model-based prediction mode includes the input structure in the second model-based prediction mode (or when the input structure in the first model-based prediction mode includes the input structure in the second model-based prediction mode and the input structure in the first model-based prediction mode is different from the input structure in the second model-based prediction mode).
[0799] For example, the first model-based prediction mode could be a model-based prediction mode that generates predicted values for CHROMA_PIX by using components 1 through 6 as inputs. On the other hand, the second model-based prediction mode could be a model-based prediction mode that generates predicted values for CHROMA_PIX by using only components 1 and 5 as inputs.
[0800] As another example, the first model-based prediction mode could be a model-based prediction mode that always uses components 2 and 5 as inputs and may also use at least one of components 1, 3, 4, or 6 as inputs. On the other hand, the second model-based prediction mode could be a model-based prediction mode that uses only components 1, 2, and 5 as inputs to produce predicted values for CHROMA_PIX.
[0801] The prediction method information merging list can be configured by taking into account the components of the target block.
[0802] The component may include at least one of the luminance component or the chrominance component.
[0803] The chromaticity component may include at least one of the Cb component, Cr component, R component, G component, or B component.
[0804] The prediction method information merging list can be configured separately for each component. For example, if the target block is a Cb component block, the prediction method information merging list can be configured by using only the mapping model used for the prediction of the Cb component.
[0805] Optionally, an integrated prediction method information merging list can be configured for multiple chromaticity components. For example, a single prediction method information merging list can be configured for Cb component blocks and Cr component blocks.
[0806] Information about whether the prediction method information merging list is used in the matching model can be transmitted / encoded / decoded using a signal. In this case, the same indicator can be used to determine whether the prediction method information merging list is used in both the luminance and chrominance components.
[0807] Optionally, in this case, whether the prediction method information merging list is used in the luminance component and whether the prediction method information merging list is used in the chrominance component can be determined from separate indicators. Similarly, whether the prediction method information merging list is used in the Cb component and whether the prediction method information merging list is used in the Cr component can be determined from separate indicators.
[0808] In the matching model, signal transmission / encoding / decoding is available to specify merging candidates from the prediction method information merging list. In this case, merging candidates for model-based prediction methods in the luminance component and merging candidates for model-based prediction methods in the chrominance component can be determined from the same index.
[0809] Optionally, in this case, merging candidates for model-based prediction methods in the luminance component and merging candidates for model-based prediction methods in the chrominance component can be determined from separate indices. Similarly, merging candidates for model-based prediction methods in the Cb component and merging candidates for model-based prediction methods in the Cr component can be determined from separate indices.
[0810] The predicted block for the current block can be obtained by performing a weighted sum over multiple blocks. In this case, multiple blocks can be generated based on multiple mapping models. Each block can be generated by a different mapping model. As an example, the predicted block for the current block can be obtained by averaging / weighting the first predicted block obtained based on the first mapping model and the second predicted block obtained based on the second mapping model.
[0811] As an example, the first mapping model can be any of a linear model, a nonlinear model, a regression model, and a filter, and the second mapping model can be another.
[0812] Information about whether multiple models are used or the number of models used can be sent / encoded / decoded using signals.
[0813] As an example, a predicted block for the target block can be generated by performing a weighted sum on the blocks generated by applying each mapping model to the target block.
[0814] Weighted sums can include averages. The weights applied to each block can be 0, real numbers greater than 0, or negative real numbers less than 0.
[0815] For example, multiple mapping models can be identified as N merging candidates with the lowest indices in the prediction method information merging list. Alternatively, for example, multiple mapping models can be identified as N merging candidates with the lowest mapping costs in the prediction method information merging list. N can be 2, 3, or a positive integer.
[0816] New mapping models can be derived from multiple mapping models. For example, a new mapping model can be derived by performing a weighted sum on multiple mapping models. For example, predictions for a target block can be performed based on the derived mapping model. For example, the derived mapping model can be added to the prediction method information merge list.
[0817] Alternatively, the final prediction block can be obtained by performing a weighted sum on the prediction blocks derived using a mapping model and the prediction blocks derived using another prediction method other than the mapping model (e.g., intra-frame prediction, inter-frame prediction, IBC, or intraTMP).
[0818] Additionally, merging candidates that include combinations of multiple prediction models can be included in the merging candidate list. When selecting a merging candidate, a prediction block can be derived according to the implementation example.
[0819] Optionally, when multiple mapping models are selected, a first mapping model can be selected from a list of merged candidates, and a second mapping model can be determined by using a predefined model or based on information sent by a signal.
[0820] When performing a weighted sum over at least two blocks, the weights used for the weighted sum can be determined on a block-by-block basis, on a predetermined region within a block, or on a sample basis within a block.
[0821] As an example, the weights of blocks can be determined based on a mapping model. Alternatively, the weights of samples belonging to a block can be determined based on a mapping model.
[0822] Optionally, the weight of each block can be determined based on the cost of the template region for the target block of each mapping model. As an example, a first mapping model can be applied to the template region of the target block to obtain a first predicted signal and a first cost for the template region. As an example, the first cost can be the SAD between the first predicted signal and the reconstructed signal in the template region. Alternatively, a second mapping model can be applied to the template region of the target block to obtain a second predicted signal and a second cost for the template region. As an example, the first cost can be the SAD between the first predicted signal and the reconstructed signal in the template region. Then, the first cost and the second cost can be compared to determine the weight of each of the first block obtained by the first mapping model and the second block obtained by the second mapping model.
[0823] As an example, if the first cost is less than the second cost, the weight assigned to the first block can have a greater value than the weight assigned to the second block. Conversely, if the first cost is greater than the second cost, the weight assigned to the second block can have a greater value than the weight assigned to the first block.
[0824] The weights of blocks and / or the weights of samples within a block can be determined simultaneously. Alternatively, the weights of blocks and / or the weights of samples within a block can be determined sequentially. Here, the sequence may include the scan order.
[0825] In this context, the dependent sample set may include sample values determined based on at least one of the following: neighboring samples of the target block; at least one neighboring sample of a block whose weight has already been determined; or the weight of a block whose weight has already been determined. For example, the weight of the same block may be determined at least twice.
[0826] The order in which the weights of the blocks are determined can be based on at least one of the following: the POC of the image to which each block belongs; the orientation of the reference image to which each reference block belongs; the index indicating the motion information of each block; or the matching cost of each block.
[0827] The independent sample set may include at least one neighboring sample from the reference block of the target block. The dependent sample set may include neighboring samples from the target block. The subject sample set may include at least one sample from the reference block of the target block.
[0828] A filter can be applied to a prediction block derived based on at least one mapping model to modify the prediction block.
[0829] Information indicating whether to modify the prediction block can be encoded and sent using a signal. This information can be a 1-bit flag. Optionally, the decision to modify the prediction block can be adaptively determined based on the size / shape of the current block or the type of mapping model used.
[0830] Figure 34 An example of a 3x3 low-frequency band filter used for predictive sample correction is shown.
[0831] Predicted samples can be corrected by applying filters of a predetermined size / shape to the prediction block. The size / shape / type of the filters can be predefined in the decoder / decoder. For example, a 3×3 low-frequency band filter can be used.
[0832] Optionally, the size / shape / type of the filter can be adaptively determined based on at least one of the current block size / shape, the type of mapping model used, or the prediction mode / intra-frame prediction mode of the independent sample set.
[0833] Optionally, multiple filter type candidates can be defined, and an index identifying one of the multiple filter type candidates can be encoded and transmitted with a signal.
[0834] When applying a filter to the left / top boundary of a block, filtering can be performed using pre-reconstructed neighbor samples. When applying a filter to the right / bottom boundary, filtering can be performed by filling the predicted samples at the right / top boundary.
[0835] Optionally, the size / shape of the filter can be set differently depending on the location of the predicted sample. Optionally, if the predicted sample is located at the block boundary, the application of the filter can be omitted.
[0836] When performing a weighted sum on each of the prediction blocks derived from multiple mapping models, a filter can be applied to the weighted prediction blocks. Alternatively, after applying the filter to each of the prediction blocks derived from multiple mapping models, the final prediction block can be obtained by performing a weighted sum on the prediction blocks to which the filter has been applied.
[0837] The order of merge candidates in the prediction method information merge list can be rearranged.
[0838] The order of the merge candidates in the prediction method information merging list can be rearranged based on at least one of the following: the type of the merge candidate; the mapping model information of the merge candidate; the target block; the neighboring samples of the target block; the reference block of the target block; or the reference sample of the target block.
[0839] Merge candidates can be categorized based on the method used to derive them. For example, the types of merge candidates can include at least one of the following: merge candidates derived from neighboring blocks; merge candidates derived from non-neighboring blocks; merge candidates derived from spatially neighboring blocks; merge candidates derived from temporally neighboring blocks; merge candidates derived from a prediction method information table; or merge candidates derived using default prediction method information.
[0840] The rearrangement order of merge candidates in the prediction method information merge list can be determined based on the matching cost of each merge candidate.
[0841] For example, the rearrangement order of merge candidates in the prediction method information merge list can follow the ascending order of the matching cost of each merge candidate.
[0842] The matching cost of a particular merge candidate can be determined based on the second and third sample sets obtained by applying the corresponding merge candidate to the first sample set.
[0843] For example, the matching cost of a particular merge candidate can be determined based on a second set of samples and a third set of samples obtained by applying the corresponding merge candidate to a first set of samples.
[0844] For example, the coefficients to be corrected can be determined based on the differences between the second and third sample sets.
[0845] For example, the first sample set may include neighboring brightness prediction / reconstruction samples of the target block, and the third sample set may include neighboring chromaticity prediction / reconstruction samples of the target block.
[0846] For example, the first sample set may include neighboring luminance / chromaticity prediction samples of the target block, and the third sample set may include neighboring luminance / chromaticity reconstruction samples of the target block.
[0847] To execute a model-based prediction method in a target block, signal transmission / encoding / decoding can be used to specify information about at least one mapping model from a prediction method information merge list. A model-based prediction method in the target block can be executed using at least one specified mapping model.
[0848] Alternatively, for example, to perform a model-based prediction method in the target block, the N mapping models with the lowest indices in the prediction method information merge list can be used. In other words, a model-based prediction method in the target block can be performed using the N mapping models with the lowest indices in the prediction method information merge list.
[0849] Alternatively, for example, to perform a model-based prediction method in the target block, N mapping models with the lowest matching cost in the prediction method information merge list can be used. In other words, a model-based prediction method in the target block can be performed by using the N mapping models with the lowest matching cost in the prediction method information merge list.
[0850] A model-based prediction merging pattern can refer to a pattern that determines information about the model-based prediction method used in a target block based on at least one of the following: 1) the model-based prediction method of the block determined by predefined rules; or 2) information about the model-based prediction method specified from a prediction method information merging list.
[0851] For example, a model-based prediction merging pattern may refer to a model-based prediction pattern that inherits from and uses at least one of the following: 1) a model-based prediction method for a block determined by predefined rules; or 2) information about a model-based prediction method specified from a prediction method information merging list.
[0852] An indicator indicating whether a target block is in a model-based predictive merging mode can be sent / encoded / decoded based on at least one of the following: the size of the target block, whether the target block is a luma component block, whether the target block is a chroma component block, the type of the target strip, the neighboring blocks of the target block, the availability of the neighboring blocks of the target block, the neighboring samples of the target block, the luma component samples of the target block, the chroma component samples of the target block, the availability of the neighboring samples of the target block, or the encoding parameters of the target block.
[0853] If the indicator indicating whether a model-based predictive merging mode is in a first value, the target block can be determined to be in a model-based predictive merging mode. For example, the first value can be 1 or true.
[0854] If the indicator indicating whether a model-based predictive merging pattern is in a second value, the target block can be determined to be not in a model-based predictive merging pattern. For example, the second value can be 0 or false.
[0855] For example, an indicator indicating whether a target block is in a model-based predictive merging mode can be signaled / encoded / decoded based on at least one of the availability of neighboring samples (or neighboring blocks) adjacent to the top boundary of the target block or the availability of neighboring samples (or neighboring blocks) adjacent to the left boundary of the target block.
[0856] As an example, an indicator indicating whether a target block is in a model-based predictive merging mode can only be signaled / encoded / decoded if both a neighboring sample (or neighboring block) adjacent to the upper boundary of the target block and a neighboring sample (or neighboring block) adjacent to the left boundary of the target block are available.
[0857] For example, if an indicator indicating whether model-based prediction is performed in the target block is true, then an indicator indicating whether the target block is in model-based prediction merging mode can be signaled / encoded / decoded.
[0858] If the target block is in a model-based prediction merging mode, information such as blocks, mapping model information, and / or merging candidates in the merging candidate list that specifies the model-based prediction mode can be signaled / encoded / decoded to inherit at least one mapping model information in the target block. This information may include indicators and / or indices for specifying merging candidates in the merging candidate list.
[0859] For example, in the target block, the mapping model information of the Cb component and the mapping model information of the Cr component can both be inherited from the merge candidate specified by the index used to specify the merge candidate in the merge candidate list.
[0860] Optionally, for example, in the target block, the mapping model information of the Cb component can be inherited from the first merge candidate specified by the first index used to specify the merge candidate in the merge candidate list, and the mapping model information of the Cr component can be inherited from the second merge candidate specified by the second index.
[0861] The first index and the second index can be the same as or different from each other. Both the first index and the second index can be sent / encoded / decoded using signals. The first merge candidate and the second merge candidate can be the same as or different from each other.
[0862] When performing model-based predictions in the target block, the mapping model adjustment offset can be applied to the mapping model in the target block.
[0863] For example, when performing a model-based prediction method in a target block, instead of using a first mapping model to perform the model-based prediction method, a second mapping model, which is the result of applying an adjustment offset of the mapping model to the first mapping model, can be used to perform the model-based prediction method.
[0864] In other words, when applying a mapping model adjustment offset in a target block, at least one of the following can be used in the target block: a first mapping model as a predictor, a mapping model adjustment offset for the first mapping model, or a second mapping model after applying the mapping model adjustment offset to the first mapping model.
[0865] In one or more mapping models, the offset adjustment will be applied to the target block. The mapping model for which the offset adjustment is applied to the target block can be a mapping model determined based on the independent sample set and dependent sample set configured for the target block. Alternatively, the mapping model for which the offset adjustment is applied to the target block can be a mapping model determined in a model-based prediction merging pattern.
[0866] If a mapping model adjustment offset is applied to multiple mapping models in a target block, the same mapping model adjustment offset can be applied to all mapping models. Alternatively, when a mapping model adjustment offset is applied to multiple mapping models in a target block, the same mapping model adjustment offset can be applied to all mapping models, or different mapping model adjustment offsets can be applied to at least two of them.
[0867] Within the target block, information related to the offset adjustment of the mapping model can be obtained by signal transmission / encoding / decoding based on at least one of the following: the size of the target block, whether the target block is a luma component block, whether the target block is a chroma component block, the type of the target strip, the neighboring blocks of the target block, the availability of the neighboring blocks of the target block, the neighboring samples of the target block, the luma component samples of the target block, the chroma component samples of the target block, the availability of the neighboring samples of the target block, or the encoding parameters of the target block.
[0868] For example, offset-related information can be adjusted using a signaling / encoding / decoding mapping model based on whether the target block is in a model-based predictive merging mode.
[0869] For example, offset-related information can be adjusted using a signaling / encoding / decoding mapping model based on an indicator indicating whether a model-based predictive merging mode is to be performed in the target block.
[0870] The mapping model adjustment offset information may include at least one of a plurality of elements. Here, the plurality of elements may include: 1) an indicator indicating whether the mapping model adjustment offset is applied to at least one mapping model in the target block; 2) an indicator indicating whether the mapping model adjustment offset is applied to the mapping model of a specific component of the target block; 3) an indicator indicating whether the mapping model adjustment offset is applied to both the mapping models of the Cb component and the Cr component of the target block; 4) the size of at least one mapping model adjustment offset applied to the mapping model in the target block; 5) the sign of at least one mapping model adjustment offset applied to the mapping model in the target block; 6) at least one mapping model adjustment offset applied to the mapping model in the target block; 7) the index of at least one mapping model adjustment offset applied to the mapping model in the target block; or 8) the number of mapping model adjustment offsets applied to the mapping model in the target block.
[0871] Additionally, when multiple mapping models are used in the target block (or when the value of the indicator indicating whether to use multiple mapping models is true or 1), multiple elements may also include information for specifying the mapping model to which the mapping model adjustment offset will be applied (9).
[0872] Additionally, when a multi-mapping model is used in the target block (or when the indicator indicating whether a multi-mapping model is used is true or 1), multiple elements may also include an indicator indicating whether the mapping model adjustment offset is applied to a specific component of the Nth mapping model (N is 1, 2, ...). The specific component may refer to one of the luminance component, Cb component, and Cr component.
[0873] As an example, when using a single mapping model in a target block, an indicator (OFFSET_BOTH0_FLAG) can be used to signal / encode / decode whether the mapping model adjustment offset is applied to the first mapping model of the Cb component and the first mapping model of the Cr component.
[0874] As an example, when using a single mapping model in a target block, at least one of the following can be used: a signal transmission / encoding / decoding indicator (OFFSET_Cr0_FLAG) indicating whether the mapping model adjustment offset is applied to the Cb component, or a signal transmission / encoding / decoding indicator (OFFSET_Cb0_FLAG) indicating whether the mapping model adjustment offset is applied to the Cr component. Whether to perform signal transmission / encoding / decoding of OFFSET_Cr0_FLAG and OFFSET_Cb0_FLAG can be determined based on OFFSET_BOTH0_FLAG. For example, signal transmission / encoding / decoding of at least one of OFFSET_Cr0_FLAG or OFFSET_Cb0_FLAG can be performed only if the value of OFFSET_BOTH0_FLAG is false or 0. Similarly, whether to perform signal transmission / encoding / decoding of OFFSET_Cr0_FLAG can be determined based on OFFSET_Cb0_FLAG. For instance, signal transmission / encoding / decoding of OFFSET_Cr0_FLAG can be performed only if the value of OFFSET_Cb0_FLAG is false or 0.
[0875] As an example, if a single mapping model is used in the target block, OFFSET_BOTH0_FLAG can be signaled / encoded / decoded. If the value of OFFSET_BOTH0_FLAG is true or 1, the mapping model adjustment offset can be applied to both the first mapping model of the Cb component and the first mapping model of the Cr component. If the value of OFFSET_BOTH0_FLAG is false or 0, OFFSET_Cb0_FLAG can be signaled / encoded / decoded. If the value of OFFSET_Cb0_FLAG is true or 1, the mapping model adjustment offset can be applied to the first mapping model of the Cb component, and the mapping model adjustment offset may not be applied to the first mapping model of the Cr component. If the value of OFFSET_Cb0_FLAG is false or 0, the mapping model adjustment offset may not be applied to the first mapping model of the Cb component, but it can be applied to the first mapping model of the Cr component, and the mapping model adjustment offset may not be applied to the first mapping model of the Cb component.
[0876] As an example, if two mapping models for the Cb component and two mapping models for the Cr component are used in the target block, an indicator (OFFSET_BOTH0_FLAG) can be used to indicate whether the mapping model adjustment offset is applied to the first mapping model of the Cb component and the first mapping model of the Cr component.
[0877] As an example, if two mapping models for the Cb component and two mapping models for the Cr component are used in the target block, at least one of the following can be used: a signal transmission / encoding / decoding indicator (OFFSET_Cr0_FLAG) indicating whether the mapping model adjustment offset is applied to the first mapping model of the Cb component, or a signal transmission / encoding / decoding indicator (OFFSET_Cb0_FLAG) indicating whether the mapping model adjustment offset is applied to the first mapping model of the Cr component. Whether to perform signal transmission / encoding / decoding of OFFSET_Cr0_FLAG and OFFSET_Cb0_FLAG can be determined based on OFFSET_BOTH0_FLAG. As an example, signal transmission / encoding / decoding of at least one of OFFSET_Cr0_FLAG or OFFSET_Cb0_FLAG can only be performed if the value of OFFSET_BOTH0_FLAG is false or 0. Whether to perform signal transmission / encoding / decoding of OFFSET_Cr0_FLAG can be determined based on OFFSET_Cb0_FLAG. For example, the signal sending / encoding / decoding of OFFSET_Cr0_FLAG can only be performed when the value of OFFSET_Cb0_FLAG is false or 0.
[0878] As an example, if two mapping models for the Cb component and two mapping models for the Cr component are used in the target block, an indicator (OFFSET_BOTH1_FLAG) can be used to indicate whether the mapping model adjustment offset is applied to the second mapping model for the Cb component and the second mapping model for the Cr component.
[0879] As an example, if two mapping models for the Cb component and two mapping models for the Cr component are used in the target block, at least one of the following can be used: a signal transmission / encoding / decoding indicator (OFFSET_Cr1_FLAG) indicating whether the mapping model adjustment offset is applied to the second mapping model of the Cb component, or a signal transmission / encoding / decoding indicator (OFFSET_Cb1_FLAG) indicating whether the mapping model adjustment offset is applied to the second mapping model of the Cr component. Whether to perform signal transmission / encoding / decoding of OFFSET_Cr1_FLAG and OFFSET_Cb1_FLAG can be determined based on at least one of OFFSET_BOTH0_FLAG, OFFSET_Cr0_FLAG, OFFSET_Cb0_FLAG, and OFFSET_BOTH1_FLAG. As an example, signal transmission / encoding / decoding of at least one of OFFSET_Cr1_FLAG or OFFSET_Cb1_FLAG can only be performed if the value of OFFSET_BOTH1_FLAG is false or 0. As another example, signal transmission / encoding / decoding of at least one of OFFSET_Cr1_FLAG or OFFSET_Cb1_FLAG can only be performed if all of the values in OFFSET_BOTH0_FLAG, OFFSET_Cr0_FLAG, OFFSET_Cb0_FLAG, and OFFSET_BOTH1_FLAG are false or 0. Whether to perform signal transmission / encoding / decoding of OFFSET_Cr1_FLAG can be determined based on OFFSET_Cb1_FLAG. For example, signal transmission / encoding / decoding of OFFSET_Cr1_FLAG can only be performed if the value of OFFSET_Cb1_FLAG is false or 0.
[0880] As an example, if two mapping models for the Cb component and two mapping models for the Cr component are used in the target block, then OFFSET_BOTH0_FLAG can be signaled / encoded / decoded. If the value of OFFSET_BOTH0_FLAG is true or 1, then the mapping model adjustment offset can be applied to both the first mapping model for the Cb component and the first mapping model for the Cr component. If the value of OFFSET_BOTH0_FLAG is false or 0, then OFFSET_Cb0_FLAG can be signaled / encoded / decoded. If the value of OFFSET_Cb0_FLAG is true or 1, then the mapping model adjustment offset can be applied to the first mapping model for the Cb component, and the mapping model adjustment offset may not be applied to the first mapping model for the Cr component. If the value of OFFSET_Cb0_FLAG is false or 0, then the mapping model adjustment offset may not be applied to the first mapping model for the Cb component, and OFFSET_Cr0_FLAG can be signaled / encoded / decoded. In the same or similar method as described above, it can be determined whether to apply the mapping model adjustment offset to at least one of the second mapping model for the Cb component or the second mapping model for the Cr component of the target block.
[0881] If a mapping model adjustment offset is used in the target block, at least one of the following information can be sent / encoded / decoded to specify the size, sign, or index of the corresponding mapping model adjustment offset used in the target block: at least one mapping model adjustment offset.
[0882] For example, an index can be used to specify at least one of the mapping model adjustment offsets used in the target block from a predetermined mapping model adjustment offset list, and the value corresponding to that index in the mapping model adjustment offset list can be used as the mapping model adjustment offset.
[0883] For example, an index can be sent / encoded / decoded to specify at least one of the mapping model adjustment offsets used in the target block from a predetermined mapping model adjustment offset list, and the result of adding or multiplying (or subtracting or dividing) the predetermined value with the value corresponding to that index in the mapping model adjustment offset list can be used as the mapping model adjustment offset.
[0884] For example, the predetermined mapping model adjustment offset list may include integer values from -N to N. Alternatively, for example, the predetermined mapping model adjustment offset list may include the result of adding or multiplying (or subtracting or dividing) each integer value from -N to N with a predetermined value. Alternatively, for example, the predetermined mapping model adjustment offset list may include a value obtained by adding 1 to the result of adding or multiplying (or subtracting or dividing) each integer value from -N to N with a predetermined value.
[0885] For example, the predetermined mapping model adjustment offset list may include integer values from -N to N, but may not include 0. Alternatively, for example, the predetermined mapping model adjustment offset list may include the result of adding or multiplying (or subtracting or dividing) each integer value from -N to N with a predetermined value, but may not include 0. Alternatively, for example, the predetermined mapping model adjustment offset list may include a value obtained by adding 1 to the result of adding or multiplying (or subtracting or dividing) each integer value from -N to N with a predetermined value, but may not include 0.
[0886] For example, the predefined mapping model adjustment offset list may include At least one of them. Optionally, for example, the predefined mapping model adjustment offset list may include predefined values and The result of adding or multiplying each value in the given value, or by combining a predetermined value with... At least one of the results of subtracting or dividing each value in the list. Optionally, for example, the predefined mapping model adjustment offset list may include at least one of the following: by subtracting or dividing a predefined value from the list. The value obtained by adding or multiplying each value in the set and then adding 1 to the result, or by adding a predetermined value to the set... The value is obtained by subtracting or dividing each value in the set and adding 1 to the result.
[0887] Optionally, for example, the predefined mapping model adjustment offset list may include At least one of 0 and 0. Optionally, for example, the predefined mapping model adjustment offset list may include at least one of the following: predefined values and 0. The result of adding or multiplying each value in the array with 0, or adding a predetermined value to... The result of subtracting or dividing each value in the matrix from or by 0. Alternatively, for example, the predefined mapping model adjustment offset list may include at least one of the following: [predefined values are then compared with 0]. The value obtained by adding or multiplying each value in the set to 0 and then adding 1 to the result; the predetermined value is then compared with... The value is obtained by subtracting or dividing each value in the matrix from 0 and adding 1 to the result.
[0888] A can be 2, 3, 4, or a positive integer. N can be 1, 2, 3, 4, or a positive integer. Preset values can be 1, 2, 4, 8, 16, 32, 64, or integers.
[0889] For example, the mapping model adjustment offset is used for the first and second components of the target block, and the same mapping model adjustment offset can always be applied to the mapping model of the first component and the mapping model of the second component.
[0890] For example, the mapping model adjustment offset is used for the first and second components of the target block, and the sa...
Claims
1. An image decoding method, the method comprising: Based on the mapping model, the predicted block of the target block is obtained by performing a prediction of the target block; and The target block is reconstructed based on the predicted block. The mapping model is determined by one of a plurality of mapping model candidates.
2. The method according to claim 1, in, The multiple mapping model candidates are determined based on at least one of the target block's spatially adjacent blocks, spatially non-adjacent blocks, or temporally adjacent blocks.
3. The method according to claim 2, in, The time-adjacent block is a block within the same frame that corresponds to the position of the target block or is adjacent to the same frame.
4. The method according to claim 3, in, The block adjacent to the co-position block is the block located to the right, bottom, or lower right of the co-position block.
5. The method according to claim 3, in, The block adjacent to the co-position block is the block located at a position where the motion vector is spaced apart from the position of the co-position block, and The interval is determined based on the neighboring motion vectors of the target block.
6. The method according to claim 2, in, When the target block is a chroma component block, the spatially adjacent block is a block indicated by either the block vector of the luminance component block corresponding to the target block or the motion vector of the luminance component block.
7. The method according to claim 2, in, The spatially adjacent blocks include at least one of the blocks that are adjacent to the top, upper left, upper right, left, or lower left of the target block.
8. The method according to claim 1, in, One of the plurality of mapping model candidates is determined by an index sent from the bit stream via a signal.
9. An image encoding method, the method comprising: Based on the mapping model, the predicted block of the target block is obtained by performing a prediction of the target block; and The target block is reconstructed based on the predicted block. The mapping model is determined by one of a plurality of mapping model candidates.
10. The method according to claim 9, in, The multiple mapping model candidates are determined based on at least one of the target block's spatially adjacent blocks, spatially non-adjacent blocks, or temporally adjacent blocks.
11. The method according to claim 10, in, The time-adjacent block is a block within the same frame that corresponds to the position of the target block or is adjacent to the same frame.
12. The method according to claim 11, in, The block adjacent to the co-position block is the block located to the right, bottom, or lower right of the co-position block.
13. The method according to claim 11, in, The block adjacent to the co-position block is the block located at a position where the motion vector is spaced apart from the position of the co-position block, and The interval is determined based on the neighboring motion vectors of the target block.
14. The method according to claim 10, in, When the target block is a chroma component block, the spatially adjacent block is a block indicated by either the block vector of the luminance component block corresponding to the target block or the motion vector of the luminance component block.
15. The method of claim 10, wherein: in, The spatially adjacent blocks include at least one of the blocks that are adjacent to the top, upper left, upper right, left, or lower left of the target block.
16. The method according to claim 9, in, The index of one of the plurality of mapping model candidates is encoded into the bitstream.
17. A non-transitory computer-readable recording medium for storing a bitstream generated by an encoding method, the method comprising: Based on the mapping model, the predicted block of the target block is obtained by performing a prediction of the target block; and The target block is reconstructed based on the predicted block. The mapping model is determined by one of a plurality of mapping model candidates.