Image decoding method, image decoding apparatus, image encoding method and image encoding apparatus

The image decoding method addresses the computational and memory challenges of neural network-based in-loop filtering by using indices to identify pixel correction values, resulting in efficient image decoding and encoding processes with reduced resource requirements.

WO2025110477A1PCT designated stage expired Publication Date: 2025-05-30SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/015390
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-23
Filing Date
2024-10-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing image encoding and decoding methods that utilize neural networks for in-loop filtering face challenges related to increased computational complexity and memory requirements.

Method used

An image decoding method and device that input a decoded image, a predicted image, and coding context information into a neural network to obtain a feature map, map indices indicating pixel correction values based on the feature map, identify and apply these correction values to obtain a corrected output image, thereby reducing computational and memory demands.

Benefits of technology

The proposed method achieves reduced memory capacity and computational load while maintaining effective image quality through the use of indices to identify pixel correction values, allowing for efficient image decoding and encoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024015390_30052025_PF_FP_ABST
    Figure KR2024015390_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an image decoding method and apparatus and an image encoding method and apparatus, which are for: acquiring a feature map for a plurality of pixels by inputting, into a neural network, a decoded image, a prediction image and coding context information for the current image including the plurality of pixels; mapping an index indicating a pixel correction value for at least one pixel from among the plurality of pixels on the basis of the feature map; identifying the pixel correction value on the basis of the index; acquiring a corrected pixel value for the at least one pixel on the basis of the identified pixel correction value; and acquiring, from the current image, an output image corrected so that the at least one pixel has the corrected pixel value.
Need to check novelty before this filing date? Find Prior Art

Description

Image decoding method, image decoding device, image encoding method, and image encoding device

[0001] The present disclosure relates to image encoding and decoding. More specifically, the present disclosure relates to image encoding and decoding that perform in-loop filtering using artificial intelligence (AI), such as a neural network.

[0002] In codecs such as H.264 AVC (Advanced Video Coding), HEVC (High Efficiency Video Coding), and VVC (Versatile Video Coding), an image is divided into blocks, and each block can be predicted and decoded through inter prediction or intra prediction.

[0003] Intra prediction is a method of compressing images by removing spatial redundancy within the image, and inter prediction is a method of compressing images by removing temporal redundancy between images.

[0004] During the encoding and decoding process, the current block is reconstructed using the predicted block and the residual block of the current block. In-loop filtering is performed on the reconstructed current block to improve the image quality of the reconstructed image.

[0005] Recently, with the advancement of hardware and artificial intelligence technology, techniques utilizing neural networks have been proposed for performing in-loop filtering. However, the use of these neural networks increases computational complexity and increases the memory requirements.

[0006] Accordingly, a method is required to reduce the amount of calculation and save memory capacity while using neural networks.

[0007] An image decoding method according to one embodiment of the present disclosure may include: inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels; mapping an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map; identifying the pixel correction value based on the index; obtaining a corrected pixel value for the at least one pixel using the identified pixel correction value; and obtaining an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0008] An image decoding device according to one embodiment of the present disclosure may include a memory storing one or more instructions; and at least one processor operating according to the one or more instructions. The at least one processor may input a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels. The at least one processor may map an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map. The at least one processor may identify the pixel correction value based on the index. The at least one processor may obtain a corrected pixel value for the at least one pixel using the identified pixel correction value. The at least one processor may obtain an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0009] An image decoding method according to one embodiment of the present disclosure may include: inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels; mapping a plurality of indices indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map; identifying a plurality of pixel correction values ​​based on the plurality of indices; obtaining a final pixel correction value for the at least one pixel using the identified plurality of pixel correction values; obtaining a corrected pixel value for the at least one pixel using the obtained final pixel correction value; and obtaining an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0010] An image decoding device according to one embodiment of the present disclosure may include a memory storing one or more instructions; and at least one processor operating according to the one or more instructions. The at least one processor may input a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels. The at least one processor may map a plurality of indices indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map. The at least one processor may identify a plurality of pixel correction values ​​based on the plurality of indices. The at least one processor may obtain a final pixel correction value for the at least one pixel using the identified plurality of pixel correction values. The at least one processor may obtain a corrected pixel value for the at least one pixel using the obtained final pixel correction value. The at least one processor may obtain an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0011] An image encoding method according to one embodiment of the present disclosure may include: inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels; mapping an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map; identifying the pixel correction value based on the index; obtaining a corrected pixel value for the at least one pixel using the identified pixel correction value; and obtaining an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0012] An image encoding device according to one embodiment of the present disclosure may include a memory storing one or more instructions; and at least one processor operating according to the one or more instructions. The at least one processor may input a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels. The at least one processor may map an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map. The at least one processor may identify the pixel correction value based on the index. The at least one processor may obtain a corrected pixel value for the at least one pixel using the identified pixel correction value. The at least one processor may obtain an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0013] According to one embodiment of the present disclosure, an image encoding method may include: inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels; mapping a plurality of indices indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map; identifying a plurality of pixel correction values ​​based on the plurality of indices; obtaining a final pixel correction value for the at least one pixel using the identified plurality of pixel correction values; obtaining a corrected pixel value for the at least one pixel using the obtained final pixel correction value; and obtaining an output image corrected so that the at least one pixel from the current image has the corrected pixel value.

[0014] An image encoding device according to one embodiment of the present disclosure may include a memory storing one or more instructions; and at least one processor operating according to the one or more instructions. The at least one processor may input a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels. The at least one processor may map a plurality of indices indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map. The at least one processor may identify a plurality of pixel correction values ​​based on the plurality of indices. The at least one processor may obtain a final pixel correction value for the at least one pixel using the identified plurality of pixel correction values. The at least one processor may obtain a corrected pixel value for the at least one pixel using the obtained final pixel correction value. The at least one processor may obtain an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0015] Figure 1 is a diagram illustrating the process of encoding and decoding an image.

[0016] Figure 2 is a diagram showing blocks divided according to a tree structure from an image.

[0017] FIG. 3 illustrates a neural network-based pixel correction process according to one embodiment of the present disclosure.

[0018] FIG. 4 illustrates a process for determining a pixel correction value based on an index according to one embodiment of the present disclosure.

[0019] FIG. 5 illustrates a process for determining a pixel correction value based on a plurality of indices and weights according to one embodiment of the present disclosure.

[0020] FIG. 6 is a diagram for explaining a process of determining a pixel correction value based on a plurality of indices and weights according to one embodiment of the present disclosure.

[0021] FIG. 7 is a diagram for explaining weight information according to one example of the present disclosure.

[0022] FIG. 8 is a diagram for explaining an image decoding method according to one embodiment of the present disclosure.

[0023] FIG. 9 is a diagram for explaining an image decoding method according to one embodiment of the present disclosure.

[0024] FIG. 10 is a diagram illustrating a configuration of an image decoding device according to one embodiment of the present disclosure.

[0025] FIG. 11 is a drawing for explaining an image encoding method according to one embodiment of the present disclosure.

[0026] FIG. 12 is a drawing for explaining an image encoding method according to one embodiment of the present disclosure.

[0027] FIG. 13 is a diagram illustrating a configuration of an image encoding device according to one embodiment of the present disclosure.

[0028] FIG. 14 is a diagram for explaining a method for training neural networks and a method for obtaining a data set according to one embodiment of the present disclosure.

[0029] FIG. 15 is a diagram illustrating a method for training neural networks and a method for acquiring data sets according to one embodiment of the present disclosure.

[0030] In this disclosure, the expression “at least one of a, b or c” may refer to “a”, “b”, “c”, “a and b”, “a and c”, “b and c”, “all of a, b and c”, or variations thereof.

[0031] The present disclosure may be subject to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail herein. However, this is not intended to limit the embodiments of the present disclosure, and it should be understood that the present disclosure encompasses all modifications, equivalents, and alternatives falling within the spirit and technical scope of the various embodiments.

[0032] In describing the embodiments, detailed descriptions of related known technologies are omitted if they are deemed to unnecessarily obscure the gist of the present disclosure. Furthermore, numbers (e.g., "first," "second," etc.) used throughout the description of the specification are merely identifiers used to distinguish one component from another.

[0033] Additionally, in the present disclosure, when a component is referred to as being “connected” or “connected” to another component, it should be understood that the component may be directly connected or connected to the other component, but may also be connected or connected via another component in between, unless there is a specific description to the contrary.

[0034] In addition, components expressed as 'unit', 'module', etc. in the present disclosure may be two or more components combined into one component, or one component may be divided into two or more components with more detailed functions. In addition, each component described below may additionally perform some or all of the functions performed by other components in addition to its own main function, and of course, some of the main functions performed by each component may be exclusively performed by other components.

[0035] A processor may include various processing circuits and / or multiple processors. For example, the term “processor” as used herein, including in the claims, may include various processing circuits, including at least one processor. At least one processor, one or more processors may be configured to perform various functions described herein, individually and / or collectively, in a distributed fashion. As used herein, “processor,” “at least one processor,” and “one or more processors” may be configured to perform multiple functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions. Furthermore, the at least one processor may include a combination of processors that perform various of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.

[0036] Additionally, in the present disclosure, 'image or picture' may mean a still image (or frame), a moving image composed of a plurality of consecutive still images, or a video.

[0037] In this disclosure, "neural network" refers to a representative example of an artificial neural network model that mimics brain neurons, and is not limited to an artificial neural network model using a specific algorithm. A neural network may also be referred to as a deep neural network.

[0038] In this disclosure, a "parameter" refers to a value used in the computational process of each layer forming a neural network. For example, it can be used when applying an input value to a given computational formula. A parameter is a value set as a result of training and can be updated using separate training data as needed.

[0039] In this disclosure, a "sample" refers to data assigned to a sampling location within one-dimensional or two-dimensional data, such as an image, block, or feature data, and is the data to be processed. For example, a sample may include a pixel within a two-dimensional image. Two-dimensional data may also be referred to as a "map."

[0040] Additionally, in the present disclosure, the term "current block" refers to a block that is currently being processed. The current block may be a slice, tile, maximum coding unit, encoding unit, prediction unit, or transformation unit segmented from the current image.

[0041] Before describing an image decoding method, an image decoding device, an image encoding method, and an image encoding device according to an embodiment of the present disclosure, the image encoding and decoding process will be described with reference to FIGS. 1 and 2.

[0042] Figure 1 is a diagram illustrating the process of encoding and decoding an image.

[0043] The encoding device (110) transmits a bitstream generated through encoding of an image to the decoding device (150), and the decoding device (150) receives and decodes the bitstream to restore the image.

[0044] Specifically, in the encoding device (110), the prediction encoding unit (115) outputs a prediction block through inter prediction and intra prediction, and the transformation and quantization unit (120) transforms and quantizes residual samples of a residual block between the prediction block and the current block to output quantized transformation coefficients. The entropy encoding unit (125) encodes the quantized transformation coefficients and outputs them as a bitstream.

[0045] The quantized transform coefficients are restored into a residual block containing residual samples in the spatial domain through the inverse quantization and inverse transformation unit (130). The restored block, which is a combination of the prediction block and the residual block, is output as a filtered block through the deblocking filtering unit (135) and the loop filtering unit (140). The restored image containing the filtered block can be used as a reference image for the next input image in the prediction encoding unit (115).

[0046] The bitstream received by the decoding device (150) is restored to a residual block including residual samples in the spatial domain through the entropy decoding unit (155) and the inverse quantization and inverse transformation unit (160). The prediction block and the residual block output from the prediction decoding unit (175) are combined to generate a restored block, and the restored block is output as a filtered block through the deblocking filtering unit (165) and the loop filtering unit (170). The restored image including the filtered block can be used as a reference image for the next image in the prediction decoding unit (175).

[0047] The loop filtering unit (140) of the encoding device (110) performs loop filtering using filter information input according to user input or system settings. The filter information used by the loop filtering unit (140) is transmitted to the decoding device (150) via the entropy encoding unit (125). The loop filtering unit (170) of the decoding device (150) can perform loop filtering based on the filter information input from the entropy decoding unit (155).

[0048] In the process of encoding and decoding an image, the image is hierarchically divided, and encoding and decoding are performed on the blocks divided from the image. The blocks divided from the image are described with reference to FIG. 2.

[0049] Figure 2 is a diagram showing blocks divided according to a tree structure from an image.

[0050] A single image (200) can be divided into one or more slices or one or more tiles. A single slice can include multiple tiles.

[0051] A slice or a tile may be a sequence of one or more Maximum Coding Units (Maximum CUs).

[0052] A single maximum coding unit may be split into one or more coding units. The coding unit may be a reference block for determining a prediction mode. In other words, it may be determined whether an intra-prediction mode or an inter-prediction mode is applied to each coding unit. In the present disclosure, a maximum coding unit may be referred to as a maximum coding block, and a coding unit may be referred to as a coding block.

[0053] The size of the coding unit may be equal to or smaller than the maximum coding unit. Since the maximum coding unit is the coding unit with the maximum size, it may also be referred to as the coding unit.

[0054] One or more prediction units for intra-prediction or inter-prediction can be determined from a coding unit. The size of the prediction unit can be the same as or smaller than the coding unit.

[0055] Additionally, one or more transformation units for transformation and quantization may be determined from the coding unit. The size of the transformation unit may be the same as or smaller than the coding unit. The transformation unit serves as a reference block for transformation and quantization, and residual samples of the coding unit may be transformed and quantized for each transformation unit within the coding unit.

[0056] In the present disclosure, a current block may be a slice, tile, maximum coding unit, encoding unit, prediction unit, or transformation unit segmented from an image (200). In addition, a lower block of the current block is a block segmented from the current block. For example, if the current block is a maximum coding unit, the lower block may be a coding unit, prediction unit, or transformation unit. In addition, an upper block of the current block is a block that includes the current block as a part. For example, if the current block is a maximum coding unit, the upper block may be a picture sequence, a picture, a slice, or a tile.

[0057] Below, in FIGS. 3 to 13, a neural network-based pixel correction method used in the in-loop filtering process during the image encoding and decoding process is described.

[0058] FIG. 3 illustrates a neural network-based pixel correction process according to one embodiment of the present disclosure.

[0059] Referring to FIG. 3, a current decoded image (301), a current predicted image (302), and coding context information (303) for the current image are input to a neural network (305). The current decoded image (301), the current predicted image (302), and the coding context information (303) may be concatenated and input to the neural network. The neural network (305) may be a shallow neural network including one hidden layer, but is not limited thereto. That is, the neural network (305) may include at least one hidden layer. A feature map for a plurality of pixels included in the current image is output from the neural network (305).

[0060] Coding context information may include quantization parameters of blocks included in the current image, a partitioning tree structure of blocks included in the current image, a partitioning type of blocks included in the current image, etc.

[0061] An index for one of a plurality of pixels is obtained by applying index mapping (310) to the feature maps. Index mapping (310) may include non-linear mapping methods such as clipping, quantization, and range mapping, for example. Index mapping converts the feature map into the form of an index.

[0062] The index represents one pixel correction value among the pixel correction values ​​included in at least one predetermined data set (315) stored in the data set storage (320). The data set may be stored in the form of a lookup table. The data set is acquired using a deep neural network used in the training process of the neural network (305). The training method of the neural network (305) is described below in FIG. 14.

[0063] A scale factor (325) indicating the extent to which the pixel correction value is reflected is multiplied (330) with respect to the pixel correction value indicated by the index. The scale factor may be determined as a value greater than or equal to 0 and less than or equal to 1, but is not limited thereto. If the pixel correction value is reflected as is, the scale factor value becomes 1, and if the pixel correction value is not reflected, the scale factor value becomes 0.

[0064] A corrected pixel value for the current image is obtained by adding (335) a pixel correction value to which a scale factor is applied to a pixel value of one of multiple pixels of the current decoded image. Based on the corrected pixel value, an output image (340) in which the current image is corrected is obtained.

[0065] According to one embodiment of the present disclosure, compared to using a deep neural network including multiple layers, a shallow neural network including a single hidden layer can be used to correct pixels on a pixel-by-pixel basis, thereby reducing memory and computational complexity. In other words, while a deep neural network has a complex structure, requires multi-scale processing, requires a large amount of memory, and is also computationally complex, a shallow neural network has the advantage of a simple structure, is processed on a single scale, requires less memory, and is computationally simple.

[0066] Utilizing a neural network with a single layer is a hardware-friendly implementation. Specifically, it requires only one layer of computation, resulting in low complexity. Avoiding intermediate hidden layers (especially when using residual or dense networks) saves memory. Furthermore, pixel-level processing allows for complete parallelization.

[0067] Furthermore, a thin neural network does not necessarily contain just one layer; it may contain at least one layer. That is, a thin neural network may contain two or more layers. However, a thin neural network may have a simpler structure than a deep neural network.

[0068] FIG. 4 illustrates a process for determining a pixel correction value based on an index according to one embodiment of the present disclosure.

[0069] Referring to FIG. 4, the current decoded image (401) for the current image is a frame with a width of 4 and a height of 4, and contains 16 pixels. The current decoded image (401) is input into a neural network (305) together with the current predicted image and coding context information, and a feature map for the current image is output. By applying index mapping (310) to the feature map, indices (411) for each of the 16 pixels included in the current decoded image (401) are obtained.

[0070] A data set (421) stored in the form of a lookup table in the data set storage (320) is acquired. Specifically, if the value of the index is 1, the pixel correction value is determined as A, which is the first value of the data set (421), if the value of the index is 2, the pixel correction value is determined as B, which is the second value of the data set (421), if the value of the index is 3, the pixel correction value is determined as C, which is the third value of the data set (421), if the value of the index is 4, the pixel correction value is determined as D, which is the fourth value of the data set (421), if the value of the index is 5, the pixel correction value is determined as E, which is the fifth value of the data set (421), if the value of the index is 6, the pixel correction value is determined as F, which is the sixth value of the data set (421), and if the value of the index is 7, the pixel correction value may be determined as G, which is the seventh value of the data set (421).

[0071] Based on the indices (411) obtained through index mapping (310) and the data set (421) obtained from the data set storage (320), pixel correction values ​​(412) indicated by the index of each of the 16 pixels are obtained.

[0072] Each of the pixel correction values ​​(412) is multiplied by a scale factor (330) and then added to each of a plurality of pixels corresponding to each of the pixel correction values ​​(412) of the current decoded image (401), thereby obtaining an output image in which the current image is corrected.

[0073] In Fig. 4, the data set (421) is described in the form of a one-dimensional array lookup table, but is not limited thereto, and may be in the form of a multi-dimensional, for example, two-dimensional, three-dimensional, four-dimensional, or five-dimensional array lookup table. In such a case, the index of the data set may also be expressed as multi-dimensional values ​​rather than one-dimensional values. For example, in the case of two dimensions, the index may be expressed as (i1, i2), and in the case of three dimensions, the index may be expressed as (i1, i2, i3).

[0074] FIG. 5 illustrates a process for determining a pixel correction value based on a plurality of indices and weights according to one embodiment of the present disclosure.

[0075] Referring to FIG. 5, the current decoded image (501), the current predicted image (502), and the coding context information (503) for the current image are input into a neural network (505), and a feature map and multiple pieces of weight information (513, 514, 515) for the current image are obtained. The neural network (505) may be a shallow neural network including one hidden layer. However, the present invention is not limited thereto. That is, the neural network (505) may include at least one hidden layer. The training method of the neural network (505) is described below in FIG. 15.

[0076] The feature map for the current image is applied to each of the index mapping methods (510, 511, 512). The feature map can have a size of HxWxC, where H represents height, W represents width, and C represents the number of channels.

[0077] In addition, the feature map for the current image may be divided into the number of index mapping methods and applied to each of the index mapping methods (510, 511, 512). In addition, the feature map may be divided and applied unequally to each of the index mapping methods. Each feature map may have a size of HxWxC1, HxWxC2, HxWxC3. H represents height, W represents width, C1 represents the number of channels to which the first index mapping method (510) is applied, C2 represents the number of channels to which the second index mapping method (511) is applied, and C3 represents the number of channels to which the third index mapping method (512) is applied.

[0078] That is, the feature map for the current image may be feature maps for each of the three index mapping methods (510, 511, 512). Alternatively, the feature map for the current image may be divided into three parts and applied to each of the three index mapping methods (510, 511, 512).

[0079] The weight information (513, 514, 515) may be weight information corresponding to each index mapping method (510, 511, 512). In addition, the weight information (513, 514, 515) may be weight information corresponding to a feature map that is input and divided into each index mapping method (510, 511, 512).

[0080] Since the weight information (513, 514, 515) includes weights for each of the plurality of pixels included in the current decoded image (501), the size of the weight information (513, 514, 515) may be HxWx1, respectively. H represents height, and W represents width.

[0081] One of the first indices obtained through the first index mapping (510) represents one pixel correction value among the pixel correction values ​​included in at least one first data set (516) predetermined from the first data set storage (520). The first data set (516) may be stored in the form of a lookup table.

[0082] The first weight information (513) corresponding to the first index mapping (510) is multiplied (526) with respect to the pixel correction value indicated by the index.

[0083] One of the second indices obtained through the second index mapping (511) represents one pixel correction value among the pixel correction values ​​included in at least one second data set (517) predetermined from the second data set storage (521). The second data set (517) may be stored in the form of a lookup table.

[0084] The second weight information (514) corresponding to the second index mapping (511) is multiplied (527) with respect to the pixel correction value indicated by the index.

[0085] One of the third indices obtained through the third index mapping (512) represents one pixel correction value among the pixel correction values ​​included in at least one third data set (518) predetermined from the third data set storage (522). The third data set (518) may be stored in the form of a lookup table.

[0086] The third weight information (515) corresponding to the third index mapping (512) is multiplied (528) with respect to the pixel correction value indicated by the index.

[0087] The pixel correction values ​​multiplied by the first weight information (513), the pixel correction values ​​multiplied by the second weight information (514), and the pixel correction values ​​multiplied by the third weight information (515) are added (530). The final pixel correction values ​​are obtained by multiplying the added pixel correction values ​​by a scale factor (525). The scale factor (525) indicates the extent to which the pixel correction values ​​are reflected, and may be determined as a value greater than or equal to 0 and less than or equal to 1, but is not limited thereto. If the pixel correction values ​​are reflected as is, the scale factor value becomes 1, and if the pixel correction values ​​are not reflected, the scale factor value becomes 0.

[0088] The final pixel correction values ​​and the pixel values ​​of the current decoded image (501) are added (540) to obtain an output image (545) in which the current image is corrected.

[0089] According to one embodiment of the present disclosure, multiple data sets are used to generate multiple pixel compensation value candidates, and the multiple pixel compensation value candidates are linearly weighted to generate a single pixel compensation value, thereby dynamically and adaptively applying pixel compensation to a pixel. By blending multiple smaller data sets, the same coding efficiency as using a larger data set is achieved, while saving memory.

[0090] FIG. 6 is a diagram for explaining a process of determining a pixel correction value based on a plurality of indices and weights according to one embodiment of the present disclosure.

[0091] Referring to Fig. 6, the current decoded image (501) for the current image is a frame with a width of 4 and a height of 4, and includes 16 pixels (600). The current decoded image (501) is input into a neural network (505) together with the current predicted image and coding context information, and a feature map for the current image and weight information for each index mapping method are output.

[0092] By applying the first index mapping (510) to the feature map, first indices (610) are obtained for each of the 16 pixels (600) included in the current decoded image (501).

[0093] A first data set (516) is acquired in a first data set storage (520). The first data set (516) is stored in the form of a lookup table (640). Specifically, if the value of the first index is 1, the pixel compensation value is determined as A1, which is the first value of the lookup table (640), if the value of the first index is 2, the pixel compensation value is determined as B1, which is the second value of the lookup table (640), if the value of the first index is 3, the pixel compensation value is determined as C1, which is the third value of the lookup table (640), if the value of the first index is 4, the pixel compensation value is determined as D1, which is the fourth value of the lookup table (640), if the value of the first index is 5, the pixel compensation value is determined as E1, which is the fifth value of the lookup table (640), if the value of the first index is 6, the pixel compensation value is determined as F1, which is the sixth value of the lookup table (640), if the value of the first index is 7, the pixel compensation value is determined as G1, which is the seventh value of the lookup table (640), and if the value of the first index is 8, the pixel compensation value It can be determined as H1, the 8th value of the lookup table (640).

[0094] According to the first indexes (610) and the lookup table (640), first pixel correction values ​​(670) corresponding to a plurality of pixels (600) of the current decoded image (501) are obtained.

[0095] Additionally, by applying a second index mapping (511) to the feature map, second indices (620) are obtained for each of the 16 pixels (600) included in the current decoded image (501).

[0096] A second data set (517) is acquired in a second data set storage (521). The second data set (517) is stored in the form of a lookup table (650). Specifically, if the value of the second index is 1, the pixel compensation value is determined as A2, which is the first value of the lookup table (650), if the value of the second index is 2, the pixel compensation value is determined as B2, which is the second value of the lookup table (650), if the value of the second index is 3, the pixel compensation value is determined as C2, which is the third value of the lookup table (650), if the value of the second index is 4, the pixel compensation value is determined as D2, which is the fourth value of the lookup table (650), if the value of the second index is 5, the pixel compensation value is determined as E2, which is the fifth value of the lookup table (650), if the value of the second index is 6, the pixel compensation value is determined as F2, which is the sixth value of the lookup table (650), if the value of the second index is 7, the pixel compensation value is determined as G2, which is the seventh value of the lookup table (650), and if the value of the second index is 8, the pixel compensation value It can be determined by H2, the 8th value of the lookup table (650).

[0097] According to the second indexes (620) and the lookup table (650), second pixel correction values ​​(680) corresponding to a plurality of pixels (600) of the current decoded image (501) are obtained.

[0098] Additionally, third index mapping (512) is applied to the feature map to obtain third indices (630) for each of the 16 pixels (600) included in the current decoded image (501).

[0099] A third data set (518) is acquired in a third data set storage (522). The third data set (815) is stored in the form of a lookup table (660). Specifically, if the value of the third index is 1, the pixel compensation value is determined as A3, which is the first value of the lookup table (660), if the value of the third index is 2, the pixel compensation value is determined as B3, which is the second value of the lookup table (660), if the value of the third index is 3, the pixel compensation value is determined as C3, which is the third value of the lookup table (660), if the value of the third index is 4, the pixel compensation value is determined as D3, which is the fourth value of the lookup table (660), if the value of the third index is 5, the pixel compensation value is determined as E3, which is the fifth value of the lookup table (660), if the value of the third index is 6, the pixel compensation value is determined as F3, which is the sixth value of the lookup table (660), if the value of the third index is 7, the pixel compensation value is determined as G3, which is the seventh value of the lookup table (660), and if the value of the third index is 8, the pixel compensation value It can be determined by H3, the 8th value of the lookup table (660).

[0100] According to the third indexes (630) and the lookup table (660), third pixel correction values ​​(690) corresponding to a plurality of pixels (600) of the current decoded image (501) are obtained.

[0101] The first pixel correction values ​​(670) are multiplied by the first weight information (513) (526), ​​the second pixel correction values ​​(680) are multiplied by the second weight information (514) (527), and the third pixel correction values ​​(690) are multiplied by the third weight information (515) (528), and then all are added (530).

[0102] The final pixel correction values ​​are obtained by multiplying the added pixel correction values ​​by a scale factor (525) (535).

[0103] An output image (545) in which the current image is corrected is obtained by adding (540) each of the final pixel correction values ​​and each of the pixel values ​​(600) of the current decoded image (501).

[0104] As illustrated in FIG. 6, the weights for each of the plurality of pixels of the current decoded image (501) of the first weight information (513) are all 0.1, the weights for each of the plurality of pixels of the current decoded image (501) of the second weight information (514) are all 0.2, and the weights for each of the plurality of pixels of the current decoded image (501) of the third weight information (515) are all 0.7, but are not limited thereto. An example of the weight information is described below in FIG. 7.

[0105] FIG. 7 is a diagram for explaining weight information according to one example of the present disclosure.

[0106] Referring to FIG. 7, the first weight information (513) of FIG. 6 may be first weight values ​​(710). That is, the first weight information (513) may not have all weight values ​​of 0.1, but may include various values. In addition, the second weight information (514) of FIG. 6 may be second weight values ​​(720). That is, the second weight information (514) may not have all weight values ​​of 0.2, but may include various values. In addition, the third weight information (515) of FIG. 6 may be third weight values ​​(730). That is, the third weight information (515) may not have all weight values ​​of 0.7, but may include various values.

[0107] According to one embodiment of the present disclosure, the sum of the weight values ​​corresponding to each pixel included in the weight information (513, 514, 515) may be 1. However, this is not limited thereto. That is, the sum of the weight values ​​corresponding to each pixel included in the weight information (513, 514, 515) may be less than 1 or greater than 1.

[0108] The method for correcting pixel values ​​according to an embodiment of the present disclosure described above with reference to FIGS. 3 to 7 can be used as an in-loop filter. For example, it can be arranged serially between in-loop filters. Specifically, the method according to an embodiment of the present disclosure can be performed before deblocking filtering, after deblocking filtering, after Sample Adaptive Offset (SAO) filtering, or after Adaptive Loop Filtering (ALF) filtering.

[0109] Additionally, the method according to one embodiment of the present disclosure may be performed before deblocking filtering, after deblocking filtering, after Constrained Directional Enhanced Filtering (CDEF) filtering, after Cross-Component Sample Offset (CCSO) filtering, or after Loop Restoration (LR) filtering.

[0110] Additionally, the method for correcting a pixel value according to one embodiment of the present disclosure may be arranged in parallel with at least one of the in-loop filters. Specifically, the method may be arranged in parallel with one of deblocking filtering, SAO filtering, and ALF filtering, or may be arranged in parallel with the deblocking filtering and the SAO filtering, or may be arranged in parallel with the SAO filtering and the ALF filtering, or may be arranged in parallel with the deblocking filtering, the SAO filtering, and the ALF filtering.

[0111] In addition, the method for correcting a pixel value according to an embodiment of the present disclosure may be arranged in parallel with at least one in-loop filter among the in-loop filters. Specifically, it may be arranged in parallel with one of deblocking filtering, CDEF filtering, CCSO filtering, and LR filtering, or arranged in parallel with deblocking filtering and CDEF filtering, or arranged in parallel with deblocking filtering, CDEF filtering, and CCSO filtering, or arranged in parallel with deblocking filtering, CDEF filtering, and CCSO filtering, or arranged in parallel with CDEF filtering and CCSO filtering, or arranged in parallel with CDEF filtering, CCSO filtering, and LR filtering, or arranged in parallel with CCSO filtering and LR filtering.

[0112] Additionally, the method for correcting pixel values ​​according to one embodiment of the present disclosure can be used as an out-loop filter, i.e., a post-filter. For example, it can be used as an enhancement filter after decoding.

[0113] The output of the neural network used in the method for correcting pixel values ​​according to one embodiment of the present disclosure is used for an index for determining a pixel correction value or pixel offset for improving the quality of video coding.

[0114] FIG. 8 is a diagram for explaining an image decoding method according to one embodiment of the present disclosure.

[0115] In step S810, the image decoding device (1000) inputs a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels.

[0116] According to one embodiment of the present disclosure, the neural network may include at least one hidden layer.

[0117] In step S830, the image decoding device (1000) maps an index indicating a pixel correction value for at least one pixel among a plurality of pixels based on the feature map.

[0118] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in a predetermined data set.

[0119] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in one of a plurality of predetermined data sets.

[0120] According to one embodiment of the present disclosure, the predetermined data set or a plurality of predetermined data sets may be stored in the form of a lookup table.

[0121] At step S850, the image decoding device (1000) identifies a pixel correction value based on the index.

[0122] In step S870, the image decoding device (1000) obtains a corrected pixel value for at least one pixel using the identified pixel correction value.

[0123] According to one embodiment of the present disclosure, an image decoding device (1000) may obtain a scale factor indicating a scale of pixel correction. The image decoding device (1000) may multiply a pixel correction value by the scale factor. That is, the degree to which the pixel correction value is applied may be modified by the scale factor. A modified pixel value for at least one pixel may be determined based on the pixel correction value and the scale factor. The scale factor may be determined and signaled by Rate-Distortion Optimization (RDO) during an encoding process. The scale factor may be determined based on information about the signaled scale factor during a decoding process.

[0124] At step S890, the image decoding device (1000) obtains an output image corrected so that at least one pixel from the current image has a corrected pixel value.

[0125] According to one embodiment of the present disclosure, a neural network can be trained by inputting a training decoded image, a training predicted image, and training coding context information for a training current image, outputting a training feature map, mapping a training index according to a predetermined index mapping method using the training feature map, inputting the training index into a second neural network including a plurality of hidden layers, obtaining a training pixel correction value, and obtaining a training output image based on the training pixel correction value.

[0126] According to one embodiment of the present disclosure, at least one data set including indices can be obtained by inputting all possible indices into a second neural network including a plurality of hidden layers.

[0127] FIG. 9 is a diagram for explaining an image decoding method according to one embodiment of the present disclosure.

[0128] In step S910, the image decoding device (1000) inputs a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels.

[0129] According to one embodiment of the present disclosure, the neural network may include at least one hidden layer.

[0130] In step S920, the image decoding device (1000) maps a plurality of indices indicating a pixel correction value for at least one pixel among a plurality of pixels based on the feature map.

[0131] According to one embodiment of the present disclosure, one of the plurality of indices may represent one of the pixel correction values ​​included in a predetermined data set.

[0132] According to one embodiment of the present disclosure, one of the plurality of indices may represent one of the pixel correction values ​​included in one of the plurality of predetermined data sets.

[0133] According to one embodiment of the present disclosure, the predetermined data set or a plurality of predetermined data sets may be stored in the form of a lookup table.

[0134] In step S930, the image decoding device (1000) identifies a plurality of pixel correction values ​​based on a plurality of indices.

[0135] In step S940, the image decoding device (1000) obtains a final pixel correction value for at least one pixel using the identified plurality of pixel correction values.

[0136] According to one embodiment of the present disclosure, an image decoding device (1000) may obtain a scale factor indicating a scale of pixel correction. The image decoding device (1000) may multiply a final pixel correction value by the scale factor. That is, the degree to which the final pixel correction value is applied may be modified by the scale factor. A modified pixel value for at least one pixel may be determined based on the final pixel correction value and the scale factor. The scale factor may be determined and signaled by Rate-Distortion Optimization (RDO) during an encoding process. The scale factor may be determined based on information about the signaled scale factor during a decoding process.

[0137] According to one embodiment of the present disclosure, the final pixel correction value can be obtained using a weighted sum of a plurality of pixel correction values.

[0138] According to one embodiment of the present disclosure, the final pixel correction value can be obtained using a simple sum of multiple pixel correction values.

[0139] In step S950, the image decoding device (1000) can obtain a corrected pixel value for at least one pixel using the obtained final pixel correction value.

[0140] At step S960, the image decoding device (1000) obtains an output image corrected so that at least one pixel from the current image has a corrected pixel value.

[0141] According to one embodiment of the present disclosure, a neural network can be trained by inputting a training decoded image, a training predicted image, and training coding context information for a training current image, outputting a training feature map, mapping each of a plurality of training indices according to a plurality of predetermined index mapping methods using the training feature map, inputting each of the plurality of training indices into each of a plurality of neural networks including a plurality of hidden layers, thereby obtaining a plurality of training pixel correction values, obtaining a training final pixel correction value using the plurality of training pixel correction values, and obtaining a training output image based on the training final pixel correction value.

[0142] According to one embodiment of the present disclosure, at least one data set including each of a plurality of indices can be obtained by inputting all possible indices into each of a plurality of neural networks including a plurality of hidden layers.

[0143] FIG. 10 is a diagram illustrating a configuration of an image decoding device according to one embodiment of the present disclosure.

[0144] Referring to FIG. 10, an image decoding device (1000) according to an embodiment of the present disclosure may include a feature map acquisition unit (1010), an index mapping unit (1020), a data set storage unit (1030), a pixel correction value identification unit (1040), and a correction pixel acquisition unit (1050). In addition, an image decoding device (1000) according to an embodiment of the present disclosure may further include a scale factor acquisition unit (1060).

[0145] The feature map acquisition unit (1010), the index mapping unit (1020), the data set storage unit (1030), the pixel correction value identification unit (1040), the corrected pixel acquisition unit (1050), and the scale factor acquisition unit (1060) may be implemented as a processor. The feature map acquisition unit (1010), the index mapping unit (1020), the data set storage unit (1030), the pixel correction value identification unit (1040), the corrected pixel acquisition unit (1050), and the scale factor acquisition unit (1060) may operate according to instructions stored in a memory.

[0146] FIG. 10 illustrates the feature map acquisition unit (1010), the index mapping unit (1020), the data set storage unit (1030), the pixel correction value identification unit (1040), the corrected pixel acquisition unit (1050), and the scale factor acquisition unit (1060) individually, but the feature map acquisition unit (1010), the index mapping unit (1020), the data set storage unit (1030), the pixel correction value identification unit (1040), the corrected pixel acquisition unit (1050), and the scale factor acquisition unit (1060) may be implemented through one processor. In this case, the feature map acquisition unit (1010), the index mapping unit (1020), the data set storage unit (1030), the pixel correction value identification unit (1040), the correction pixel acquisition unit (1050), and the scale factor acquisition unit (1060) may be implemented as a dedicated processor, or may be implemented through a combination of a general-purpose processor such as an application processor (AP), a central processing unit (CPU), or a graphic processing unit (GPU) and software. In addition, in the case of a dedicated processor, it may include a memory for implementing an embodiment of the present disclosure, or a memory processing unit for utilizing an external memory.

[0147] The feature map acquisition unit (1010), the index mapping unit (1020), the data set storage unit (1030), the pixel correction value identification unit (1040), the corrected pixel acquisition unit (1050), and the scale factor acquisition unit (1060) may be implemented by a plurality of processors. In this case, the feature map acquisition unit (1010), the index mapping unit (1020), the data set storage unit (1030), the pixel correction value identification unit (1040), the corrected pixel acquisition unit (1050), and the scale factor acquisition unit (1060) may be implemented by a combination of dedicated processors, or may be implemented by a combination of software and a plurality of general-purpose processors such as an AP, a CPU, or a GPU. In addition, the processor may include an artificial intelligence-dedicated processor. As another example, the artificial intelligence-dedicated processor may be configured as a separate chip from the processor.

[0148] The feature map acquisition unit (1010) inputs the current decoded image, the current predicted image, and coding context information for the current image into a neural network to acquire a feature map for a plurality of pixels included in the current image.

[0149] The index mapping unit (1020) maps indices for each of a plurality of pixels included in the current image based on the feature map.

[0150] The data set storage unit (1030) transfers a data set including pixel correction values ​​indicated by the index mapped in the index mapping unit (1020) to the pixel correction value identification unit (1040).

[0151] The pixel correction value identification unit (1040) identifies the pixel correction value based on the index mapped in the index mapping unit (1020) and the data set transferred from the data set storage unit (1030).

[0152] The correction pixel acquisition unit (1050) obtains an output image in which the current image is corrected by using the pixel correction value identified from the pixel correction value identification unit (1040) and the pixel value of the current decoded image.

[0153] According to one embodiment of the present disclosure, the scale factor acquisition unit (1060) can acquire a scale factor from a bitstream. The correction pixel acquisition unit (1050) can acquire an output image in which the current image is corrected using the pixel correction value identified by the pixel correction value identification unit (1040), the pixel value of the current decoded image, and the scale factor.

[0154] According to one embodiment of the present disclosure, when multiple index mapping methods are applied, the feature map acquisition unit (1010) can additionally acquire weight information from a neural network. The index mapping unit (1020) can map multiple indices for each of multiple pixels included in a current image by applying multiple index mapping methods based on the feature map. The data set storage unit (1030) can transfer multiple data sets including pixel correction values ​​indicated by the multiple indices mapped by the index mapping unit (1020) to the pixel correction value identification unit (1040). The pixel correction value identification unit (1040) can identify the pixel correction value based on the multiple indices mapped by the index mapping unit (1020) and the multiple data sets transferred from the data set storage unit (1030). The corrected pixel acquisition unit (1050) can acquire an output image in which the current image is corrected by using the pixel correction value identified by the pixel correction value identification unit (1040), the pixel value of the current decoded image, and the weight information. Additionally, the scale factor acquisition unit (1060) can acquire a scale factor from a bitstream. The correction pixel acquisition unit (1050) can acquire an output image in which the current image is corrected using the pixel correction value identified from the pixel correction value identification unit (1040), the pixel value of the current decoded image, weight information, and the scale factor.

[0155] FIG. 11 is a drawing for explaining an image encoding method according to one embodiment of the present disclosure.

[0156] In step S1110, the image encoding device (1300) inputs a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels.

[0157] According to one embodiment of the present disclosure, the neural network may include at least one hidden layer.

[0158] In step S1130, the image encoding device (1300) maps an index indicating a pixel correction value for at least one pixel among a plurality of pixels based on the feature map.

[0159] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in a predetermined data set.

[0160] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in one of a plurality of predetermined data sets.

[0161] According to one embodiment of the present disclosure, the predetermined data set or a plurality of predetermined data sets may be stored in the form of a lookup table.

[0162] At step S1150, the image encoding device (1300) identifies a pixel correction value based on the index.

[0163] In step S1170, the image encoding device (1300) obtains a corrected pixel value for at least one pixel using the identified pixel correction value.

[0164] According to one embodiment of the present disclosure, the video encoding device (1300) can generate a scale factor indicating a scale of pixel correction. The video encoding device (1300) can multiply a pixel correction value by the scale factor. That is, the degree to which the pixel correction value is applied can be modified by the scale factor. A modified pixel value for at least one pixel can be determined based on the pixel correction value and the scale factor. The scale factor can be generated and signaled by Rate-Distortion Optimization (RDO) during the encoding process.

[0165] In step S1190, the image encoding device (1300) obtains an output image corrected so that at least one pixel from the current image has a corrected pixel value.

[0166] According to one embodiment of the present disclosure, a neural network can be trained by inputting a training decoded image, a training predicted image, and training coding context information for a training current image, outputting a training feature map, mapping a training index according to a predetermined index mapping method using the training feature map, inputting the training index into a second neural network including a plurality of hidden layers, obtaining a training pixel correction value, and obtaining a training output image based on the training pixel correction value.

[0167] According to one embodiment of the present disclosure, at least one data set including indices can be obtained by inputting all possible indices into a second neural network including a plurality of hidden layers.

[0168] FIG. 12 is a drawing for explaining an image encoding method according to one embodiment of the present disclosure.

[0169] In step S1210, the image encoding device (1300) inputs a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels.

[0170] According to one embodiment of the present disclosure, the neural network may include at least one hidden layer.

[0171] In step S1220, the image encoding device (1300) maps a plurality of indices indicating a pixel correction value for at least one pixel among a plurality of pixels based on the feature map.

[0172] According to one embodiment of the present disclosure, one of the plurality of indices may represent one of the pixel correction values ​​included in a predetermined data set.

[0173] According to one embodiment of the present disclosure, one of the plurality of indices may represent one of the pixel correction values ​​included in one of the plurality of predetermined data sets.

[0174] According to one embodiment of the present disclosure, the predetermined data set or a plurality of predetermined data sets may be stored in the form of a lookup table.

[0175] In step S1230, the image encoding device (1300) identifies a plurality of pixel correction values ​​based on a plurality of indices.

[0176] In step S1240, the image encoding device (1300) obtains a final pixel correction value for at least one pixel using the identified plurality of pixel correction values.

[0177] According to one embodiment of the present disclosure, the image encoding device (1300) can obtain a scale factor indicating a scale of pixel correction. The image encoding device (1300) can multiply a final pixel correction value by the scale factor. That is, the degree to which the final pixel correction value is applied can be modified by the scale factor. A modified pixel value for at least one pixel can be determined based on the final pixel correction value and the scale factor. The scale factor can be generated and signaled by Rate-Distortion Optimization (RDO) during the encoding process.

[0178] According to one embodiment of the present disclosure, the final pixel correction value can be obtained using a weighted sum of a plurality of pixel correction values.

[0179] According to one embodiment of the present disclosure, the final pixel correction value can be obtained using a simple sum of multiple pixel correction values.

[0180] At step S1250, the image encoding device (1300) can obtain a corrected pixel value for at least one pixel using the obtained final pixel correction value.

[0181] At step S1260, the image encoding device (1300) obtains an output image corrected so that at least one pixel from the current image has a corrected pixel value.

[0182] According to one embodiment of the present disclosure, a neural network can be trained by inputting a training decoded image, a training predicted image, and training coding context information for a training current image, outputting a training feature map, mapping each of a plurality of training indices according to a plurality of predetermined index mapping methods using the training feature map, inputting each of the plurality of training indices into each of a plurality of neural networks including a plurality of hidden layers, thereby obtaining a plurality of training pixel correction values, obtaining a training final pixel correction value using the plurality of training pixel correction values, and obtaining a training output image based on the training final pixel correction value.

[0183] According to one embodiment of the present disclosure, at least one data set including each of a plurality of indices can be obtained by inputting all possible indices into each of a plurality of neural networks including a plurality of hidden layers.

[0184] FIG. 13 is a diagram illustrating a configuration of an image encoding device according to one embodiment of the present disclosure.

[0185] Referring to FIG. 13, an image encoding device (1300) according to an embodiment of the present disclosure may include a feature map acquisition unit (1310), an index mapping unit (1320), a data set storage unit (1330), a pixel correction value identification unit (1340), and a correction pixel acquisition unit (1350). In addition, an image encoding device (1300) according to an embodiment of the present disclosure may further include a scale factor generation unit (1360).

[0186] The feature map acquisition unit (1310), the index mapping unit (1320), the data set storage unit (1330), the pixel correction value identification unit (1340), the corrected pixel acquisition unit (1350), and the scale factor generation unit (1360) may be implemented as a processor. The feature map acquisition unit (1310), the index mapping unit (1320), the data set storage unit (1330), the pixel correction value identification unit (1340), the corrected pixel acquisition unit (1350), and the scale factor generation unit (1360) may operate according to instructions stored in a memory.

[0187] FIG. 13 illustrates the feature map acquisition unit (1310), the index mapping unit (1320), the data set storage unit (1330), the pixel correction value identification unit (1340), the correction pixel acquisition unit (1350), and the scale factor generation unit (1360) individually, but the feature map acquisition unit (1310), the index mapping unit (1320), the data set storage unit (1330), the pixel correction value identification unit (1340), the correction pixel acquisition unit (1350), and the scale factor generation unit (1360) may be implemented through one processor. In this case, the feature map acquisition unit (1310), the index mapping unit (1320), the data set storage unit (1330), the pixel correction value identification unit (1340), the correction pixel acquisition unit (1350), and the scale factor generation unit (1360) may be implemented as a dedicated processor, or may be implemented through a combination of a general-purpose processor such as an application processor (AP), a central processing unit (CPU), or a graphic processing unit (GPU) and software. In addition, in the case of a dedicated processor, it may include a memory for implementing an embodiment of the present disclosure, or a memory processing unit for utilizing an external memory.

[0188] The feature map acquisition unit (1310), the index mapping unit (1320), the data set storage unit (1330), the pixel correction value identification unit (1340), the correction pixel acquisition unit (1350), and the scale factor generation unit (1360) may be implemented by a plurality of processors. In this case, the feature map acquisition unit (1310), the index mapping unit (1320), the data set storage unit (1330), the pixel correction value identification unit (1340), the correction pixel acquisition unit (1350), and the scale factor generation unit (1360) may be implemented by a combination of dedicated processors, or may be implemented by a combination of software and a plurality of general-purpose processors such as an AP, a CPU, or a GPU. In addition, the processor may include an artificial intelligence-dedicated processor. As another example, the artificial intelligence-dedicated processor may be configured as a separate chip from the processor.

[0189] The feature map acquisition unit (1310) inputs the current decoded image, the current predicted image, and coding context information for the current image into a neural network to acquire a feature map for multiple pixels included in the current image.

[0190] The index mapping unit (1320) maps indices for each of a plurality of pixels included in the current image based on the feature map.

[0191] The data set storage unit (1330) transfers a data set including pixel correction values ​​indicated by the index mapped in the index mapping unit (1320) to the pixel correction value identification unit (1340).

[0192] The pixel correction value identification unit (1340) identifies the pixel correction value based on the index mapped in the index mapping unit (1320) and the data set transferred from the data set storage unit (1330).

[0193] The correction pixel acquisition unit (1350) obtains an output image in which the current image is corrected by using the pixel correction value identified from the pixel correction value identification unit (1340) and the pixel value of the current decoded image.

[0194] According to one embodiment of the present disclosure, the scale factor generation unit (1360) can generate a scale factor. The correction pixel acquisition unit (1350) can obtain an output image in which the current image is corrected using the pixel correction value identified from the pixel correction value identification unit (1340), the pixel value of the current decoded image, and the scale factor.

[0195] According to one embodiment of the present disclosure, when multiple index mapping methods are applied, the feature map acquisition unit (1310) can additionally acquire weight information from a neural network. The index mapping unit (1320) can map multiple indices for each of multiple pixels included in a current image by applying multiple index mapping methods based on the feature map. The data set storage unit (1330) can transfer multiple data sets including pixel correction values ​​indicated by the multiple indices mapped by the index mapping unit (1320) to the pixel correction value identification unit (1340). The pixel correction value identification unit (1340) can identify the pixel correction value based on the multiple indices mapped by the index mapping unit (1320) and the multiple data sets transferred from the data set storage unit (1330). The corrected pixel acquisition unit (1350) can acquire an output image in which the current image is corrected by using the pixel correction value identified by the pixel correction value identification unit (1340), the pixel value of the current decoded image, and the weight information. In addition, the scale factor generation unit (1360) can generate a scale factor. The correction pixel acquisition unit (1350) can obtain an output image in which the current image is corrected by using the pixel correction value identified from the pixel correction value identification unit (1340), the pixel value of the current decoded image, weight information, and the scale factor.

[0196] FIG. 14 is a diagram for explaining a method for training neural networks and a method for obtaining a data set according to one embodiment of the present disclosure.

[0197] Referring to FIG. 14, the current image for training (1400) corresponds to the current image described above, the current decoded image for training (1401) corresponds to the current decoded image (301) described above, the current predicted image for training (1402) corresponds to the current predicted image (302) described above, the training coding context information (1403) for the current image for training (1400) corresponds to the coding context information (303) for the current image described above, and the training output image (1430) corresponds to the output image (340) described above.

[0198] In a training process according to one embodiment of the present disclosure, neural networks can be trained so that the training output image (1430) is as similar as possible to the training current image (1400). To this end, as illustrated in FIG. 14, loss information (1445) can be used to train the neural networks.

[0199] Specifically, the training process of the neural networks will be described. The training current decoded image (1401), the training current predicted image (1402), and the training coding context information (1403) for the training current image (1400) are input to the first neural network (1405), and a feature map for the training current image (1400) is output. The first neural network (1405) is a neural network including at least one hidden layer.

[0200] Index mapping (1410) is performed using a feature map to obtain indices representing a pixel correction value for one of a plurality of pixels included in a current training image (1400). The indices are input to a second neural network (1415) to obtain pixel correction values. The second neural network (1415) is a deep neural network including a plurality of hidden layers. The index mapping (1410) used in the training process of the neural networks may be an approximate index mapping in which an index including arbitrary noise is obtained. The index mapping (1410) method may include a non-linear mapping such as clipping, stochastic quantization, or noisy quantization. A deep neural network including a plurality of layers is used to map the indices to actual residual pixel values. The receptive field of the deep neural network must have a size of 1x1. This is to well reflect local features.

[0201] A training output image (1430) is obtained by adding (1420) the acquired pixel correction values ​​and the current decoded training image (1401). Loss information (1440) is obtained by comparing (1435) the training current image (1400) and the training output image (1430).

[0202] The loss information (1440) may correspond to a difference between a current training image (1400) and a training output image (1430). In one embodiment of the present disclosure, the difference between the current training image (1400) and the training output image (1430) may include at least one of an L1-norm value, an L2-norm value, a Structural Similarity (SSIM) value, a Peak Signal-To-Noise Ratio-Human Vision System (PSNR-HVS) value, a Multiscale SSIM (MS-SSIM) value, a Variance Inflation Factor (VIF) value, or a Video Multimethod Assessment Fusion (VMAF) value between the current training image (1400) and the training output image (1430).

[0203] Since the loss information (1440) is related to the quality of the training output image (1430), it may also be referred to as quality loss information.

[0204] The first neural network (1405) and the second neural network (1415) can be trained so that the final loss information derived from the loss information (1440) is reduced or minimized.

[0205] In one embodiment of the present disclosure, the first neural network (1405) and the second neural network (1415) can reduce or minimize the final loss information by changing the values ​​of preset parameters.

[0206] In one embodiment of the present disclosure, the final loss information can be calculated according to the following mathematical expression 1.

[0207] [Mathematical Formula 1]

[0208] Final loss information = a*loss information

[0209] In mathematical expression 1, a is a weight applied to loss information (1440).

[0210] According to mathematical expression 1, it can be seen that the first neural network (1405) and the second neural network (1415) are trained in a direction in which the training output image (1430) becomes as similar as possible to the training current image (1400).

[0211] The training process described with reference to FIG. 14 can be performed by a training device. The training device can be, for example, an image encoding device (1300) or a separate server.

[0212] The trained second neural network (1415) is used to obtain a data set stored in the data set storage (1445). Specifically, all possible combinations of indices (1450) are input into the trained second neural network (1415), and data sets are output. The output data sets are input into the data set storage (1445) and used.

[0213] The indices (1450) of all possible combinations are determined according to the quantized range of each index. For example, if the feature map output from a thin neural network including one hidden layer has four channels ({c1, c2, c3, c4}) and ck|k=1, 2, 3, 4∈[-2, 2], then [-2, -1, 0, 1, 2] can be mapped to [1, 2, 3, 4, 5]. If the indices for the data set are four indices ({i1, i2, i3, i4}), then ik|k=1, 2, 3, 4∈[1, 5], so the value of the index can be one of five values ​​from 1 to 5. The output value determined according to the value of the index can be a value corresponding to table[i1][i2][i3][i4]. Here, the size of the lookup table becomes 4x5.

[0214] The parameters for the first neural network (1405) obtained as a result of training can be stored in the image encoding device (1300). In the case of the second neural network (1415), it is used to obtain data sets stored in the data set storage (1445), and since the data set storage (1445) is used thereafter and the second neural network (1415) is not used, the parameters for the second neural network (1415) do not need to be stored.

[0215] FIG. 15 is a diagram illustrating a method for training neural networks and a method for acquiring data sets according to one embodiment of the present disclosure.

[0216] Referring to FIG. 15, the current image for training (1500) corresponds to the current image described above, the current decoded image for training (1501) corresponds to the current decoded image (301) described above, the current predicted image for training (1502) corresponds to the current predicted image (302) described above, the training coding context information (1503) for the current image for training (1500) corresponds to the coding context information (303) for the current image described above, and the training output image (1545) corresponds to the output image (340) described above.

[0217] In a training process according to one embodiment of the present disclosure, neural networks can be trained so that the training output image (1545) is as similar as possible to the training current image (1500). To this end, as illustrated in FIG. 15, loss information (1555) can be used to train the neural networks.

[0218] Specifically, the training process of the neural networks will be described. The training current decoded image (1501), the training current predicted image (1502), and the training coding context information (1503) for the training current image (1500) are input to the first neural network (1505), and the feature map and weight information (1513, 1514, 1515) for the training current image (1500) are output. The first neural network (1505) is a neural network including at least one hidden layer.

[0219] The feature map may be three feature maps for each of the three index mapping methods (1510, 1511, 1512). Alternatively, the feature map may be divided into three parts and applied to each of the three index mapping methods (1510, 1511, 1512).

[0220] The weight information (1513, 1514, 1515) may be weight information corresponding to each index mapping method (1510, 1511, 1512). In addition, the weight information (1513, 1514, 1515) may be weight information corresponding to a feature map that is input and divided into each index mapping method (1510, 1511, 1512).

[0221] First index mapping (1510) is performed using a feature map to obtain first indices indicating a pixel correction value for one pixel among a plurality of pixels included in a current training image (1500). Second index mapping (1511) is performed using the feature map to obtain second indices indicating a pixel correction value for one pixel among a plurality of pixels included in a current training image (1500). Third index mapping (1512) is performed using the feature map to obtain third indices indicating a pixel correction value for one pixel among a plurality of pixels included in a current training image (1500).

[0222] The index mapping methods (1510, 1511, 1512) used in the training process of neural networks may be approximate index mappings in which feature maps are obtained as indices containing arbitrary noise. The index mapping methods (1510, 1511, 1512) may include nonlinear mappings such as clipping, stochastic quantization, or noisy quantization.

[0223] First indices are input into a second neural network (1516) to obtain first pixel correction values, and the first pixel correction values ​​are multiplied by first weight information (1513) (1526) to obtain first pixel correction values ​​multiplied by weight. Second indices are input into a third neural network (1517) to obtain second pixel correction values, and the second pixel correction values ​​are multiplied by second weight information (1514) (1527) to obtain second pixel correction values ​​multiplied by weight. Third indices are input into a fourth neural network (1518) to obtain third pixel correction values, and the third pixel correction values ​​are multiplied by third weight information (1515) (1528) to obtain third pixel correction values ​​multiplied by weight. Each of the second neural network (1516), the third neural network (1517), and the fourth neural network (1518) is a deep neural network including a plurality of hidden layers. The final pixel correction values ​​are obtained by adding (1530) the first pixel correction values ​​multiplied by weights, the second pixel correction values ​​multiplied by weights, and the third pixel values ​​multiplied by weights. The training output image (1545) is obtained by adding (1540) the final pixel correction values ​​and the sample values ​​included in the current decoded training image (1501). The training current image (1500) and the training output image (1545) are compared (1550) to obtain loss information (1555).

[0224] The loss information (1555) may correspond to a difference between a current training image (1500) and a training output image (1545). In one embodiment of the present disclosure, the difference between the current training image (1500) and the training output image (1545) may include at least one of an L1-norm value, an L2-norm value, a Structural Similarity (SSIM) value, a Peak Signal-To-Noise Ratio-Human Vision System (PSNR-HVS) value, a Multiscale SSIM (MS-SSIM) value, a Variance Inflation Factor (VIF) value, or a Video Multimethod Assessment Fusion (VMAF) value between the current training image (1500) and the training output image (1545).

[0225] Since the loss information (1555) is related to the quality of the training output image (1545), it may also be referred to as quality loss information.

[0226] The first neural network (1505), the second neural network (1516), the third neural network (1517), and the fourth neural network (1518) can be trained so that the final loss information derived from the loss information (1555) is reduced or minimized.

[0227] In one embodiment of the present disclosure, the first neural network (1505), the second neural network (1516), the third neural network (1517), and the fourth neural network (1518) can reduce or minimize the final loss information by changing the values ​​of preset parameters.

[0228] In one embodiment of the present disclosure, the final loss information can be calculated according to the following mathematical expression 2.

[0229] [Equation 2]

[0230] Final loss information = b*loss information

[0231] In mathematical expression 2, b is a weight applied to the loss information (1555).

[0232] According to mathematical expression 2, it can be seen that the first neural network (1505), the second neural network (1516), the third neural network (1517), and the fourth neural network (1518) are trained in a direction in which the training output image (1555) becomes as similar as possible to the training current image (1500).

[0233] The training process described with reference to FIG. 15 can be performed by a training device. The training device can be, for example, an image encoding device (1300) or a separate server.

[0234] The trained second neural network (1516) is used to obtain a first data set stored in the first data set storage (1565). Specifically, all possible combinations of table indexes (1560) are input into the trained second neural network (1516), and the first data sets are output. The output first data sets are input into the first data set storage (1565) and used.

[0235] The trained third neural network (1517) is used to obtain a second data set stored in the second data set storage (1575). Specifically, all possible combinations of table indexes (1570) are input into the trained third neural network (1517), and the second data sets are output. The output second data sets are input into the second data set storage (1575) and used.

[0236] The trained fourth neural network (1518) is used to obtain a third data set stored in the third data set storage (1585). Specifically, all possible combinations of table indexes (1580) are input into the trained fourth neural network (1518), and the third data sets are output. The output third data sets are input into the third data set storage (1585) and used.

[0237] The parameters for the first neural network (1505) obtained as a result of training can be stored in the image encoding device (1300). In the case of the second neural network (1516), it is used to obtain the first data sets stored in the first data set storage (1565), and since the first data set storage (1565) is used thereafter and the second neural network (1516) is not used, the parameters for the second neural network (1516) do not need to be stored. In the case of the third neural network (1517) and the fourth neural network (1518), they are used to obtain the second data sets stored in the second data set storage (1575) and the third data sets stored in the third data set storage (1585), respectively, and thereafter, the second data set storage (1575) and the third data set storage (1585) are used and the third neural network (1517) and the fourth neural network (1518) are not used, so parameters for the third neural network (1517) and the fourth neural network (1518) do not need to be stored.

[0238] An image decoding method according to one embodiment of the present disclosure may include: inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels; mapping an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map; identifying the pixel correction value based on the index; obtaining a corrected pixel value for the at least one pixel using the identified pixel correction value; and obtaining an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0239] An image decoding method according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by identifying pixel correction values ​​based on indices.

[0240] According to one embodiment of the present disclosure, the neural network may include at least one hidden layer.

[0241] An image decoding method according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using a neural network including one layer.

[0242] An image decoding method according to one embodiment of the present disclosure further includes a step of obtaining a scale factor indicating a scale of pixel correction; and the modified pixel value can be determined based on the identified pixel correction value and the scale factor.

[0243] In an image decoding method according to one embodiment of the present disclosure, the degree of application of a pixel correction value can be appropriately determined by using a scale factor.

[0244] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in a predetermined data set.

[0245] An image decoding method according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by using an index indicating one of the pixel correction values ​​included in a predetermined data set.

[0246] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in one of a plurality of predetermined data sets.

[0247] An image decoding method according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by using an index indicating one of pixel correction values ​​included in one of a plurality of predetermined data sets.

[0248] According to one embodiment of the present disclosure, the predetermined data set or a plurality of predetermined data sets may be stored in the form of a lookup table.

[0249] In an image decoding method according to one embodiment of the present disclosure, memory capacity and computational amount can be reduced by using a lookup table.

[0250] According to one embodiment of the present disclosure, the neural network may be trained by inputting a training decoded image, a training predicted image, and training coding context information for a training current image, outputting a training feature map, mapping a training index according to a predetermined index mapping method using the training feature map, inputting the training index into a second neural network including a plurality of hidden layers, obtaining a training pixel correction value, and obtaining a training output image based on the training pixel correction value.

[0251] As described above, the image decoding method according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using a trained neural network.

[0252] According to one embodiment of the present disclosure, at least one data set including the index can be obtained by inputting all possible indices into the second neural network.

[0253] An image decoding method according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using at least one data set obtained by inputting all possible indices into a second neural network.

[0254] An image decoding method according to one embodiment of the present disclosure may include the steps of: inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels; mapping an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map; identifying the pixel correction value based on the index; obtaining a corrected pixel value for the at least one pixel using the identified pixel correction value; and obtaining an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0255] An image decoding method according to one embodiment of the present disclosure can achieve the same coding efficiency as using a large data set by mixing multiple small data sets, while reducing memory capacity.

[0256] According to one embodiment of the present disclosure, the final pixel correction value can be obtained using a weighted sum of the plurality of pixel correction values.

[0257] An image decoding method according to one embodiment of the present disclosure can achieve the same coding efficiency as using a large data set while reducing memory capacity by using a weighted sum of the plurality of pixel correction values.

[0258] An image decoding device according to one embodiment of the present disclosure may include a memory storing one or more instructions; and at least one processor operating according to the one or more instructions. The at least one processor may input a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels. The at least one processor may map an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map. The at least one processor may identify the pixel correction value based on the index. The at least one processor may obtain a corrected pixel value for the at least one pixel using the identified pixel correction value. The at least one processor may obtain an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0259] An image decoding device according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by identifying pixel correction values ​​based on indices.

[0260] According to one embodiment of the present disclosure, the neural network may include at least one hidden layer.

[0261] An image decoding device according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using a neural network including one layer.

[0262] An image decoding device according to one embodiment of the present disclosure obtains a scale factor indicating a scale of pixel correction, and the modified pixel value can be determined based on the identified pixel correction value and the scale factor.

[0263] An image decoding device according to one embodiment of the present disclosure can appropriately determine the degree of application of a pixel correction value by using a scale factor.

[0264] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in a predetermined data set.

[0265] An image decoding device according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by using an index indicating one of pixel correction values ​​included in a predetermined data set.

[0266] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in one of a plurality of predetermined data sets.

[0267] An image decoding device according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by using an index indicating one of pixel correction values ​​included in one of a plurality of predetermined data sets.

[0268] According to one embodiment of the present disclosure, the predetermined data set or a plurality of predetermined data sets may be stored in the form of a lookup table.

[0269] An image decoding device according to one embodiment of the present disclosure can reduce memory capacity and computational effort by using a lookup table.

[0270] According to one embodiment of the present disclosure, the neural network may be trained by inputting a training decoded image, a training predicted image, and training coding context information for a training current image, outputting a training feature map, mapping a training index according to a predetermined index mapping method using the training feature map, inputting the training index into a second neural network including a plurality of hidden layers, obtaining a training pixel correction value, and obtaining a training output image based on the training pixel correction value.

[0271] As described above, the image decoding device according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using a trained neural network.

[0272] According to one embodiment of the present disclosure, at least one data set including the index can be obtained by inputting all possible indices into the second neural network.

[0273] An image decoding device according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using at least one data set obtained by inputting all possible indices into a second neural network.

[0274] An image decoding device according to one embodiment of the present disclosure may include a memory storing one or more instructions; and at least one processor operating according to the one or more instructions. The at least one processor may input a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels. The at least one processor may map a plurality of indices indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map. The at least one processor may identify a plurality of pixel correction values ​​based on the plurality of indices. The at least one processor may obtain a final pixel correction value for the at least one pixel using the identified plurality of pixel correction values. The at least one processor may obtain a corrected pixel value for the at least one pixel using the obtained final pixel correction value. The at least one processor may obtain an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0275] An image decoding device according to one embodiment of the present disclosure can achieve the same coding efficiency as using a large data set by mixing a plurality of small data sets, while reducing the capacity of memory.

[0276] According to one embodiment of the present disclosure, the final pixel correction value can be obtained using a weighted sum of the plurality of pixel correction values.

[0277] An image decoding device according to one embodiment of the present disclosure can achieve the same coding efficiency as using a large data set while reducing memory capacity by using a weighted sum of the plurality of pixel correction values.

[0278] An image encoding method according to one embodiment of the present disclosure may include: inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels; mapping an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map; identifying the pixel correction value based on the index; obtaining a corrected pixel value for the at least one pixel using the identified pixel correction value; and obtaining an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0279] An image encoding method according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by identifying pixel correction values ​​based on indices.

[0280] According to one embodiment of the present disclosure, the neural network may include at least one hidden layer.

[0281] An image encoding method according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using a neural network including one layer.

[0282] An image encoding method according to one embodiment of the present disclosure further includes a step of generating a scale factor indicating a scale of pixel correction; and the modified pixel value can be determined based on the identified pixel correction value and the scale factor.

[0283] In an image encoding method according to one embodiment of the present disclosure, the degree of application of a pixel correction value can be appropriately determined by using a scale factor.

[0284] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in a predetermined data set.

[0285] An image encoding method according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by using an index indicating one of pixel correction values ​​included in a predetermined data set.

[0286] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in one of a plurality of predetermined data sets.

[0287] In an image encoding method according to one embodiment of the present disclosure, memory capacity and computational amount can be reduced by using an index indicating one of pixel correction values ​​included in one of a plurality of predetermined data sets.

[0288] According to one embodiment of the present disclosure, the predetermined data set or a plurality of predetermined data sets may be stored in the form of a lookup table.

[0289] In an image encoding method according to one embodiment of the present disclosure, memory capacity and computational amount can be reduced by using a lookup table.

[0290] According to one embodiment of the present disclosure, the neural network may be trained by inputting a training encoded image, a training predicted image, and training coding context information for a training current image, outputting a training feature map, mapping a training index according to a predetermined index mapping method using the training feature map, inputting the training index into a second neural network including a plurality of hidden layers, obtaining a training pixel correction value, and obtaining a training output image based on the training pixel correction value.

[0291] As described above, the image encoding method according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using a trained neural network.

[0292] According to one embodiment of the present disclosure, at least one data set including the index can be obtained by inputting all possible indices into the second neural network.

[0293] In an image encoding method according to one embodiment of the present disclosure, memory capacity and computational amount can be reduced by using at least one data set obtained by inputting all possible indices into a second neural network.

[0294] An image encoding method according to one embodiment of the present disclosure may include the steps of: inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels; mapping an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map; identifying the pixel correction value based on the index; obtaining a corrected pixel value for the at least one pixel using the identified pixel correction value; and obtaining an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0295] An image encoding method according to one embodiment of the present disclosure can achieve the same coding efficiency as using a large data set by mixing a plurality of small data sets, while reducing memory capacity.

[0296] According to one embodiment of the present disclosure, the final pixel correction value can be obtained using a weighted sum of the plurality of pixel correction values.

[0297] An image encoding method according to one embodiment of the present disclosure can achieve the same coding efficiency as using a large data set while reducing memory capacity by using a weighted sum of the plurality of pixel correction values.

[0298] An image encoding device according to one embodiment of the present disclosure may include a memory storing one or more instructions; and at least one processor operating according to the one or more instructions. The at least one processor may input a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels. The at least one processor may map an index indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map. The at least one processor may identify the pixel correction value based on the index. The at least one processor may obtain a corrected pixel value for the at least one pixel using the identified pixel correction value. The at least one processor may obtain an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0299] An image encoding device according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by identifying pixel correction values ​​based on indices.

[0300] According to one embodiment of the present disclosure, the neural network may include at least one hidden layer.

[0301] An image encoding device according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using a neural network including one layer.

[0302] An image encoding device according to one embodiment of the present disclosure generates a scale factor indicating a scale of pixel correction, and the modified pixel value can be determined based on the identified pixel correction value and the scale factor.

[0303] An image encoding device according to one embodiment of the present disclosure can appropriately determine the degree of application of a pixel correction value by using a scale factor.

[0304] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in a predetermined data set.

[0305] An image encoding device according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by using an index indicating one of pixel correction values ​​included in a predetermined data set.

[0306] According to one embodiment of the present disclosure, the index may represent one of the pixel correction values ​​included in one of a plurality of predetermined data sets.

[0307] An image encoding device according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by using an index indicating one of pixel correction values ​​included in one of a plurality of predetermined data sets.

[0308] According to one embodiment of the present disclosure, the predetermined data set or a plurality of predetermined data sets may be stored in the form of a lookup table.

[0309] An image encoding device according to one embodiment of the present disclosure can reduce memory capacity and computational effort by using a lookup table.

[0310] According to one embodiment of the present disclosure, the neural network may be trained by inputting a training encoded image, a training predicted image, and training coding context information for a training current image, outputting a training feature map, mapping a training index according to a predetermined index mapping method using the training feature map, inputting the training index into a second neural network including a plurality of hidden layers, obtaining a training pixel correction value, and obtaining a training output image based on the training pixel correction value.

[0311] As described above, the image encoding device according to one embodiment of the present disclosure can reduce memory capacity and computational amount by using a trained neural network.

[0312] According to one embodiment of the present disclosure, at least one data set including the index can be obtained by inputting all possible indices into the second neural network.

[0313] An image encoding device according to one embodiment of the present disclosure can reduce memory capacity and computational complexity by using at least one data set obtained by inputting all possible indices into a second neural network.

[0314] An image encoding device according to one embodiment of the present disclosure may include a memory storing one or more instructions; and at least one processor operating according to the one or more instructions. The at least one processor may input a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels. The at least one processor may map a plurality of indices indicating a pixel correction value for at least one pixel among the plurality of pixels based on the feature map. The at least one processor may identify a plurality of pixel correction values ​​based on the plurality of indices. The at least one processor may obtain a final pixel correction value for the at least one pixel using the identified plurality of pixel correction values. The at least one processor may obtain a corrected pixel value for the at least one pixel using the obtained final pixel correction value. The at least one processor may obtain an output image corrected such that the at least one pixel from the current image has the corrected pixel value.

[0315] An image encoding device according to one embodiment of the present disclosure can achieve the same coding efficiency as using a large data set by mixing a plurality of small data sets, while reducing the capacity of memory.

[0316] According to one embodiment of the present disclosure, the final pixel correction value can be obtained using a weighted sum of the plurality of pixel correction values.

[0317] An image encoding device according to one embodiment of the present disclosure can achieve the same coding efficiency as using a large data set while reducing memory capacity by using a weighted sum of the plurality of pixel correction values.

[0318] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0319] According to one embodiment of the present disclosure, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

Claims

1. A step of inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels (S810); A step of mapping an index representing a pixel correction value for at least one pixel among the plurality of pixels based on the feature map (S830); A step of identifying the pixel compensation value based on the above index (S850); A step of obtaining a corrected pixel value for at least one pixel based on the identified pixel correction value (S870); and An image decoding method, comprising: a step (S890) of obtaining an output image corrected so that at least one pixel from the current image has the corrected pixel value; 2. In paragraph 1, A method for decoding an image, wherein the neural network includes at least one hidden layer.

3. In paragraph 1 or 2, A step of obtaining a scale factor representing a scale of pixel correction; further comprising; An image decoding method, wherein the modified pixel value is determined based on the identified pixel correction value and the scale factor.

4. In any one of paragraphs 1 to 3, An image decoding method, wherein the above index represents one of the pixel correction values ​​included in a predetermined data set.

5. In any one of paragraphs 1 to 3, An image decoding method, wherein the above index represents one of the pixel correction values ​​included in one of a plurality of predetermined data sets.

6. In paragraph 4 or 5, An image decoding method, wherein the above-determined data set or the above-determined plurality of data sets are stored in the form of a lookup table.

7. In any one of paragraphs 1 to 6, The above neural network takes as input the training decoded image, the training predicted image, and the training coding context information for the training current image and outputs the training feature map. Using the above training feature map, the training index is mapped according to a predefined index mapping method, By inputting the above training index into a second neural network including multiple hidden layers, a training pixel correction value is obtained, An image decoding method, wherein the training is performed by obtaining a training output image based on the above training pixel correction values.

8. In paragraph 7, An image decoding method, wherein at least one data set including the above index is obtained by inputting all possible indices into the second neural network.

9. A step of inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels (S910); A step of mapping a plurality of indices representing a pixel correction value for at least one pixel among the plurality of pixels based on the feature map (S920); A step of identifying a plurality of pixel correction values ​​based on the plurality of indices (S930); A step of obtaining a final pixel compensation value for at least one pixel by using the identified plurality of pixel compensation values ​​(S940); A step of obtaining a corrected pixel value for at least one pixel by using the obtained final pixel correction value (S950); An image decoding method, comprising: a step of obtaining an output image corrected so that at least one pixel from the current image has the corrected pixel value (S960); 10. In paragraph 9, An image decoding method, wherein the final pixel compensation value is obtained by using a weighted sum of the plurality of pixel compensation values.

11. A step of inputting a decoded image, a predicted image, and coding context information for a current image including a plurality of pixels into a neural network to obtain a feature map for the plurality of pixels (S1110); A step of mapping an index representing a pixel correction value for at least one pixel among the plurality of pixels based on the feature map (S1130); A step of identifying the pixel compensation value based on the above index (S1150); A step of obtaining a corrected pixel value for at least one pixel based on the identified pixel correction value (S1170); and An image encoding method, comprising the step of obtaining an output image corrected such that at least one pixel from the current image has the corrected pixel value (S1190).

12. In paragraph 11, A method for encoding an image, wherein the neural network includes at least one hidden layer.

13. In paragraph 11 or 12, further comprising a step of generating a scale factor representing the scale of pixel correction; An image encoding method, wherein the modified pixel value is determined based on the identified pixel correction value and the scale factor.

14. In one of the clauses 11 to 13, A method for encoding an image, wherein the above index represents one of the pixel correction values ​​included in a predetermined data set.

15. In one of the clauses 11 to 13, An image encoding method, wherein the above index represents one of the pixel correction values ​​included in one of a plurality of predetermined data sets.

Citation Information

Patent Citations

  • Pop-up retial franchising and complex econmic system

    US20220292543A1

  • Automated Prediction of Pixel Error Noticeability

    US20230028420A1

  • System and method for generating soil moisture data from satellite imagery using deep learning model

    US20230064454A1

  • Dynamic obstacle avoidance method based on real-time local grid map construction

    US20230161352A1

  • KR20230141616A