Image decoding method, image decoding device, image encoding method, and image encoding device for performing in-loop filtering using neural network

By using neural networks for in-loop filtering in image encoding and decoding, and generating filtering neural network layer information using the feature map state information of the reference image, the problems of image quality improvement and difference reduction in the prior art are solved, and more efficient image processing results are achieved.

CN121753343APending Publication Date: 2026-03-27SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing image encoding and decoding technologies struggle to effectively utilize in-loop filtering to improve image quality and reduce the difference between the original and decoded images when employing artificial intelligence.

Method used

By using a neural network for in-loop filtering, layer information is generated and applied to encode and decode the image. The layer information of the filtering neural network is generated using the state information of the feature map from the reference image, and the information of the current image is input into these layers for processing, thus realizing in-loop filtering.

Benefits of technology

It improves image quality and reduces the difference between the original and decoded images. It reconstructs details through in-loop filtering of neural networks, thereby enhancing the image encoding and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121753343A_ABST
    Figure CN121753343A_ABST
Patent Text Reader

Abstract

There is provided an image decoding method including: generating layer information to be used in at least one layer within a filtering neural network based on first state information obtained from a feature map corresponding to a first reference image of a current image or second state information obtained from a feature map corresponding to a second reference image of the current image; inputting information corresponding to the current image into at least one layer in the filtering neural network; and obtaining an image to which in-loop filtering has been applied to the current image by inputting information corresponding to the current image to at least one layer of the used layer information, in which state information of the current image is obtained from a feature map corresponding to the current image and has a lower resolution than a resolution of the current image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to the field of image encoding and decoding, and to a method and apparatus of encoding and decoding an image by performing in-loop filtering on a current block included in a current image using a neural network. BACKGROUND

[0002] In codecs such as H.264 Advanced Video Coding (AVC) and High Efficiency Video Coding (HEVC), an image can be divided into blocks, and each of the blocks can be predictively encoded and predictively decoded through inter prediction or intra prediction.

[0003] Intra prediction is a method of compressing an image by removing spatial redundancy within an image, and inter prediction is a method of compressing an image by removing temporal redundancy between images.

[0004] As a representative example of inter prediction, there is motion estimation encoding. Motion estimation encoding predicts a block of a current image by using a reference image. A reference block most similar to the current block can be searched for in a predefined search range by using a predefined evaluation function. The current block is predicted based on the reference block, and a residual block is generated and encoded by subtracting a prediction block generated as a prediction result from the current block.

[0005] In order to derive a motion vector indicating a reference block in a reference image, a motion vector of a previously encoded block can be used as a motion vector predictor of the current block. A differential motion vector that is a difference between the motion vector of the current block and the motion vector predictor of the current block is signaled to the decoder side through a predefined method.

[0006] Recently, a technology of encoding / decoding an image by using artificial intelligence (AI) has been proposed. Accordingly, there is a need for a method of efficiently encoding / decoding an image by using AI (e.g., a neural network). SUMMARY

[0007] Solution to the problem In an embodiment, an image decoding method can include generating layer information to be used in at least one layer within a filtering neural network based on first state information obtained from a feature map corresponding to a first reference image of a current image or second state information obtained from a feature map corresponding to a second reference image of the current image. The image decoding method can include inputting information corresponding to the current image to the at least one layer within the filtering neural network. The image decoding method can include obtaining an image to which in-loop filtering has been applied to the current image by inputting the information corresponding to the current image to the at least one layer having used the layer information. The state information of the current image can be obtained from a feature map corresponding to the current image and can have a lower resolution than a resolution of the current image.

[0008] In an embodiment, an image decoding apparatus can include at least one memory storing at least one instruction, and at least one processor configured to operate according to the at least one instruction. The at least one processor can be further configured to generate layer information to be used in at least one layer within a filtering neural network based on first state information obtained from a feature map corresponding to a first reference picture of a current picture or second state information obtained from a feature map corresponding to a second reference picture of the current picture. The at least one processor can be further configured to input information corresponding to the current picture to the at least one layer within the filtering neural network. The at least one processor can be further configured to obtain a picture to which in-loop filtering has been applied to the current picture by inputting the information corresponding to the current picture to the at least one layer which has used the layer information. State information of the current picture can be obtained from a feature map corresponding to the current picture and can have a lower resolution than a resolution of the current picture.

[0009] In an embodiment, an image encoding method can include generating layer information to be used in at least one layer within a filtering neural network based on first state information obtained from a feature map corresponding to a first reference picture of a current picture or second state information obtained from a feature map corresponding to a second reference picture of the current picture. In an embodiment, the image encoding method can include inputting information corresponding to the current picture to the at least one layer within the filtering neural network. In an embodiment, the image encoding method can include obtaining a picture to which in-loop filtering has been applied to the current picture by inputting the information corresponding to the current picture to the at least one layer which has used the layer information. State information of the current picture can be obtained from a feature map corresponding to the current picture and can have a lower resolution than a resolution of the current picture.

[0010] In an embodiment, an image encoding apparatus can include at least one memory storing at least one instruction, and at least one processor configured to operate according to the at least one instruction. The at least one processor can be further configured to generate layer information to be used in at least one layer within a filtering neural network based on first state information obtained from a feature map corresponding to a first reference picture of a current picture or second state information obtained from a feature map corresponding to a second reference picture of the current picture. The at least one processor can be further configured to input information corresponding to the current picture to the at least one layer within the filtering neural network. The at least one processor can be further configured to obtain a picture to which in-loop filtering has been applied to the current picture by inputting the information corresponding to the current picture to the at least one layer which has used the layer information. State information of the current picture can be obtained from a feature map corresponding to the current picture and can have a lower resolution than a resolution of the current picture.

[0011] In an embodiment, a computer-readable recording medium having a bitstream recorded thereon can be provided. The bitstream can include prediction information used to predict a current picture. Based on the prediction information, layer information to be used in at least one layer within a filtering neural network can be generated based on first state information obtained from a feature map corresponding to a first reference picture or second state information obtained from a feature map corresponding to a second reference picture. Information corresponding to the current picture can be input to the at least one layer within the filtering neural network. A picture to which in-loop filtering has been applied to the current picture can be obtained by inputting the information corresponding to the current picture to the at least one layer that has used the layer information. State information of the current picture can be obtained from a feature map corresponding to the current picture and can have a lower resolution than a resolution of the current picture.

[0012] Advantages of the Invention According to the present disclosure, details can be reconstructed and image quality can be improved through in-loop filtering using a neural network. In addition, according to the present disclosure, a difference between an original picture and a decoded picture can be reduced. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 FIG. 1 shows a block diagram of an image encoding and decoding system that performs in-loop filtering according to an embodiment of the present disclosure.

[0014] Figure 2 FIG. 2 is a diagram showing blocks divided from an image according to a tree structure according to an embodiment of the present disclosure.

[0015] Figure 3 FIG. 3 is a diagram for describing a group of pictures (GOP) structure according to an embodiment of the present disclosure.

[0016] Figure 4 FIG. 4 is a diagram for describing a relationship between a current block and a reference block according to an embodiment of the present disclosure.

[0017] Figure 5 FIG. 5 is a diagram showing an in-loop filtering operation according to an embodiment of the present disclosure.

[0018] Figure 6 FIG. 6 shows a block diagram of an in-loop filtering unit according to an embodiment of the present disclosure.

[0019] Figure 7 FIG. 7 is a diagram for describing a process of outputting an intermediate picture from a filtering neural network according to an embodiment of the present disclosure.

[0020] Figures 8a to 8c FIG. 8 is a diagram for describing a filtering neural network using layer information according to an embodiment of the present disclosure.

[0021] Figure 9is a diagram illustrating a structure of a feature extraction neural network included in a filtering neural network according to an embodiment of the disclosure.

[0022] Figure 10 is a diagram illustrating a structure of a regression neural network included in a filtering neural network according to an embodiment of the disclosure.

[0023] Figures 11a to 11c is a diagram illustrating a structure of a layer information neural network according to an embodiment of the disclosure.

[0024] Figure 12 is a diagram illustrating a method of training a filtering neural network according to an embodiment of the disclosure.

[0025] Figure 13 is a block diagram illustrating a configuration of an image decoding apparatus according to an embodiment of the disclosure.

[0026] Figure 14 is a flowchart of an image decoding method according to an embodiment of the disclosure.

[0027] Figure 15 is a block diagram illustrating a configuration of an image encoding apparatus according to an embodiment of the disclosure.

[0028] Figure 16 is a flowchart of an image encoding method according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0029] Throughout the disclosure, the expression "at least one of a, b, or c" indicates only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.

[0030] It will be understood that the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "configuration surface" can also include instances in which one or more such surfaces are indicated.

[0031] In describing the disclosure, descriptions of technical contents well known in the technical field to which the disclosure belongs and not directly related to the disclosure will be omitted. By omitting unnecessary descriptions, the disclosure can be described more clearly without obscuring the gist of the disclosure. The terms used herein are defined by considering the functions in the disclosure, but the terms can vary according to the intention of users or those of ordinary skill in the art, precedents, etc. Therefore, the definitions should be made based on the contents of the entire specification.

[0032] For the same reason, some elements in the drawings are exaggerated, omitted, or schematically shown. In addition, the sizes of the elements do not completely reflect actual sizes. The same reference numerals are assigned to the same or corresponding elements in the drawings.

[0033] Advantages and features of the present disclosure and methods of accomplishing the same will be clarified by referring to embodiments described in detail below with reference to the accompanying drawings. However, the present disclosure is not limited to the following embodiments and can be implemented in various forms. The embodiments presented below are provided so that the present disclosure will be thorough and complete, and will fully convey the inventive idea of the present disclosure to those skilled in the art. The embodiments of the present disclosure can be defined by the claims. Throughout the specification, like reference numerals denote like elements. Also, in describing the embodiments of the present disclosure, detailed description of functions or configurations that are determined to make the gist of the present disclosure unnecessarily obscure can be omitted herein. The terms used herein are defined by considering the functions in the present disclosure, but the terms can vary according to the intention of users or those skilled in the art, precedents, etc. Therefore, the definition should be made based on the contents of the entire specification.

[0034] It will be understood that the blocks in the flowcharts and combinations of the flowcharts in the present disclosure can be performed by one or more computer programs including computer-executable instructions. The one or more computer programs can be stored in a single memory or can be segmented and stored in a plurality of different memories.

[0035] In embodiments, it will be understood that the individual blocks of the flowcharts and combinations of the flowcharts can be performed by computer program instructions. The computer program instructions can be installed on a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, and the instructions to be executed by the processor of the computer or other programmable data processing apparatus can generate a device for performing the functions described in the block(s) of the flowchart. Because the computer program instructions can also be stored in a computer-executable or computer-readable memory that can instruct the computer or other programmable data processing apparatus to implement functions in a specific manner, the instructions stored in the computer-executable or computer-readable memory can also produce an article of manufacture containing the instruction device for performing the functions described in the block(s) of the flowchart. The computer program instructions can also be installed on a computer or other programmable data processing apparatus.

[0036] In addition, each block in the flowcharts can represent a part of a module, a segment, or code including one or more executable instructions for performing a specific logical function(s). In embodiments, the functions mentioned in the blocks can not occur in order. For example, two blocks shown in succession can actually be executed substantially concurrently or at times in the reverse order, depending on involved functions.

[0037] All functions or operations described in the disclosure can be processed by a single processor or a combination of processors. The single processor or the combination of processors is a circuit that performs processing and can include a circuit such as an application processor (AP), a communication processor (CP), a graphics processing unit (GPU), a neural processing unit (NPU), a micro processing unit (MPU), a system on chip (SoC), or an integrated chip (IC).

[0038] The term "unit" used in the embodiments of the disclosure refers to a software element or a hardware element such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC), and the "unit" can perform a specific function. However, the term "unit" is not limited to software or hardware. The term "unit" can be configured in an addressable storage medium, or can be configured to reproduce one or more processors. In an embodiment, the term "unit" can include elements such as software elements, object-oriented software elements, class elements, and task elements, processes, functions, attributes, procedures, sub-routines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided by specific elements or specific "units" can be combined to reduce the number thereof or can be divided into additional elements. Furthermore, in an embodiment, a "unit" can include one or more processors.

[0039] In the disclosure, an "image" can mean a still image, a picture, a frame, a moving image composed of a plurality of consecutive still images, or a video.

[0040] In the disclosure, a "neural network" is a representative example of an artificial neural network model that simulates a brain nerve, and is not limited to an artificial neural network model using a specific algorithm. The neural network can be referred to as at least one of a deep neural network, a convolutional neural network, and a transformer neural network.

[0041] In the disclosure, a "current image" refers to an image that is currently to be processed, and a "previous image" refers to an image that is processed before the current image.

[0042] In the disclosure, a "sample point" refers to data to be processed as data assigned to a sampling position within an image, a feature map, or feature data. For example, a sample point can include a pixel within a two-dimensional image.

[0043] In the disclosure, "state information" can be information that comprehensively represents abstract features of an image, such as a main object included in a specific image, a shape of an object, a direction of a line, and a color distribution.

[0044] Figure 1 A block diagram of an image encoding and decoding system that performs in-loop filtering according to an embodiment of the disclosure is shown.

[0045] In an embodiment, the image encoding apparatus 110 transmits a bitstream generated by image encoding to the image decoding apparatus 150, and the image decoding apparatus 150 reconstructs an image by receiving and decoding the bitstream.

[0046] In an embodiment, in the image encoding apparatus 110, the prediction encoding unit 115 outputs a prediction block through inter prediction and intra prediction, and the transform and quantization unit 120 outputs quantized transform coefficients by transforming and quantizing residual samples of a residual block between the prediction block and a current block. The entropy encoding unit 125 outputs a bitstream by encoding the quantized transform coefficients.

[0047] In an embodiment, the quantized transform coefficients are reconstructed into a residual block including residual samples in a spatial domain by the inverse quantization and inverse transform unit 130. A reconstructed block that is a combination of the prediction block and the residual block is output as a filtered block by the in-loop filter unit 135. A reconstructed image including the filtered block can be used as a reference image for a next input image in the prediction encoding unit 115.

[0048] In an embodiment, a bitstream received by the image decoding apparatus 150 is reconstructed into a residual block including residual samples in a spatial domain by the entropy decoding unit 155 and the inverse quantization and inverse transform unit 160. The residual block and a prediction block output from the prediction decoding unit 170 are combined to generate a reconstructed block, and the reconstructed block is output as a filtered block by the in-loop filter unit 165. A reconstructed image including the filtered block can be used as a reference image for a next image in the prediction decoding unit 170.

[0049] In an embodiment, the in-loop filter unit 135 of the image encoding apparatus 110 performs in-loop filtering by using filter information input according to a user input or a system setting. The filter information used by the in-loop filter unit 135 is transmitted to the image decoding apparatus 150 through the entropy encoding unit 125. The in-loop filter unit 165 of the image decoding apparatus 150 can perform in-loop filtering based on filter information input from the entropy decoding unit 155.

[0050] In an embodiment, in the image encoding and decoding process, an image is hierarchically divided, and encoding and decoding are performed on blocks divided from the image. Referring to Figure 2 A block divided from an image is described.

[0051] Figure 2 is a diagram illustrating a block divided from an image according to a tree structure.

[0052] In an embodiment, one image 200 can be divided into one or more slices or one or more parallel blocks. One slice can include a plurality of parallel blocks. On the other hand, in the present specification, a video can have the same meaning as an image or a picture.

[0053] In an embodiment, one slice or one tile can be a sequence of one or more largest coding units (CUs).

[0054] In an embodiment, one largest CU can be divided into one or more CUs. The CU can be a reference block for determining a prediction mode. In other words, it can be determined whether an intra prediction mode is applied to each CU or whether an inter prediction mode is applied to each CU. In the present disclosure, the largest CU can be referred to as a largest coding block, and the CU can be referred to as a coding block.

[0055] In an embodiment, a size of the CU can be equal to or smaller than the largest CU. Since the largest CU is a CU having the largest size, the largest CU can also be referred to as a CU.

[0056] In an embodiment, one or more prediction units for intra prediction or inter prediction can be determined from the CU. A size of the prediction unit can be equal to or smaller than the CU.

[0057] In an embodiment, one or more transform units for transformation and quantization can be determined from the CU. A size of the transform unit can be equal to or smaller than the CU. The transform unit is a reference block for transformation and quantization, and the residual samples of the CU can be transformed and quantized for each transform unit within the CU.

[0058] In an embodiment, the current block of the present disclosure can be a slice, a tile, a largest CU, a CU, a prediction unit, or a transform unit divided from the picture 200. Further, a sub-block of the current block can be a block divided from the current block. For example, when the current block is a largest CU, the sub-block can be a CU, a prediction unit, or a transform unit. In addition, a parent block of the current block is a block including the current block as a part. For example, when the current block is a largest CU, the parent block can be a picture sequence, a picture, a slice, or a tile.

[0059] Figure 3 is a diagram for describing a group of pictures (GOP) structure according to an embodiment of the present disclosure.

[0060] Referring to Figure 3 A plurality of pictures can be compressed with reference to other pictures within the same group.

[0061] In an embodiment, the compression order can be determined differently from the order of the actual images, and thus, an image that is earlier or later in time can be used as a reference image. For example, the order of the actual images can be 0, 1, 2,..., 16, but compression can be performed in the order of 0, 16, 8,..., 15. For the zeroth image, since there is no image to be referenced, only an intra prediction method is used, and for the subsequent images, an inter prediction method is additionally used. In this case, the structure of 16 images from the first image to the sixteenth image is referred to as a GOP, and the GOP structure is repeatedly used. At this time, the size of the GOP is 16. The plurality of images included in the GOP can be distinguished into a plurality of layers based on the compression order or the reference relationship. For example, the eighth image, the twelfth image, and the fourteenth image, which refer to the sixteenth image, can be determined as a higher layer than the sixteenth image.

[0062] In an embodiment, in order to improve compression efficiency when compressing images, different values of a quantization parameter (QP) can be used for each image. For example, a reference image can reduce the QP to improve image quality and thus increase reference efficiency, and a reference image can increase the QP to have a lower bit rate. An image in a lower layer can use a lower QP because the image in the lower layer is more frequently referenced, and an image in a higher layer can use a relatively higher QP. An image in the highest layer is not referenced in other image compression, and thus can use the highest QP. In such an embodiment, a reference image can improve the efficiency of image compression by highlighting good image quality of the reference image.

[0063] Figure 4 is a diagram for describing a relationship between a current block and a reference block according to an embodiment of the disclosure.

[0064] Referring to Figure 4 , the relationship between the current block and the reference block can be forward reference, bi-directional reference, backward reference, or 2 forward reference.

[0065] In an embodiment, the relationship between the current block and the reference block can be forward reference 410. When the time order of the current image is T, the time order of the reference image can be T-M earlier than the current image.

[0066] In an embodiment, the relationship between the current block and the reference block can be bi-directional reference 420. When the time order of the current image is T, the time order of the reference image can be T-M earlier than the current image and T+N later than the current image.

[0067] In an embodiment, the relationship between the current block and the reference block can be backward reference 430. When the time order of the current image is T, the time order of the reference image can be T+N later than the current image.

[0068] In an embodiment, the relationship between the current block and the reference block can be 2 forward references 440. When the temporal order of the current picture is T, the temporal order of the reference picture can be T-M and T-N earlier than the current picture.

[0069] In an embodiment, the relationship between the current block and the reference block can be 2 backward references (not shown). When the temporal order of the current picture is T, the temporal order of the reference picture can be T+M and T+N later than the current picture.

[0070] In an embodiment, the picture encoding apparatus can determine pictures that are to be used as reference pictures of the current picture in the picture decoding apparatus, and can determine motion information about sub-blocks included in the current picture.

[0071] In an embodiment, the picture decoding apparatus can determine motion information based on the relationship between the current picture and the reference picture. For example, from among a plurality of pieces of motion information about a plurality of sub-blocks of the current block, motion information for which the current picture and the reference picture satisfy a pre-defined condition can be obtained.

[0072] In an embodiment, the reference picture can be determined from among pictures of a lower layer than the current picture in the GOP.

[0073] In an embodiment, the reference picture can be determined from among pictures of the lowest layer in the GOP. For example, only the zeroth picture of layer 1 can be determined as the reference picture.

[0074] In an embodiment, it can be determined whether to downsize each of a plurality of pictures within the GOP. For example, pictures using intra prediction can be decoded in their original size without being downsized, and other pictures can be downsized and decoded. For example, pictures using intra prediction and pictures included in a specific layer (e.g., the lowest layer) can be decoded in their original size without being downsized, and other pictures can be downsized and decoded.

[0075] In an embodiment, to decode the current picture, the picture decoding apparatus can use a first reference picture included in a first reference picture list of the current picture, wherein the first reference picture list includes at least one picture having a picture order count (POC) value less than a POC value of the current picture. Alternatively, to decode the current picture, the picture decoding apparatus can use a second reference picture included in a second reference picture list of the current picture, wherein the second reference picture list includes at least one picture having a POC value greater than the POC value of the current picture.

[0076] On the other hand, the disclosure is not limited to the disclosed examples, both the first reference picture and the second reference picture can be pictures included in the first reference picture list of the current picture, or can be pictures included in the second reference picture list of the current picture.

[0077] Figure 5 FIG. 1 is a diagram illustrating an in-loop filtering operation according to an embodiment of the disclosure.

[0078] In an embodiment, the in-loop filtering unit 135 of the image encoding apparatus 110 or the in-loop filtering unit 165 (hereinafter, an in-loop filtering unit 500) of the image decoding apparatus 150 can include at least one of a deblocking filtering unit 510, a sample adaptive offset filtering unit 520, and an adaptive loop filtering unit 530.

[0079] In an embodiment, the deblocking filtering unit 510 can determine whether to perform deblocking filtering, and can perform deblocking filtering. The sample adaptive offset filtering unit 520 can determine whether to perform sample adaptive offset filtering, and can perform sample adaptive offset filtering. The adaptive loop filtering unit 530 can determine whether to perform adaptive loop filtering, and can perform adaptive loop filtering.

[0080] In an embodiment, the in-loop filtering unit 500 can store a filtered picture filtered through the deblocking filtering unit 510, the sample adaptive offset filtering unit 520, and the adaptive loop filtering unit 530 in a decoded picture buffer 540. Accordingly, at least one picture stored in the decoded picture buffer 540 can have undergone in-loop filtering, and can be used as a reference picture for generating a prediction block in inter prediction.

[0081] In an embodiment, deblocking filtering can be filtering applied to a boundary of a transform block by using a deblocking filter to reduce block artifacts occurring during a process of performing transform, prediction, and quantization.

[0082] In an embodiment, a filter length for deblocking filtering can be determined based on a size of an image component or a minimum transform block on both sides. A 4-tap, 8-tap, or 14-tap filter can be applied to a transform block of a luma sample, and a 6-tap filter can be applied to a transform block of a chroma sample. On the other hand, a size or length of a tap for deblocking filtering is not limited to the disclosed examples.

[0083] In embodiments, the filter length for deblocking filtering can be determined based on the smoothness or boundary conditions of the boundaries between transform blocks. For example, when filtering image data, the image decoding device 150 may identify regions including high-dispersion samples as edges and may not perform deblocking filtering to prevent the edges of objects from being blurred. The image decoding device 150 may calculate the smoothness within the image data or identify whether boundary conditions are met, and may filter the boundaries of transform blocks only if the boundaries of the transform blocks are identified as flat. Boundary conditions may vary depending on whether the boundary is a vertical or horizontal boundary, whether the transform block is used for luminance samples, or whether the transform block is used for chrominance samples. The image decoding device 150 may obtain information associated with the boundary conditions from the bitstream, and the boundary conditions may be preset. Furthermore, filter coefficients associated with deblocking filtering may be preset.

[0084] In an embodiment, sample adaptive offset filtering may involve classifying samples from neighboring blocks and applying an offset filter to the classified samples to reduce ringing artifacts. For example, sample adaptive offset filtering may be a filter that reduces the error between the reconstructed image and the original image by adding an offset to at least one sample included in an image that has already undergone deblocking filtering.

[0085] In an embodiment, the sample adaptive offset filtering can determine the sample characteristics within a block as one of not using sample adaptive offset filtering, performing edge offset filtering, or performing offset filtering.

[0086] In this embodiment, edge offset filtering is performed when a block is identified to have an edge in a specific direction and there is an error in the sample points along that edge direction. The image decoding device 150 can perform edge offset filtering on the block by obtaining the class corresponding to the block and four edge offset values ​​for the purpose of edge offset filtering.

[0087] In this embodiment, band-off filtering is a filtering method that classifies samples within a block into brightness bands with similar brightness values ​​and uses the offsets of multiple consecutive bands. The image decoding device 150 can perform band-off filtering by acquiring information about the start time of the bands and the offset values ​​of each of the multiple bands in order to obtain information about the multiple bands.

[0088] In an embodiment, the image decoding device 150 can perform sample adaptive offset filtering on the current block by using information about sample adaptive offset filtering for neighboring blocks. For example, the image decoding device 150 can perform sample adaptive offset filtering on the current block by using information about sample adaptive offset filtering for the upper or left block.

[0089] In an embodiment, the adaptive loop filtering can be filtering applied to the current block by using a filter determined or obtained adaptively based on characteristics of the current block. The adaptive loop filtering can be filtering performed by using at least one of at least one pre-defined filter set, at least one APS filter set obtained from an adaptive parameter set (APS), and a filter determined based on a class of the current block.

[0090] In an embodiment, the adaptive loop filtering can be filtering determined based on information obtained from a bitstream and using one of a pre-defined filter set or at least one APS filter set obtained from an APS. Further, the filter to be used in the adaptive loop filtering can be determined to be one of a plurality of classes based on characteristics such as directionality and activity of samples within the current block. On the other hand, the characteristics such as directionality and activity of samples within the current block can be determined by using gradients of the current block.

[0091] On the other hand, the in-loop filtering unit 500 is not limited to the disclosed examples, and can additionally process an image in parallel with at least one of the deblocking filtering unit 510, the sample adaptive offset filtering unit 520, and the adaptive loop filtering unit 530, or additional filtering can be performed through a configuration added before or after at least one of the filtering units. For example, the in-loop filtering unit 500 can use an output image obtained from a filtering neural network in order to obtain an image to which in-loop filtering has been applied to the current image.

[0092] In an embodiment, the in-loop filtering unit 500 can include a filtering neural network 550 that additionally filters an image by using state information. For example, the filtering neural network 550 can additionally process a reconstructed image in parallel with the deblocking filtering unit.

[0093] In an embodiment, the filtering neural network 550 can be a neural network that performs filtering on a current image by using state information corresponding to each of a reconstructed image of the current image, at least one image included in the decoded picture buffer 540, and at least one image included in the state information buffer 560.

[0094] In an embodiment, the filtering neural network 550 can generate an output image of a current image based on obtaining information corresponding to the current image. For example, the filtering neural network 550 can obtain at least one of a reconstructed image of the current image, a prediction image of the current image, a first reference image of the current image, a second reference image of the current image, at least one quantization parameter for the current image, and partition information about the current image, as the information corresponding to the current image.

[0095] In an embodiment, the in-loop filtering unit 500 can obtain an image to which in-loop filtering has been applied by inputting information corresponding to the current image to the filtering neural network 550. For example, the filtering neural network 550 can generate a feature map corresponding to the current image based on receiving the information corresponding to the current image, and can generate an output image from the feature map corresponding to the current image. Also, the in-loop filtering unit 500 can obtain an image to which in-loop filtering has been applied by using the output image.

[0096] In an embodiment, the in-loop filtering unit 500 can determine or generate layer information to be used for at least one layer within the filtering neural network by using state information corresponding to at least one reference image included in the state information buffer 560. For example, the in-loop filtering unit 500 can determine or generate a kernel or an attention map to be used in the filtering neural network based on first state information corresponding to a first reference image and second state information corresponding to a second reference image stored in the state information buffer 560.

[0097] In an embodiment, the in-loop filtering unit 500 can input information corresponding to the current image to the filtering neural network 550 based on the layer information. The in-loop filtering unit 500 can obtain an output image of the current image from the filtering neural network 550 based on inputting the information corresponding to the current image to the filtering neural network 550.

[0098] In an embodiment, the image decoding apparatus 150 can perform an addition operation on the deblocking filtered image obtained from the deblocking filtering unit 510 and the output image obtained from the filtering neural network 550. The addition operation can refer to element-wise summation. The image decoding apparatus 150 can input a result of performing the addition operation on the deblocking filtered image and the output image to the sample adaptive offset filtering unit 520. The result of performing the addition operation on the deblocking filtered image and the output image can be stored in the decoded picture buffer 540 through the sample adaptive offset filtering unit 520 and the adaptive loop filtering unit 530. Also, the image stored in the decoded picture buffer 540 through the sample adaptive offset filtering unit 520 and the adaptive loop filtering unit 530 can be used as a reference image, and state information corresponding to the image stored in the decoded picture buffer 540 can be stored in the state information buffer 560.

[0099] On the other hand, the configuration of the filtering neural network 550 in parallel with the deblocking filter unit is merely an example, and the disclosure is not limited to the disclosed example. The filtering neural network 550 can be configured in parallel with a plurality of filter units or connected in series. Also, instead of simply performing an addition operation on the deblocking filtered image and the output image obtained from the filtering neural network 550, an image obtained by performing a weighted sum by assigning a pre-defined weight to each of the deblocking filtered image and the output image can be input as an input of the sample adaptive offset filter unit 520. The filtering neural network 550 can be appropriately trained according to a position where the filtering neural network 550 is placed.

[0100] Hereinafter, referring to Figure 12 The training of the filtering neural network 550 is described in detail.

[0101] Figure 6 is a diagram illustrating an in-loop filtering operation according to an embodiment of the disclosure.

[0102] In an embodiment, the in-loop filter unit 135 of the image encoding apparatus 110 or the in-loop filter unit 165 of the image decoding apparatus 150 (hereinafter, an in-loop filter unit 600) can include at least one of a deblocking filter unit 610, a sample adaptive offset filter unit 620, and an adaptive loop filter unit 630.

[0103] In an embodiment, the deblocking filter unit 610, the sample adaptive offset filter unit 620, and the adaptive loop filter unit 630 can correspond to the deblocking filter unit 510, the sample adaptive offset filter unit 520, and the adaptive loop filter unit 530 of FIG. 5, respectively, and thus the same description is omitted. Figure 5

[0104] In an embodiment, the in-loop filter unit 600 can additionally process an image in parallel with at least one of the deblocking filter unit 610, the sample adaptive offset filter unit 620, and the adaptive loop filter unit 630, or additional filtering can be performed through a configuration added before or after at least one of the filter units. For example, the in-loop filter unit 600 can include a filtering neural network that generates an output image used to obtain an in-loop filtered image.

[0105] In an embodiment, the filtering neural network 650 can be a neural network that performs filtering by using the reconstructed image, the deblocking filtered image obtained by processing the reconstructed image in the deblocking filter unit, at least one image included in the decoded picture buffer 640, and the state information included in the state information buffer 660.

[0106] ​In an embodiment, the filtering neural network 650 can generate an output image of the current picture based on obtaining information corresponding to the current picture. For example, the filtering neural network 650 can obtain at least one of the following as the information corresponding to the current picture: a reconstructed image of the current picture, a deblocking filtered image of the current picture, a prediction image of the current picture, a first reference image of the current picture, a second reference image of the current picture, at least one quantization parameter used for the current picture, and partition information about the current picture.

[0107] In an embodiment, the in-loop filtering unit 600 can generate or determine layer information to be used in the filtering neural network by using the state information corresponding to the reference picture included in the state information buffer 660. For example, the in-loop filtering unit 600 can generate or determine a kernel or an attention map to be used in the filtering neural network based on the state information stored in the state information buffer 660.

[0108] In an embodiment, the in-loop filtering unit 600 can input information corresponding to the current picture to the filtering neural network 650 based on the layer information. The in-loop filtering unit 600 can obtain an output image of the current picture from the filtering neural network 650 based on inputting the information corresponding to the current picture to the filtering neural network 650.

[0109] In an embodiment, the in-loop filtering unit 600 can input the output image obtained from the filtering neural network 650 to the sample adaptive offset filtering unit 620. Through the sample adaptive offset filtering unit 620 and the adaptive loop filtering unit 630, the output image is an image to which in-loop filtering has been applied, and can be stored in the decoded picture buffer 640.

[0110] On the other hand, the filtering neural network 650 is configured to be connected in series after deblocking filtering, which is only an example, but the disclosure is not limited to the disclosed example. For example, the filtering neural network 650 can be placed between the sample adaptive offset filtering unit 620 and the adaptive loop filtering unit 630, and can be appropriately trained according to the location of placement.

[0111] Figure 7 is a diagram for describing a process of generating an output image from a filtering neural network according to an embodiment of the disclosure.

[0112] In an embodiment, the in-loop filtering unit can include a filtering neural network 720. The filtering neural network 720 can include a feature extraction neural network 721 and a regression neural network 726. In addition, the filtering neural network 720 can be a neural network that generates an output image used to obtain an in-loop filtered image.

[0113] In an embodiment, when the first reference picture or the second reference picture is used to decode the current picture, the filtering neural network 720 can use the first state information corresponding to the first reference picture or the second state information corresponding to the second reference picture. For example, the filtering neural network 720 can use only the first reference picture to decode the current picture, and in this case, only the first state information can be used.

[0114] On the other hand, the first state information corresponding to the first reference picture can be obtained from a feature map corresponding to the first reference picture. On the other hand, the second state information corresponding to the second reference picture can be obtained from a feature map corresponding to the second reference picture.

[0115] In an embodiment, the first reference picture of the current picture or the second reference picture of the current picture can be a picture included in the first reference picture list or the second reference picture list. For example, the first reference picture can be included in the first reference picture list, and the second reference picture can be included in the second reference picture list. Alternatively, the first reference picture and the second reference picture can be pictures included in the second reference picture list. Alternatively, the first reference picture and the second reference picture can be pictures included in the first reference picture list. However, the disclosure is not limited to the disclosed examples.

[0116] In an embodiment, the image decoding apparatus can determine the filtering neural network 720 used to filter the current picture by using the first state information included in the first reference state list 742 and corresponding to the first reference picture included in the first reference picture list. For example, the image decoding apparatus can input the first state information and the second state information to the layer information neural network 750. The image decoding apparatus can generate layer information to be used in at least one neural network within the filtering neural network 720 based on inputting the first state information and the second state information to the layer information neural network 750.

[0117] On the other hand, for convenience of explanation, an example in which only the first reference image or the second reference image is used as a reference image of the current image has been described, but the disclosure is not limited to the disclosed example. State information corresponding to a reference block included in at least one reference image to be used for predicting the current image can be input to the layer information neural network, and layer information for the current image can be generated by using at least one piece of input state information. For example, for a first current block and a second current block included in the current image, when the first current block is predicted by using a first reference block included in the first reference image and the second current block is predicted by using a second reference block included in the second reference image and a third reference block included in a third reference image, the current image can generate layer information for the current image based on inputting state information corresponding to the first reference block, state information corresponding to the second reference block, and state information corresponding to the third reference block to the layer information neural network 750.

[0118] In an embodiment, the image decoding apparatus can generate or obtain layer information 755 used in at least one layer within the filtering neural network 720 based on first state information obtained from a feature map corresponding to a first reference image of the current image or second state information obtained from a feature map corresponding to a second reference image of the current image.

[0119] In an embodiment, the filtering neural network can include a feature extraction neural network that extracts a feature of the current image from information corresponding to the current image and generates a feature map corresponding to the current image, and a regression neural network used to obtain an image to which in-loop filtering has been applied from the feature map corresponding to the current image. The layer information 755 can include at least one of a kernel and an attention map to be used in at least one layer included in the feature extraction neural network 721 and the regression neural network 726.

[0120] In an embodiment, the filtering neural network can include at least one of a deep neural network, a convolutional neural network, and a transformer neural network.

[0121] In an embodiment, the filtering neural network 720 can generate the output image 760 of the current image based on obtaining information corresponding to the current image. For example, the filtering neural network 720 can obtain at least one of the following as the information corresponding to the current image 710: a reconstructed image of the current image, a deblocking filtered image of the current image, a prediction image of the current image, a first reference image of the current image, a second reference image of the current image, at least one quantization parameter for the current image, and partition information about the current image. On the other hand, the present disclosure is not limited to the disclosed examples, and at least one of a sample adaptive offset filtered image and an adaptive loop filtered image of the current image can be obtained as the information corresponding to the current image 710 according to a structure of an in-loop filter.

[0122] In an embodiment, the image decoding apparatus can input the information corresponding to the current image 710 to the filtering neural network 720. The image decoding apparatus can input the information corresponding to the current image to at least one layer within the filtering neural network. In the image decoding apparatus, the feature extraction neural network 721 included in the filtering neural network 720 can obtain, generate, or extract a feature map 725 corresponding to the current image by obtaining the information corresponding to the current image 710. The regression neural network 726 included in the filtering neural network 720 can obtain the feature map 725 from the feature extraction neural network 721 as input. The image decoding apparatus can generate or obtain the output image 760 based on inputting the feature map 725 to the regression neural network 726.

[0123] In an embodiment, the image decoding apparatus can obtain the intermediate feature map 725 corresponding to the current image from the filtering neural network 720, and can input the intermediate feature map 725 to the state encoding neural network. For example, the image decoding apparatus can input the intermediate feature map 725 obtained or extracted from the feature extraction neural network 721 included in the filtering neural network 720 to the state encoding neural network 730.

[0124] In an embodiment, the image decoding apparatus can generate or obtain the state information 735 of the current image from the state encoding neural network 730 by inputting the feature map 725 obtained from the filtering neural network 720 to the state encoding neural network 730.

[0125] In an embodiment, the state information 735 of the current image can be information obtained from a feature map corresponding to the current image. Because the state information 735 of the current image includes abstract information rather than high-resolution information, it can require less memory capacity than information about images stored in a decoded picture buffer. In addition, because the state information 735 is an output of a hidden layer trained by optimizing a loss function using a gradient descent method, the state information 735 can be abstract information that is not clear about what the information exactly represents.

[0126] In an embodiment, the state information 735 of the current picture can have a lower resolution than a resolution of the current picture depending on characteristics of the abstract information representing the current picture.

[0127] In an embodiment, the image decoding apparatus can store the generated or obtained state information 735 in a state information buffer 740. The image decoding apparatus can also use the state information 735 of the current picture stored in the state information buffer 740 when the current picture is used as a reference picture.

[0128] On the other hand, in order to facilitate the explanation of the disclosure, the process of obtaining an output picture of the current picture by inputting information corresponding to the current picture to the filtering neural network and obtaining current state information corresponding to the current picture has been mainly described, but the disclosure is not limited to the disclosed examples, and the same operation can be performed by inputting information corresponding to the first reference picture or the second reference picture to the filtering neural network.

[0129] Hereinafter, referring to Figures 8a to 8c The layer information applied to the filtering neural network is described in detail. In addition, referring to Figures 9 to 11b The neural network according to the disclosure is described in detail.

[0130] Figures 8a to 8c is a diagram for describing a filtering neural network using layer information according to an embodiment of the disclosure.

[0131] In an embodiment, the image decoding apparatus can configure the filtering neural network based on the layer information. For example, the image decoding apparatus can generate or obtain an attention map as the layer information by using the layer information neural network. The attention map can be at least one of a channel attention map 832, a spatial attention map 834, and a full attention map 836. The channel attention map 832 can be a one-dimensional attention map, the spatial attention map 834 can be a two-dimensional attention map, and the full attention map 836 can be a three-dimensional attention map.

[0132] Referring to Figure 8a The feature map can be output by using the channel attention map 832 generated as the layer information included in at least one neural layer 810 in the filtering neural network. Hereinafter, the output of the neural layer 810 can be referred to as an output feature map.

[0133] In an embodiment, the neural layer 810 included in the filtering neural network can be a layer that performs an attention operation by using the channel attention map 832 obtained, generated, or output by the layer information neural network. The neural layer 810 can be a simple diagram indicating that the channel attention map 832 is applied to a feature map that has only performed a convolution operation.

[0134] In an embodiment, the neural layer 810 can perform a convolution operation and an attention operation. For example, the image decoding device can perform a convolution operation by applying weights and biases to an input feature map using a convolution kernel through the neural layer 810, and can perform an attention operation by applying element-wise multiplication with an attention map to a result of performing the convolution operation. Through the performance of the convolution operation and the attention operation, an output feature map can be generated or output through one neural layer 810. This can be expressed as Equation 1 below.

[0135] [Equation 1] Output feature map = activation(weight input feature map + bias) x (attention map) In an embodiment, the activation can denote an activation function and can be one of a sigmoid function, a rectified linear unit function (ReLu function), a hyperbolic tangent function (Tanh function), and a softmax function according to the design of a neural network, but the disclosure is not limited to the disclosed examples, and other activation functions can be used as technology advances.

[0136] In an embodiment, for the result 820 of performing a convolution operation on an input feature map, the channel attention map 832 can be a map that is used to emphasize a portion to be focused or concentrated in one-dimensional directions among a height direction, a width direction, and a channel direction. Furthermore, when the channel attention map 832 has a different size from the result 820 of performing a convolution operation on an input feature map in a certain dimension, the channel attention map 832 can be applied after being adjusted to the same or similar size as the result of performing a convolution operation on an input feature map. For example, when the result 820 of performing a convolution operation on an input feature map is expressed in the size of C W H and the channel attention map 832 is c' 1 1, the channel attention map 832 can be adjusted to the size of C 1 1 and applied to all W H channels.

[0137] Referring to Figure 8b , an output feature map can be output by using the spatial attention map 834 generated as layer information included in at least one neural layer 810 in the filtering neural network.

[0138] In an embodiment, the neural layer 810 included in the filtering neural network can be a layer that performs an attention operation by using the spatial attention map 834 obtained, generated, or output through the layer information neural network. Hereinafter, because Figure 8b the neural layer 810 is a layer that performs an attention operation by using the spatial attention map 834, it can be referred to as an attention layer. Figure 8aThe only difference in neural layer 810 is that spatial attention map 834 is used as attention map, so its same redundant description is omitted.

[0139] In an embodiment, for the result 820 of performing a convolution operation on the input feature map, the spatial attention map 834 can be a map used to emphasize portions that will be focused or concentrated in two-dimensional directions in the height, width, and channel directions. Furthermore, when the spatial attention map 834 has a different size in a specific dimension than the result 820 of performing a convolution operation on the input feature map, the spatial attention map 834 can be applied after being adjusted to the same or similar size as the result 820 of performing a convolution operation on the input feature map. For example, when the result 820 of performing a convolution operation on the input feature map has a C... W The size of H is represented by the spatial attention diagram 834, which is 1. w' At time h', the spatial attention map 834 can be adjusted to 1. W The size of H and it is applied to a size with W All C spaces of size H.

[0140] Reference Figure 8c The feature map can be output using a full attention map 836 generated as layer information included in at least one neural layer 810 in the filtered neural network.

[0141] In an embodiment, the neural layer 810 included in the filtering neural network may be a layer that performs attention operations by using a full attention map 836 obtained, generated, or output by a layer information neural network. In the following, because... Figure 8c neural layer 810 and Figure 8a The only difference in neural layer 810 is that the full attention map 836 is used as the attention map, so its same redundant description is omitted.

[0142] In an embodiment, for the result 820 of performing a convolution operation on the input feature map, the full attention map 836 can be a map used to emphasize portions that will be focused or concentrated in three dimensions: height, width, and channel. Furthermore, when the full attention map 836 has a different size in a specific dimension than the result 820 of performing a convolution operation on the input feature map, the full attention map 836 can be applied after being adjusted to the same or similar size as the result 820 of performing a convolution operation on the input feature map. For example, when the intermediate feature map as the result 820 of performing a convolution operation on the input feature map is C... W The dimension of H is represented and full attention figure 836 is c' w' h' time, the full attention map 836 can be adjusted to C W H size, and for each position of the result 820 of performing a convolution operation on the input feature map, the output feature map can be obtained or generated by performing an element-wise multiplication using a corresponding sample of the full attention map 836.

[0143] In an embodiment, the full attention map 836 can be generated in the same or similar size as the intermediate feature map for performing the element-wise multiplication.

[0144] On the other hand, the attention map is not limited to Figures 8a to 8c the examples disclosed in the disclosure, and other methods can be applied.

[0145] Figure 9 is a diagram illustrating a structure of a feature extraction neural network included in a filtering neural network according to an embodiment of the disclosure.

[0146] Referring to Figure 9 , the feature extraction neural network can obtain information 910 corresponding to a current image as input, wherein the information 910 corresponding to the current image includes at least one of a reconstructed image, a deblocking filtered image, a prediction image, a first reference image, a second reference image, at least one quantization parameter, and partition information. In addition, the feature extraction neural network can output an intermediate feature map 930 based on obtaining the information 910 corresponding to the current image.

[0147] In an embodiment, the feature extraction neural network can be a deep neural network, a convolutional neural network, and a transformer neural network, or can be a combination of a deep neural network, a convolutional neural network, and a transformer neural network. On the other hand, although Figure 9 an example in which the feature extraction neural network is a convolutional neural network is illustrated, the disclosure is not limited thereto, and other neural network models can be used.

[0148] In an embodiment, the feature extraction neural network can include at least one neural layer 920. On the other hand, for convenience of explanation, the at least one neural layer 920 included in the feature extraction neural network is illustrated as including only a convolution layer, but the disclosure is not limited to the disclosed example, and the at least one neural layer 920 can be configured to include at least one of a convolution layer, an activation layer, an attention layer, a pooling layer, a dropout layer, a batch normalization layer, a recurrent layer, and an embedding layer. On the other hand, the at least one neural layer 920 can be referred to as at least one layer.

[0149] In an embodiment, the feature extraction neural network can be a neural network including a structure of a combination of a repeated convolution layer, an activation layer, and an attention layer.

[0150] In an embodiment, the at least one layer included in the feature extraction neural network can include at least one neural layer for generating a feature map corresponding to the current image. The feature map corresponding to the current image can have the same resolution as the resolution of the current image. Alternatively, the feature map corresponding to the current image can have a higher or lower resolution than the resolution of the current image.

[0151] In an embodiment, the at least one layer included in the feature extraction neural network can include at least one down-sampling layer. The feature extraction neural network can include at least one down-sampling layer for generating a feature map corresponding to the current image, wherein the feature map has a lower resolution than the resolution of the current image.

[0152] In an embodiment, the image decoding apparatus can obtain a feature map corresponding to the current image based on inputting information corresponding to the current image to the feature extraction neural network. In an embodiment, the image decoding apparatus can obtain a feature map corresponding to the current image in response to inputting information corresponding to the current image to the feature extraction neural network. Further, state information of the current image can be obtained from the feature map corresponding to the current image, and can have a lower resolution than the resolution of the current image. This can be based on a down-sampling layer included in the feature extraction neural network.

[0153] On the other hand, the size of the intermediate feature map generated as an output of the feature extraction neural network can be expressed as c w h, where w and h can be different from W and H, respectively, which are the sizes of the image.

[0154] Figure 10 is a diagram illustrating a structure of a regression neural network included in a filtering neural network according to an embodiment of the disclosure.

[0155] Referring to Figure 10 , the regression neural network can obtain the feature map 1010 obtained from the feature extraction neural network as input. Further, the regression neural network can generate an output image 1030 to be used to obtain an in-loop filtered image based on the obtaining of the feature map 1010.

[0156] In an embodiment, the regression neural network can be a deep neural network, a convolutional neural network, and a transformer neural network, or can be a combination of a deep neural network, a convolutional neural network, a recurrent neural layer, and a transformer neural network. On the other hand, although Figure 10 An example in which the regression neural network is a convolutional neural network is illustrated, but the disclosure is not limited thereto, and other neural network models can be used.

[0157] In an embodiment, the regressive neural network may include at least one neural layer 1020. Alternatively, for ease of illustration, the at least one neural layer 1020 included in the regressive neural network is shown as comprising only convolutional layers; however, this disclosure is not limited to the disclosed example, and the at least one neural layer 1020 may be configured to include at least one of convolutional layers, activation layers, attention layers, pooling layers, dropout layers, batch normalization layers, recursive layers, and embedding layers. Alternatively, the at least one neural layer 1020 may be referred to as at least one layer.

[0158] In an embodiment, the regressive neural network may be a neural network with a structure that includes a combination of repeated convolutional layers, activation layers, and attention layers.

[0159] In an embodiment, at least one layer included in the regressive neural network can receive a feature map corresponding to the current image as input and generate an output image. The resolution of the feature map corresponding to the current image can be equal to the resolution of the output image. The resolution of the feature map corresponding to the current image can be lower or higher than the resolution of the output image.

[0160] In an embodiment, at least one layer included in a regression neural network may include at least one upsampling layer. The regression neural network may include at least one upsampling layer for generating an output image with a higher resolution than the feature map corresponding to the current image. For example, when a feature extraction neural network includes at least one downsampling layer for generating a feature map corresponding to the current image with a lower resolution than the current image, the regression neural network may include at least one upsampling layer for generating an output image with a higher resolution than the feature map corresponding to the current image.

[0161] In one embodiment, the regressive neural network can generate an output image based on obtaining a feature map corresponding to the current image. This output image will be used to obtain an image in which in-loop filtering has been applied. The image decoding device can obtain the in-loop filtered image of the current image by performing an addition operation on the output image obtained from the regressive neural network and the deblocked filtered image of the current image.

[0162] In one embodiment, the image decoding device may obtain an output image from a regression neural network that takes a feature map corresponding to the current image as input. Furthermore, the output image may have a higher resolution than the feature map corresponding to the current image.

[0163] On the other hand, the size of the output image 1030 generated as the output of the regression neural network can be represented as C. W H, where W and H represent the dimensions of the image, and c represents information about the color channels. For example, a color channel can represent information about each of the R, G, and B channels (RGB channels) or the Y, U, and V channels (YUV channels).

[0164] In an embodiment, a state-coded neural network may be similar to... Figure 10 Regression neural networks. For example, state-encoding neural networks can obtain regression neural networks. Figure 10 The same input is used for a regressive neural network. However, the output of a state-encoding neural network can be state information about the image. However, this disclosure is not limited to the disclosed examples.

[0165] In embodiments, the state-encoded neural network may include at least one of a deep neural network, a convolutional neural network, and a transformer neural network. To reduce the size of the data by taking into account the abstract nature of state information, the state-encoded neural network may include downsampling layers (such as strided convolutional layers and pooling layers).

[0166] Figures 11a to 11c This is a diagram illustrating the structure of a layered information neural network according to an embodiment of the present disclosure.

[0167] Reference Figures 11a to 11c The layer information neural network can obtain the state information corresponding to the image stored in the state information buffer as input. For example, layer information neural networks 1120, 1122, and 1124 can obtain the state information corresponding to the reference image of the current image in order to generate or output the layer information 1130 of the current image. In the following, the reference image may be the first reference image and / or the second reference image of this disclosure.

[0168] In an embodiment, the image decoding apparatus can generate or output layer information by inputting multiple pieces of state information corresponding to a reference block into layer information neural networks 1120, 1122, and 1124, wherein the reference block is used to generate a prediction block that includes the block in the current image. Therefore, the image decoding apparatus can generate or output layer information using all or part of at least one piece of state information corresponding to at least one reference image used to perform prediction on the current image. Alternatively, the layer information can be generated as frame units, stripe units, parallel block units, maximum CUs, CUs, or prediction units, but this disclosure is not limited to the disclosed examples.

[0169] However, although examples have been described using state information corresponding to the current block used for prediction of the current block of the prediction unit, this disclosure is not limited to the disclosed examples, and layer information can be generated, determined, or output at the frame unit level. For example, when generating layer information at the frame unit level, the layer information can be generated by using state information corresponding to each reference block included in the frame.

[0170] In an embodiment, when inter-frame prediction is performed via unidirectional reference (such as forward reference or backward reference) to generate a predicted block for the current block included in the current image, the image decoding device can obtain state information in the first reference state list 1112 corresponding to a reference block in a reference image included in the first reference image list, or state information in the second reference state list 1114 corresponding to a reference block in a reference image included in the second reference image list, as input to the layer information neural networks 1120, 1122 and 1124.

[0171] In an embodiment, when performing bidirectional prediction (such as bidirectional reference, 2 forward reference, or 2 backward reference) to generate a predicted block of the current block included in the current image, the image decoding device may generate or output layer information based on inputting multiple state information corresponding to the two reference blocks into the layer information neural networks 1120, 1122, and 1124, wherein the two reference blocks are respectively included in two reference images in the first reference image list and the second reference image list.

[0172] In an embodiment, at least one image included in the first reference image list for predicting the current image may include a first reference image, and the state information corresponding to the first reference image may be referred to as first state information. Furthermore, at least one image included in the second reference image list for predicting the current image may include a second reference image, and the state information corresponding to the second reference image may be referred to as second state information. Additionally, the first reference state list may be a list including multiple pieces of state information corresponding to images included in the first reference image list. The second reference state list may be a list including multiple pieces of state information corresponding to images included in the second reference image list.

[0173] In an embodiment, the layer information neural networks 1120, 1122 and 1124 may be based on multiple state information corresponding to the image stored in the state information buffer as input, and the output will be applied to the layer information 1130 of at least one layer in the filter neural network.

[0174] In the embodiments, reference is made to Figure 11a A layered information neural network can output information with c 1 An attention map or kernel of size 1. For example, a layer information neural network 1120 may include at least one downsampling layer. The at least one downsampling layer may include at least one strided convolutional layer, and the layer information neural network may output a channel attention map or kernel.

[0175] In the embodiments, reference is made to Figure 11bThe layered information neural network 1122 can output information with c 1 A size 1 attention map or kernel. For example, a layer information neural network may include at least one downsampling layer. At least one downsampling layer may include at least one pooling layer. According to... Figure 11b The layered information neural network can output channel attention maps or kernels.

[0176] In the embodiments, reference is made to Figure 11c The layered information neural network 1124 can output c w An attention map of size h. For example, a layer information neural network may include at least one upsampling layer. The at least one upsampling layer may include at least one of an interpolation layer, a deconvolution layer, and a pixel shuffle layer. The layer information neural network may output a spatial attention map or a full attention map.

[0177] In an embodiment, layer information 1130 may include at least one of a convolutional kernel and an attention map. For example, the layer information neural network may output only a convolutional kernel, only an attention map, or both a convolutional kernel and an attention map. Furthermore, layer information 1130 may include layer information 1130 to be applied to each of the feature extraction neural network and the regression neural network included in the filtering neural network, and the layer information 1130 to be applied to the feature extraction neural network may be different from the layer information 1130 to be applied to the regression neural network.

[0178] In embodiments, the layer information neural network can be a deep neural network, a convolutional neural network, and a transformer neural network, or it can be a combination of a deep neural network, a convolutional neural network, a recurrent neural layer, and a transformer neural network. On the other hand, although Figure 11a The layer information neural network shown is an example of a convolutional neural network, but this disclosure is not limited thereto, and other neural network models may be used.

[0179] In an embodiment, the layer information neural network may include at least one neural layer 1120, 1122, and 1124. On the other hand, for ease of illustration, at least one neural layer 1120 included in the layer information neural network is shown as including only stride convolutional layers, but this disclosure is not limited to the disclosed example, and at least one neural layer 1120 may be configured to include at least one of convolutional layers, activation layers, attention layers, pooling layers, dropout layers, batch normalization layers, recursive layers, and embedding layers.

[0180] For example, refer to Figures 11a to 11cA layer information neural network can be a neural network with a structure that includes a combination of repeated convolutional layers and activation layers, a neural network with a structure that includes a combination of repeated convolutional layers and downsampling layers, or a neural network with a structure that includes a combination of repeated convolutional layers and upsampling layers.

[0181] In an embodiment, the layer information neural network may include at least one of stride convolutional layers, pooling layers, interpolation layers, deconvolutional layers, and pixel rearrangement layers.

[0182] For example, considering the characteristics of having relatively small kernel sizes to perform scanning of input state information, layer information neural networks can be configured using downsampling layers (such as strided convolutional layers and pooling layers). Additionally, kernels can be generated using downsampling layers in layer information neural networks.

[0183] For example, refer to Figures 11a to 11c A layered information neural network may include performing downsampling to generate a layer with c 1 A layer of channel attention map of size 1, and may include performing upsampling to generate a channel attention map with c w The full attention diagram 836 has a size of h. Furthermore, the layer information neural networks 1120, 1122, and 1124 can appropriately adjust the size of the output layer information 1130 by including downsampling layers or upsampling layers. On the other hand, the layer performing downsampling may include at least one of strided convolutional layers and pooling layers, and the layer performing upsampling may include at least one of interpolation layers, deconvolutional layers, and pixel rearrangement layers.

[0184] On the other hand, strided convolutional layers are a type of convolutional layer where, when performing convolution operations, a stride greater than 1 can be used. The stride size represents the interval at which the kernel moves the input feature map. Therefore, the size of the output feature map can be adjusted according to the stride size and can be used to reduce the dimensionality of the input data. Furthermore, because the use of strided convolutional layers reduces computation, the training and inference speed of layer-based neural networks can be improved.

[0185] On the other hand, pooling layers can be used to reduce the dimensionality of the input while preserving important information (e.g., by selecting and outputting the maximum value within the pooling window to emphasize the most prominent features of the input, or by calculating and outputting the average of all values ​​within the pooling window). Furthermore, pooling layers can be used to reduce the computational complexity of the network and prevent overfitting by excluding learnable weights.

[0186] On the other hand, interpolation layers are layers used to generate outputs with increased dimensions compared to the input, and can be layers that scale and increase the size of the data obtained as input.

[0187] On the other hand, a deconvolution layer is a layer that expands the dimensions of an image by performing the inverse process of convolution, and can be a layer with learnable parameters.

[0188] On the other hand, a pixel rearrangement layer is a layer that improves resolution by rearranging the pixels in the input, and can also be a layer that increases the number of channels in the input and enlarges the size of the image by rearranging the channels at the pixel level.

[0189] On the other hand, this disclosure is not limited to the disclosed examples, and has x y A kernel of size z (where x, y, and z are natural numbers) can be generated as the output of a layer information neural network, and has c w Attention maps of size h (where c, w, and h are natural numbers) can be generated as the output of a layer information neural network.

[0190] Figure 12 This is a diagram illustrating a method for training a filtered neural network according to an embodiment of the present disclosure.

[0191] In an embodiment, the filtering neural network 1220 may be trained based on information obtained about at least two images with a reference relationship, such that the in-loop filtered image obtained by the in-loop filtering unit becomes closer to the original image.

[0192] In an embodiment, the filtering neural network 1220 can output a training output image 1230 by taking information 1210 corresponding to at least two images with a reference relationship as input. The filtering neural network 1220 can be trained by comparing the training output image 1230 with a ground truth output image 1240 through a loss function 1250 to make the training output image 1230 closer to the ground truth output image 1240. On the other hand, the output image that makes the in-loop filtered image closer to the original image according to the structure of the in-loop filtering unit including the filtering neural network 1220 can be previously determined as the ground truth output image 1240.

[0193] In an embodiment, the filtering neural network, the state coding neural network, and the layer information neural network can be trained together by inputting information corresponding to the first reference image or the second reference image into the filtering neural network and inputting information corresponding to the current image into the filtering neural network.

[0194] For example, the filtering neural network 1220 can receive information corresponding to a first reference image and information 1210 corresponding to the current image. The filtering neural network 1220 can generate a feature map corresponding to the first reference image using the information corresponding to the first reference image. Furthermore, the encoding neural network can receive the feature map corresponding to the first reference image as input and can generate first state information for the first reference image. The filtering neural network 1220 can generate a feature map corresponding to a second reference image using the information corresponding to a second reference image. Furthermore, the encoding neural network can receive the feature map corresponding to the second reference image as input and can generate second state information for the second reference image. Additionally, the layer information neural network can generate layer information to be applied to the filtering neural network 1220 using either the first state information corresponding to the first reference image or the second state information corresponding to the second reference image. An output image as a training output image 1230 can be generated by inputting the information corresponding to the current image into the filtering neural network 1220, which includes at least one layer to which the generated layer information has been applied. The filtering neural network 1220 can be trained by comparing the training output image 1230 with the benchmark real output image 1240 so that the training output image 1230 is closer to the benchmark real output image 1240.

[0195] On the other hand, for ease of explanation, the filter neural network 1220 is represented and shown as being trained, but the feature extraction neural network, regression neural network, state encoding neural network, and layer information neural network included in the filter neural network of this disclosure can be trained as a whole. Alternatively, the neural network can be trained individually for each preset benchmark true value among the feature extraction neural network, regression neural network, state encoding neural network, and layer information neural network.

[0196] Figure 13 This is a block diagram illustrating the configuration of an image decoding apparatus according to an embodiment of the present disclosure.

[0197] Reference Figure 13 The image decoding device 1300 may include an acquisition unit 1310 and a prediction decoding unit 1330. Figure 13 The obtaining unit 1310 shown in the figure can be connected with Figure 1 The entropy decoding unit 155 shown corresponds to the one described above, and the prediction decoding unit 1330 can be associated with... Figure 1 The in-loop filtering unit 165 and the prediction decoding unit 170 shown in the figure correspond to each other.

[0198] In an embodiment, the acquisition unit 1310 and the prediction decoding unit 1330 may be implemented as at least one processor. In an embodiment, the acquisition unit 1310 and the prediction decoding unit 1330 may operate according to instructions stored in a memory.

[0199] In one embodiment, the image decoding apparatus 1300 may include a memory for storing the input and output data of the acquisition unit 1310 and the prediction decoding unit 1330. Furthermore, the image decoding apparatus 1300 may also include a memory control unit for controlling the data input and output of the memory.

[0200] In an embodiment, at least one processor is configured to control a series of processes such that the image decoding device 1300 operates according to the embodiments described below, and may be implemented as one or more processors. The one or more processors included in the processor may be circuits, such as a SoC or IC. The one or more processors included in the processor may be general-purpose processors (such as a central processing unit (CPU), microprocessor unit (MPU), AP, or digital signal processor (DSP)), dedicated graphics processors (such as a GPU or vision processing unit (VPU)), dedicated artificial intelligence processors (such as an NPU), or dedicated communication processors (such as a CP). When the one or more processors included in the processor 140 are dedicated artificial intelligence processors, the dedicated artificial intelligence processor may be designed with a hardware architecture specifically designed to process a particular artificial intelligence model.

[0201] In embodiments, the processor can write data to or read data stored in memory. Specifically, the processor can execute at least one instruction or program stored in memory to process data according to predefined operating rules or an artificial intelligence model. Therefore, the processor can perform the operations described in the embodiments of this disclosure, and unless otherwise stated, the embodiments of this disclosure are described as being performed by the image decoding device 1300 or detailed elements included in the image decoding device 1300 (see [link to documentation]). Figure 13 The operations performed by 1310 and 1330 can be considered to be performed by the processor.

[0202] In embodiments, the memory may be configured to store various programs or data, and may be configured as a storage medium such as read-only memory (ROM), random access memory (RAM), hard disk, optical disc read-only memory (CD-ROM), and digital versatile optical disc (DVD) or a combination of storage media. The memory may be configured to be included within the processor, rather than existing separately. The memory may include volatile memory, non-volatile memory, or a combination of volatile and non-volatile memory. The memory may store a program or at least one instruction for performing operations according to the embodiments described below. The memory may provide stored data to the processor in response to a request from the processor.

[0203] In an embodiment, the obtaining unit 1310 may obtain a bitstream generated as a result of encoding the image.

[0204] In an embodiment, the bitstream may include the result of encoding the current block. The bitstream may include multiple pieces of information used to reconstruct the current block. The current block may be the largest CU, CU, transform unit, prediction unit, or filtering unit divided from the current image to be decoded. Furthermore, the current block may be a block at a predefined position being processed in the currently being performed encoding or decoding operation. The current sample may be any sample included in the current block.

[0205] In an embodiment, the prediction decoding unit 1330 can generate a prediction image to reconstruct the current image, and can generate or determine the prediction image of the current block based on prediction information included in the bitstream corresponding to at least one of the sequence parameter set, picture parameter set, video parameter set, strip header, and strip segment header. The prediction information can be information used to predict the current image.

[0206] In an embodiment, the obtaining unit 1310 can receive a bitstream from an image encoding device via a network.

[0207] In an embodiment, the obtaining unit 1310 may obtain a bit stream from a data storage medium, including magnetic media (such as hard disks, floppy disks, and magnetic tapes), optical recording media (such as CD-ROMs and DVDs), magneto-optical media (such as optical floppy disks), etc.

[0208] In this embodiment, the obtaining unit 1310 can obtain syntax elements for decoding the image from the bitstream. The values ​​corresponding to the syntax elements can be included in the bitstream according to the hierarchical structure of the image. The obtaining unit 1310 can obtain the binary bits corresponding to the syntax elements by performing entropy decoding on the bitstream.

[0209] In embodiments, the bitstream may include prediction information used to predict the current image. For example, the bitstream may include information indicating the prediction mode of the current block within the current image. The prediction mode of the current block may include an inter-frame mode. An inter-frame mode is a mode that predicts or reconstructs the current block based on a reference image to reduce temporal redundancy between images. Furthermore, when the prediction mode of the current block is an inter-frame mode, the prediction information may include information for determining a reference block. For example, the information for determining a reference block may include the index and motion information of a reference image within a list of reference frames, but this disclosure is not limited to the disclosed examples.

[0210] In an embodiment, the prediction decoding unit 1330 can generate a prediction block for the current block by performing inter-frame prediction on the current block according to the prediction mode of the current block, and can reconstruct the current block by using the prediction block.

[0211] In an embodiment, when the prediction decoding unit 1330 reconstructs the current block based on a reference image, the prediction decoding unit 1330 may use one reference image (e.g., one-way prediction) or two reference images (e.g., two-way prediction). Whether the current block is predicted one-way or two-way may be determined based on explicit information included in the bitstream, or may be implicitly determined from the prediction patterns of neighboring blocks associated with the current block.

[0212] In an embodiment, when the current block is bidirectionally predicted, the prediction decoding unit 1330 may use motion information included in the bitstream to determine a reference block for the current block, or may perform implicit determination from reference blocks of neighboring blocks associated with the current block. The motion information for the current block may include at least one of a reference image index, a motion vector, a differential motion vector, and a reference direction, and may include all multiple pieces of information used to predict the motion vector of the current block.

[0213] In an embodiment, the prediction decoding unit 1330 can make the current image closer to the original image by performing in-loop filtering on the reconstructed image of the current image.

[0214] In this embodiment, the prediction decoding unit 1330 can obtain filter information from the sequence parameter set of the bitstream, the image parameter set, the strip header, or the strip data. The filter information may include whether to perform deblocking filtering, sample offset adaptive filtering, and adaptive loop filtering, as well as the filter parameters used to perform each filter.

[0215] In one embodiment, the prediction decoding unit 1330 may include an in-loop filtering unit. The in-loop filtering unit may include a filtering neural network comprising a feature extraction neural network and a regression neural network. Furthermore, the in-loop filtering unit may use a state-encoding neural network and a layer information neural network to generate layer information applied to at least one layer within the filtering neural network.

[0216] In one embodiment, the prediction decoding unit 1330 can input information corresponding to the current image into a filtering neural network to generate an output image used to obtain an in-loop filtered image of the current image. The prediction decoding unit 1330 can obtain the output image from the filtering neural network, which is used to obtain an image with the in-loop filter applied to the current image.

[0217] In an embodiment, the prediction decoding unit can use state information corresponding to a reference block for each block in the current image to obtain an output image. For example, state information of a reference image can be determined using reference blocks corresponding to all blocks included in the current image; this state information is obtained as input to a layer information neural network to generate layer information for a filtering neural network to be applied to the current image. On the other hand, because references have already been made... Figure 7 and Figures 11a to 11cThe operation of the prediction decoding unit to generate layer information using a layer information neural network is described in detail, so the same redundant descriptions are omitted.

[0218] For example, when the current block is being predicted bidirectionally, the prediction decoding unit can perform a weighted sum by applying weights to a first reference block and a second reference block, respectively, and generate a predicted block for the current block. The state information corresponding to each of the reference blocks included in the current image can be obtained by applying weights to a weighted sum of multiple state information points corresponding to the two reference blocks, respectively, for bidirectional prediction.

[0219] On the other hand, for ease of explanation, a method for using reference block state information for blocks of a prediction unit has been described. However, this disclosure is not limited to the disclosed examples, and the prediction decoding unit can generate layer information by using state information corresponding to the reference block used for each of the prediction frame units, strip units, parallel block units, maximum CU, CU, and prediction units. For example, to generate layer information at the frame unit level, the prediction decoding unit can generate layer information by inputting state information corresponding to the reference block of the block included in the current image into a layer information neural network.

[0220] In an embodiment, the prediction decoding unit can obtain layer information applied to the filtering neural network from the layer information neural network. The prediction decoding unit can obtain at least one of a convolutional kernel and an attention map as layer information applied to the filtering neural network. The prediction decoding unit can use the obtained layer information in the filtering neural network. For example, the prediction decoding unit can be configured to include a convolutional layer in which the filtering neural network uses the convolutional kernel and attention map included in the layer information. The convolutional kernel and attention map of at least one convolutional layer included in the filtering neural network can be determined based on the obtained layer information. On the other hand, because it has already been referenced... Figure 7 and Figures 8a to 8c The operation of configuring the filtering neural network in the prediction decoding device by using layer information is described in detail, so the same redundant descriptions are omitted.

[0221] In an embodiment, the prediction decoding unit may input information corresponding to the current image into the filtering neural network. The information corresponding to the current image includes at least one of the following: the reconstructed image of the current image, the predicted image of the current image, the first reference image of the current image, the second reference image of the current image, at least one quantization parameter for the current image, and segmentation information for the current image.

[0222] In an embodiment, the prediction decoding unit may obtain a feature map corresponding to the current image by inputting information corresponding to the current image into a filtering neural network. The prediction decoding unit may also obtain an output image used to obtain an in-loop filtered image by inputting information corresponding to the current image into a filtering neural network. For example, the prediction decoding unit may obtain a feature map corresponding to the current image by inputting information corresponding to the current image into a feature extraction neural network included in the filtering neural network. The prediction decoding unit may obtain an output image from a regression neural network included in the filtering neural network that takes the feature map corresponding to the current image as input, wherein the output image is used to obtain the in-loop filtered image. On the other hand, since it has already referred to... Figure 7 and Figures 9 to 10 An example structure of a filtering neural network is described in detail, so its redundant descriptions are omitted.

[0223] In an embodiment, the prediction decoding unit can obtain the state information of the current image by using a feature map corresponding to the current image obtained from a filtering neural network. For example, the prediction decoding unit can obtain the feature map corresponding to the current image from a feature extraction neural network included in the filtering neural network. Furthermore, the prediction decoding unit can obtain the state information of the current image by inputting the feature map corresponding to the current image into a state encoding neural network, wherein the state information of the current image has a lower resolution than the reconstructed or predicted image of the current image. The state information of the current image can be obtained by inputting the feature map corresponding to the current image into the state encoding neural network and can be stored in a state information buffer.

[0224] On the other hand, because Figure 7 The description of state-encoded neural networks is also disclosed in the literature, so the same redundant descriptions are omitted.

[0225] Figure 14 This is a flowchart of an image decoding method according to an embodiment of the present disclosure.

[0226] In operation S1410, the image decoding device 1300 may generate layer information to be used in at least one layer of the filtering neural network based on first state information obtained from a feature map corresponding to a first reference image of the current image or second state information obtained from a feature map corresponding to a second reference image of the current image.

[0227] In an embodiment, the image decoding apparatus 1300 may obtain prediction information from the bitstream. The prediction information may include information indicating that a first reference image and / or a second reference image are used to decode the current image. Specifically, the prediction information may include information indicating that the first reference image and / or the second reference image are used to decode a block included in the current image, and may include information indicating that a first reference block indicated by a first motion vector and / or a second reference block indicated by a second motion vector are used.

[0228] In this embodiment, the image decoding device 1300 can obtain first state information based on a feature map corresponding to a first reference image of the current image. The image decoding device 1300 can also obtain second state information based on a feature map corresponding to a second reference image of the current image. The image decoding device 1300 can obtain the first state information corresponding to the first reference image by inputting the feature map corresponding to the first reference image into a state encoding neural network. Similarly, the image decoding device 1300 can obtain the second state information corresponding to the second reference image by inputting the feature map corresponding to the second reference image into a state encoding neural network.

[0229] In an embodiment, the image decoding device 1300 may input information corresponding to a first reference image into a filtering neural network. The image decoding device 1300 may obtain a feature map corresponding to the first reference image based on inputting the information corresponding to the first reference image into the filtering neural network. For example, the image decoding device 1300 may obtain the feature map corresponding to the first reference image from a feature extraction neural network included in the filtering neural network. The image decoding device 1300 may obtain first state information indicating state information of the first reference image by inputting the feature map corresponding to the first reference image into a state encoding neural network. The first state information may have a lower resolution than that of the first reference image.

[0230] In an embodiment, the image decoding device 1300 may input information corresponding to a second reference image into a filtering neural network. The image decoding device 1300 may obtain a feature map corresponding to the second reference image based on inputting the information corresponding to the second reference image into the filtering neural network. For example, the image decoding device 1300 may obtain the feature map corresponding to the second reference image from a feature extraction neural network included in the filtering neural network. The image decoding device 1300 may obtain second state information indicating the state information of the second reference image by inputting the feature map corresponding to the second reference image into a state encoding neural network. The second state information may have a lower resolution than the second reference image. On the other hand, it has already been referenced... Figures 1 to 13The operation of the image decoding device 1300 in generating state information of the current image is described in detail. Since the operation of generating state information of the current image can correspond to the operation of generating first state information of the first reference image and / or second state information of the second reference image, redundant descriptions of the same operation are omitted.

[0231] In this embodiment, the image decoding device 1300 may store first state information and / or second state information in a state information buffer. The image decoding device 1300 may use the first state information and / or second state information to obtain an image in which in-loop filtering has been applied to the current image.

[0232] In an embodiment, when a block included in the current image uses a block included in the first reference image and / or the second reference image as a reference block, the image decoding device 1300 can generate or obtain a predicted image by using the reference block. The image decoding device 1300 can input multiple pieces of state information corresponding to the reference block used to generate the predicted image of the current image into the layer information neural network.

[0233] In one embodiment, the image decoding device 1300 can generate layer information to be applied to at least one layer within a filtering neural network by inputting first state information and second state information into the layer information neural network. The layer information may include at least one kernel and attention map applied to at least one layer within the filtering neural network.

[0234] On the other hand, because it has already been referenced Figures 1 to 13 The layer information is described in detail, so the redundant detailed descriptions provided above have been omitted.

[0235] In operation S1420, the image decoding device 1300 can input information corresponding to the current image into at least one layer in the filtering neural network.

[0236] In an embodiment, the image decoding device 1300 may apply the generated layer information to a filtering neural network. For example, a feature extraction neural network included in the filtering neural network may include neural layers using convolutional kernels included in the layer information. Optionally, a regression neural network included in the filtering neural network may include neural layers that include attention maps included in the layer information. However, this disclosure is not limited to the disclosed examples. Furthermore, the layer information applied to the feature extraction neural network and the layer information applied to the regression neural network may be different because the purposes of the respective neural networks are different from each other.

[0237] In an embodiment, the image decoding device 1300 may input information corresponding to the current image into at least one layer of the filtered neural network containing used layer information. The image decoding device 1300 may input information corresponding to the current image into the filtered neural network including at least one layer containing used layer information. The information corresponding to the current image may include at least one of the following: a reconstructed image of the current image, a deblocking filtered image of the current image, a sample adaptive offset filtered image of the current image, an adaptive loop filtered image of the current image, a predicted image of the current image, a first reference image of the current image, a second reference image of the current image, at least one quantization parameter for the current image, and segmentation information for the current image.

[0238] In one embodiment, the image decoding device 1300 can input information corresponding to the current image into a feature extraction neural network included in a filtering neural network. On the other hand, because it has already referenced... Figures 1 to 13 The operation of the image decoding device 1300 in inputting the generated layer information into the filtering neural network and at least one layer within the filtering neural network is described in detail, therefore, the same descriptions thereof are omitted.

[0239] During operation S1430, the image decoding device 1300 can obtain an image to which in-loop filtering has been applied by inputting information corresponding to the current image into at least one layer of the used layer information.

[0240] In one embodiment, the image decoding device 1300 may obtain an output image from a filtering neural network, which is used to obtain an in-loop filtered image that has been subjected to in-loop filtering on the current image.

[0241] In one embodiment, in response to the image decoding device 1300 inputting information corresponding to the current image into a filtering neural network, a feature extraction neural network included in the filtering neural network can generate a feature map corresponding to the current image. Furthermore, the feature map corresponding to the current image obtained from the feature extraction neural network can be input into a regression neural network included in the filtering neural network. The image decoding device 1300 can obtain an output image from the regression neural network that receives the feature map corresponding to the current image.

[0242] In an embodiment, when the filtering neural network and the deblocking filtering unit are configured in parallel within the in-loop filtering unit, the image decoding device 1300 can obtain the in-loop filtered image based on performing an addition operation on the output image, which is the output of the filtering neural network, and the deblocked filtered image. For example, the image decoding device 1300 can obtain an image with applied in-loop filtering by performing sample adaptive loop filtering and adaptive loop filtering on the result of performing an addition operation on the output image and the deblocked filtered image.

[0243] In an embodiment, when the filtering neural network is placed in series after the deblocking filtering unit within the in-loop filtering unit, the image decoding device 1300 can input the deblocked filtered image as one of multiple pieces of information corresponding to the current image into the filtering neural network. The image decoding device 1300 can obtain the image with applied in-loop filtering by performing sample adaptive offset filtering and adaptive loop filtering on the output image obtained from the filtering neural network.

[0244] On the other hand, this disclosure is not limited to the disclosed examples. The position of the filtering neural network can be adjusted according to the design of the in-loop filtering unit, and the filtering neural network can be trained according to the position of the filtering neural network to generate an appropriate output image.

[0245] On the other hand, because it has already been referenced Figures 1 to 13 The operation of generating the output image from the filtering neural network is described in detail, so its redundant descriptions are omitted.

[0246] Figure 15 This is a block diagram illustrating the configuration of an image encoding apparatus according to an embodiment of the present disclosure.

[0247] Reference Figure 15 The image encoding device 1500 may include a prediction encoding unit 1510 and a generation unit 1530. Figure 15 The prediction coding unit 1510 shown can be used with Figure 1 The prediction coding unit 115 shown corresponds to the one shown, and the generation unit 1530 can be associated with... Figure 1 The in-loop filtering unit and predictive coding unit shown in the figure correspond to each other.

[0248] The prediction coding unit 1510 and the generation unit 1530 according to the embodiment can be implemented as at least one processor. In the embodiment, the prediction coding unit 1510 and the generation unit 1530 can operate according to at least one instruction stored in at least one memory.

[0249] On the other hand, since at least one processor and at least one memory of the image encoding device 1500 correspond to at least one processor and at least one memory of the image decoding device 1300, respectively, their identical descriptions are omitted. In the following embodiments, unless otherwise stated, they are described as consisting of detailed components included in the image encoding device 1500 or the image decoding device 1300. Figure 15 The operations performed by 1510 and 1530 can be considered to be performed by the processor.

[0250] In this embodiment, the image encoding device 1500 may perform operations that are performed in the image decoding device 1300 or in units included in the image decoding device 1300. Redundant descriptions thereof are omitted.

[0251] In an embodiment, the image encoding apparatus 1500 may include at least one memory storing input and output data of the predictive encoding unit 1510 and the generation unit 1530. Furthermore, the image encoding apparatus 1500 may include a memory control unit for controlling the data input and output of the memory.

[0252] In an embodiment, the prediction coding unit 1510 may determine prediction information for generating a prediction block within the current block of the current image. For example, the prediction information may include information indicating that the prediction mode of the current block is an inter-frame mode. An inter-frame mode is a mode for predicting or reconstructing the current block based on a reference image to reduce temporal redundancy between images. The current block may be the largest CU, CU, transform unit, or prediction unit divided from the current image to be encoded. Furthermore, when the prediction mode of the current block is an inter-frame mode, the prediction information may include information for determining a reference block. For example, the information for determining a reference block may include the index and motion information of a reference image within a list of reference frames, but this disclosure is not limited to the disclosed examples.

[0253] In an embodiment, the prediction coding unit 1510 may perform inter-frame prediction on the current block based on the prediction information of the determined current block, and may encode the current block by using the prediction block generated as a result of performing the inter-frame prediction.

[0254] In an embodiment, when the prediction coding unit 1510 encodes the current block based on a reference image, the prediction coding unit 1510 may use one reference image (e.g., one-way prediction) or two reference images (e.g., two-way prediction). Whether the current block is predicted one-way or two-way may also be included in the bitstream as prediction information determined for the current block, and may be represented as flags or indexes. When it is implicitly determined from the prediction patterns of neighboring blocks associated with the current block whether one-way or two-way prediction is performed on the current block, information about neighboring blocks may be included in the bitstream as flags or indexes.

[0255] In an embodiment, when the current block is bidirectionally predicted, the prediction coding unit 1510 can identify motion information of a reference block used to determine the current block, and the motion information can be included in the prediction information. The motion information of the current block can be included in the bitstream. The motion information of the current block may include at least one of a reference image index, a motion vector, a differential motion vector, and a reference direction, and may include all multiple pieces of information used to predict the motion vector of the current block.

[0256] In this embodiment, the prediction information may be included in the sequence parameter set, frame parameter set, strip header, or strip data of the bitstream.

[0257] In this embodiment, the process of encoding the current block to generate an indexable information enables the image decoding device 1300 to reconstruct the information of the current block. The information generated by encoding can be included in the bitstream.

[0258] In this embodiment, when a prediction block is generated by performing bidirectional prediction on the current block, the prediction coding unit 1510 can encode the current block using the prediction block. The prediction coding unit 1510 can generate residual data corresponding to the difference between the prediction block and the current block. When the prediction block is determined to be the current block, residual data may not be generated.

[0259] In one embodiment, the generation unit 1530 may generate a bitstream including the result of encoding the image. The bitstream may include the result of encoding the current block. The generation unit 1530 may transmit the bitstream to the image decoding device 1300 via a network.

[0260] In an embodiment, the generation unit 1530 may store the bitstream in a data storage medium, including magnetic media (such as hard disks, floppy disks, and magnetic tapes), optical recording media (such as CD-ROMs and DVDs), magneto-optical media (such as optical-magnetic floppy disks), etc. The generation unit 1530 may generate a bitstream comprising syntax elements generated by encoding an image. Values ​​corresponding to the syntax elements may be included in the bitstream according to the hierarchical structure of the image. The bitstream generated when the generation unit 1530 performs entropy encoding on the syntax elements may be included in the bitstream.

[0261] In an embodiment, the predictive coding unit 1510 may perform the same filtering as the in-loop filtering performed in the predictive decoding unit. On the other hand, although not explicitly disclosed, it is understood that... Figures 1 to 14 The operations performed in the publicly disclosed prediction decoding unit or image decoding device 1300 can also be performed in the prediction coding unit 1510.

[0262] In an embodiment, the predictive coding unit 1510 may determine filter information for filtering performed in the in-loop filtering unit. For example, the filter information may include information used to perform deblocking filtering, sample offset adaptive filtering, and adaptive loop filtering included in the in-loop filtering unit. The filter information may include whether deblocking filtering, sample offset adaptive filtering, and adaptive loop filtering are performed, and filter parameters for performing each filter.

[0263] In this embodiment, the predictive coding unit 1510 may perform intra-loop filtering on the current block or current image based on filter information determined for the current block or current image, and may encode the current block using an intra-loop filtered block or intra-loop filtered image generated as a result of performing intra-loop filtering. For example, the predictive coding unit 1510 may store the intra-loop filtered image in a frame buffer and may encode the image by using the intra-loop filtered image as a reference image. Alternatively, the intra-loop filtered image may be an image to which intra-loop filtering has been applied.

[0264] In an embodiment, filter information may be included in the sequence parameter set, frame parameter set, strip header, or strip data of the bitstream.

[0265] In one embodiment, the predictive coding unit 1510 can input information corresponding to the current image into a filtering neural network to generate an output image used to obtain an in-loop filtered image of the current image. The predictive coding unit 1510 can obtain the output image of the current image from the filtering neural network.

[0266] In an embodiment, the predictive coding unit 1510 may include an in-loop filtering unit. The in-loop filtering unit may include a filtering neural network, which includes a feature extraction neural network and a regression neural network. Furthermore, the in-loop filtering unit may use a state-encoding neural network and a layer information neural network to generate layer information that will be used in at least one layer within the filtering neural network.

[0267] In an embodiment, the predictive coding unit 1510 may use state information corresponding to a reference block for each block in the current image to obtain an output image. For example, state information of a reference image may be determined by using reference blocks corresponding to all blocks included in the current image, and this state information of the reference image is obtained as input to a layer information neural network to generate layer information for a filtering neural network to be applied to the current image. On the other hand, since references have already been made... Figure 7 and Figures 11a to 11c The operation of the predictive coding unit 1510 in generating layer information by using a layer information neural network is described in detail, so the same redundant descriptions are omitted.

[0268] For example, when the current block is being predicted bidirectionally, the prediction coding unit 1510 can perform a weighted sum by applying weights to a first reference block and a second reference block, respectively, and generate a predicted block for the current block. The state information corresponding to each of the reference blocks included in the current image can be obtained by applying weights to a weighted sum of multiple state information corresponding to the two reference blocks, respectively, for the two reference blocks used for bidirectional prediction.

[0269] On the other hand, for ease of explanation, a method for using reference block state information for blocks of prediction units has been described. However, this disclosure is not limited to the disclosed examples, and the prediction coding unit 1510 can generate layer information by using state information corresponding to the reference block used for each of the prediction frame units, strip units, parallel block units, maximum CU, CU, and prediction units. For example, in order to generate layer information at the frame unit level, the prediction coding unit 1510 can generate layer information by inputting state information corresponding to the reference block of the block included in the current image into the layer information neural network.

[0270] In an embodiment, the predictive coding unit 1510 may obtain layer information from the layer information neural network that will be used in at least one layer within the filtering neural network. The predictive coding unit 1510 may obtain at least one of a convolutional kernel and an attention map as layer information applied to the filtering neural network. The predictive coding unit 1510 may use the obtained layer information in the filtering neural network. For example, the predictive coding unit 1510 may use a convolutional kernel and an attention map to configure the filtering neural network to include convolutional layers using the convolutional kernel and attention map included in the layer information. The convolutional kernel and attention map included in at least one convolutional layer in the filtering neural network may be determined based on the obtained layer information. On the other hand, since reference has been made... Figure 7 and Figures 8a to 8c The operation of the predictive decoding device in configuring the filtering neural network using layer information is described in detail, so the same redundant descriptions are omitted.

[0271] In an embodiment, the prediction coding unit 1510 may input information corresponding to the current image into the filtering neural network. The information corresponding to the current image includes at least one of the following: a reconstructed image of the current image, a predicted image of the current image, a first reference image of the current image, a second reference image of the current image, at least one quantization parameter for the current image, and segmentation information for the current image.

[0272] In an embodiment, the predictive coding unit 1510 may obtain a feature map corresponding to the current image by inputting information corresponding to the current image into a filtering neural network. The predictive coding unit 1510 may also obtain an output image used to obtain an image for which in-loop filtering has been applied, based on inputting information corresponding to the current image into the filtering neural network. For example, the predictive coding unit 1510 may obtain a feature map corresponding to the current image by inputting information corresponding to the current image into a feature extraction neural network included in the filtering neural network. The predictive coding unit 1510 may obtain an output image from a regression neural network included in the filtering neural network based on the input feature map corresponding to the current image. On the other hand, since it has already referenced... Figure 7 and Figures 9 to 10An example structure of a filtering neural network is described in detail, so its redundant descriptions are omitted.

[0273] In an embodiment, the predictive coding unit 1510 can obtain the state information of the current image by using a feature map corresponding to the current image obtained from a filtering neural network. For example, the predictive coding unit 1510 can obtain the feature map corresponding to the current image from a feature extraction neural network included in the filtering neural network. Furthermore, the predictive coding unit 1510 can obtain the state information of the current image by inputting the feature map corresponding to the current image into a state coding neural network, wherein the state information of the current image has a lower resolution than the reconstructed or predicted image of the current image. The state information of the current image can be obtained by inputting the feature map corresponding to the current image into the state coding neural network and can be stored in a state information buffer.

[0274] On the other hand, because Figure 7 The description of state-encoded neural networks is also disclosed in the literature, so the same redundant descriptions are omitted.

[0275] Figure 16 This is a flowchart of an image encoding method according to an embodiment of the present disclosure.

[0276] In operation S1610, the image encoding device 1500 may generate layer information to be used in at least one layer of the filtering neural network based on first state information obtained from a feature map corresponding to a first reference image of the current image or second state information obtained from a feature map corresponding to a second reference image of the current image.

[0277] In an embodiment, the image encoding apparatus 1500 can generate layer information using determined prediction information. The prediction information may include information indicating that a first reference image and / or a second reference image are used to decode the current image. Specifically, the prediction information may include information indicating that the first reference image and / or the second reference image are used to decode a block included in the current image, and may include information indicating that a first reference block indicated by a first motion vector and / or a second reference block indicated by a second motion vector are used. The image encoding apparatus 1500 can determine the prediction information and generate a bitstream including the determined prediction information.

[0278] In this embodiment, the image encoding device 1500 may obtain first state information based on a feature map corresponding to a first reference image of the current image. The image encoding device 1500 may obtain second state information based on a feature map corresponding to a second reference image of the current image. The image encoding device 1500 may obtain the first state information corresponding to the first reference image by inputting the feature map corresponding to the first reference image into a state encoding neural network. The image encoding device 1500 may obtain the second state information corresponding to the second reference image by inputting the feature map corresponding to the second reference image into a state encoding neural network.

[0279] In an embodiment, the image encoding device 1500 may input information corresponding to a first reference image into a filtering neural network. The image encoding device 1500 may obtain a feature map corresponding to the first reference image based on inputting the information corresponding to the first reference image into the filtering neural network. For example, the image encoding device 1500 may obtain the feature map corresponding to the first reference image from a feature extraction neural network included in the filtering neural network. The image encoding device 1500 may obtain first state information indicating state information of the first reference image by inputting the feature map corresponding to the first reference image into a state encoding neural network. The first state information may have a lower resolution than the first reference image.

[0280] In an embodiment, the image encoding device 1500 may input information corresponding to a second reference image into a filtering neural network. The image encoding device 1500 may obtain a feature map corresponding to the second reference image based on inputting the information corresponding to the second reference image into the filtering neural network. For example, the image encoding device 1500 may obtain the feature map corresponding to the second reference image from a feature extraction neural network included in the filtering neural network. The image encoding device 1500 may obtain second state information indicating the state information of the second reference image by inputting the feature map corresponding to the second reference image into a state encoding neural network. The second state information may have a lower resolution than the second reference image. On the other hand, information already referenced... Figures 1 to 13 The operation of the image encoding apparatus 1500 in generating state information of the current image is described in detail. Since the operation of generating state information of the current image can correspond to the operation of generating first state information of the first reference image and / or second state information of the second reference image, redundant descriptions of the same operation are omitted.

[0281] In this embodiment, the image encoding device 1500 may store first state information and / or second state information in a state information buffer. The image encoding device 1500 may use the first state information and / or second state information to obtain an image in which in-loop filtering has been applied to the current image.

[0282] In an embodiment, when a block included in the current image uses a block included in the first reference image and / or the second reference image as a reference block, the image encoding device 1500 can generate or obtain a predicted image by using the reference block. The image encoding device 1500 can input multiple pieces of state information corresponding to the reference block used to generate the predicted image of the current image into the layer information neural network.

[0283] In an embodiment, the image encoding device 1500 may generate layer information to be applied to at least one layer within a filtering neural network by inputting first state information and second state information into a layer information neural network. The layer information may include at least one kernel and attention map applied to at least one layer within the filtering neural network.

[0284] On the other hand, because it has already been referenced Figures 1 to 13 The layer information is described in detail, so the redundant detailed descriptions provided above have been omitted.

[0285] During operation S1620, the image encoding device 1500 can input information corresponding to the current image into at least one layer in the filtering neural network.

[0286] In an embodiment, the image encoding device 1500 may apply the generated layer information to a filtering neural network. For example, the feature extraction neural network included in the filtering neural network may include neural layers using convolutional kernels included in the layer information. Optionally, the regression neural network included in the filtering neural network may include neural layers that include attention maps included in the layer information. However, this disclosure is not limited to the disclosed examples. Furthermore, the layer information applied to the feature extraction neural network and the layer information applied to the regression neural network may be different because the purposes of the respective neural networks are different from each other.

[0287] In an embodiment, the image encoding device 1500 may input information corresponding to the current image into at least one layer of the filtered neural network containing used layer information. The image encoding device 1500 may input information corresponding to the current image into the filtered neural network including at least one layer containing used layer information. The information corresponding to the current image may include at least one of the following: a reconstructed image of the current image, a deblocking filtered image of the current image, a sample adaptive offset filtered image of the current image, an adaptive loop filtered image of the current image, a predicted image of the current image, a first reference image of the current image, a second reference image of the current image, at least one quantization parameter for the current image, and segmentation information for the current image.

[0288] In one embodiment, the image encoding device 1500 can input information corresponding to the current image into a feature extraction neural network included in a filtering neural network. On the other hand, because it has already referenced... Figures 1 to 13The operation of the image encoding device 1500 in inputting the generated layer information into the filter neural network and at least one layer within the filter neural network is described in detail, therefore, the same descriptions thereof are omitted.

[0289] In operation S1630, the image encoding device 1500 can obtain an image to which in-loop filtering has been applied by inputting information corresponding to the current image into at least one layer of the used layer information.

[0290] In one embodiment, the image encoding device 1500 may obtain an output image from a filtering neural network, which is used to obtain an in-loop filtered image that has been subjected to in-loop filtering on the current image.

[0291] In this embodiment, in response to the image encoding device 1500 inputting information corresponding to the current image into a filtering neural network, a feature extraction neural network included in the filtering neural network can generate a feature map corresponding to the current image. Furthermore, the feature map corresponding to the current image obtained from the feature extraction neural network can be input into a regression neural network included in the filtering neural network. The image encoding device 1500 can obtain an output image from the regression neural network that receives the feature map corresponding to the current image.

[0292] In an embodiment, when the filtering neural network and the deblocking filtering unit are configured in parallel within the in-loop filtering unit, the image encoding device 1500 can obtain the in-loop filtered image based on performing an addition operation on the output image, which is the output of the filtering neural network, and the deblocked filtered image. For example, the image encoding device 1500 can obtain an image with applied in-loop filtering by performing sample adaptive loop filtering and adaptive loop filtering on the result of performing an addition operation on the output image and the deblocked filtered image.

[0293] In an embodiment, when the filtering neural network is placed in series after the deblocking filtering unit within the in-loop filtering unit, the image encoding device 1500 can input the deblocked filtered image as one of multiple pieces of information corresponding to the current image into the filtering neural network. The image encoding device 1500 can obtain the image with applied in-loop filtering by performing sample adaptive offset filtering and adaptive loop filtering on the output image obtained from the filtering neural network.

[0294] On the other hand, this disclosure is not limited to the disclosed examples. The position of the filtering neural network can be adjusted according to the design of the in-loop filtering unit, and the filtering neural network can be trained according to the position of the filtering neural network to generate an appropriate output image.

[0295] On the other hand, because it has already been referenced Figures 1 to 13 The operation of generating the output image from the filtering neural network is described in detail, so its redundant descriptions are omitted.

[0296] On the other hand, because it has already been referenced Figures 1 to 15 The operation of generating the output image from the filtering neural network is described in detail, so redundant descriptions are omitted. Furthermore, besides... Figure 16 The operation is performed outside of the image encoding device 1500. Figure 16 Operation can be combined with Figures 1 to 14 The public response.

[0297] In an embodiment, the image decoding method may include operation S1410: generating layer information to be used in at least one layer within the filtering neural network 720 based on first state information obtained from a feature map corresponding to a first reference image of the current image or second state information obtained from a feature map corresponding to a second reference image of the current image. The image decoding method may include operation S1420: inputting information 710 corresponding to the current image into at least one layer within the filtering neural network 720. The image decoding method may include operation S1430: obtaining an image to which in-loop filtering has been applied by inputting information 710 corresponding to the current image into at least one layer that has used the layer information. The state information 735 of the current image may be obtained from the feature map 725 corresponding to the current image and may have a lower resolution than the current image.

[0298] In an embodiment, layer information may include at least one kernel and attention map applied to at least one layer within the filtered neural network 720.

[0299] In this embodiment, the first reference image may be an image included in a first reference frame list of the current image. The second reference image may be an image included in a second reference frame list of the current image. The POC value of the first reference image may be less than the POC value of the current image. The POC value of the second reference image may be greater than the POC value of the current image.

[0300] In an embodiment, the first reference image and the second reference image are images included in a second reference image list of the current image, and each of the POC values ​​of the first reference image and the second reference image may be greater than the POC value of the current image.

[0301] In this embodiment, the state information 735 of the current image can be obtained by inputting the feature map 725 corresponding to the current image into the state encoding neural network 730. The state information 735 of the current image can be stored in the state information buffer 740.

[0302] In one embodiment, first state information can be obtained by inputting a feature map corresponding to the first reference image into a state-encoding neural network 730. The first state information can be stored in a state information buffer 740.

[0303] In one embodiment, the second state information can be obtained by inputting a feature map corresponding to the second reference image into a state encoding neural network 730. The second state information can be stored in a state information buffer 740.

[0304] In an embodiment, the filtering neural network 720 may include a feature extraction neural network 721 and a regression neural network 726, wherein the feature extraction neural network 721 extracts features of the current image from information corresponding to the current image and generates a feature map corresponding to the current image, and the regression neural network 726 is used to obtain an image from the feature map corresponding to the current image, in which in-loop filtering has been applied to the current image.

[0305] In an embodiment, the information 710 corresponding to the current image may include at least one of the following: a reconstructed image of the current image, a predicted image of the current image, a first reference image, a second reference image, at least one quantization parameter for the current image, and segmentation information for the current image.

[0306] In an embodiment, the image decoding method may obtain the in-loop filtered image of the current image by performing an addition operation on the output image 760 obtained from the regression neural network 726 and the deblocked filtered image of the current image.

[0307] In an embodiment, the filtering neural network may include at least one of a deep neural network, a convolutional neural network, and a transformer neural network.

[0308] In an embodiment, the feature extraction neural network 721 may include at least one downsampling layer for generating a feature map 725 corresponding to the current image, the feature map 725 having a lower resolution than the current image.

[0309] In an embodiment, the regression neural network 726 may include at least one upsampling layer for generating an output image having a higher resolution than the feature map 725 corresponding to the current image.

[0310] In an embodiment, the state-encoded neural network 730 may include at least one downsampling layer, and the at least one downsampling layer may include at least one of a strided convolutional layer and a pooling layer.

[0311] In an embodiment, the layer information neural network used to generate layer information may include at least one of stride convolutional layers, pooling layers, interpolation layers, deconvolutional layers, and pixel rearrangement layers.

[0312] In an embodiment, the filtering neural network, the state coding neural network, and the layer information neural network can be trained together by inputting information corresponding to the first reference image or the second reference image into the filtering neural network and inputting information corresponding to the current image into the filtering neural network.

[0313] In an embodiment, the image decoding apparatus 1300 may include at least one memory storing at least one instruction and at least one processor operating according to the at least one instruction. The at least one processor may generate layer information for use in at least one layer within the filtering neural network 720 based on first state information obtained from a feature map corresponding to a first reference image of the current image or second state information obtained from a feature map corresponding to a second reference image of the current image. The at least one processor may input information 710 corresponding to the current image into at least one layer within the filtering neural network 720. The at least one processor may obtain an image to which in-loop filtering has been applied by inputting information 710 corresponding to the current image into at least one layer that has used the layer information. The state information 735 of the current image may be obtained from a feature map 725 corresponding to the current image and may have a lower resolution than the current image.

[0314] In one embodiment, the image encoding method may include operation S1610: generating layer information to be used in at least one layer of the filtering neural network 720 based on first state information obtained from a feature map corresponding to a first reference image of the current image or second state information obtained from a feature map corresponding to a second reference image of the current image. In another embodiment, the image encoding method may include operation S1620: inputting information 710 corresponding to the current image into at least one layer of the filtering neural network 720.

[0315] In an embodiment, the image encoding method may include operation S1630: obtaining an image to which in-loop filtering has been applied by inputting information 710 corresponding to the current image into at least one layer that has used layer information. The state information 735 of the current image may be obtained from a feature map 725 corresponding to the current image and may have a lower resolution than that of the current image.

[0316] In an embodiment, the image encoding apparatus 1500 may include at least one memory storing at least one instruction and at least one processor operating according to the at least one instruction. The at least one processor may generate layer information to be used in at least one layer within the filtering neural network 720 based on first state information obtained from a feature map corresponding to a first reference image of the current image or second state information obtained from a feature map corresponding to a second reference image of the current image. The at least one processor may input information 710 corresponding to the current image into at least one layer within the filtering neural network 720. The at least one processor may obtain an image to which in-loop filtering has been applied by inputting information 710 corresponding to the current image into at least one layer that has used the layer information. The state information 735 of the current image may be obtained from a feature map 725 corresponding to the current image and may have a lower resolution than the current image.

[0317] In an embodiment, a computer-readable recording medium on which a bitstream is recorded may be provided. The bitstream may include prediction information used to predict a current image. Based on the prediction information, layer information to be used in at least one layer within a filtering neural network 720 may be generated based on first state information obtained from a feature map corresponding to a first reference image or second state information obtained from a feature map corresponding to a second reference image. Information 710 corresponding to the current image may be input to at least one layer in the filtering neural network 720. An image to which in-loop filtering has been applied to the current image may be obtained by inputting information 710 corresponding to the current image to at least one layer that has used the layer information. The state information 735 of the current image may be obtained from a feature map 725 corresponding to the current image and may have a lower resolution than that of the current image.

[0318] Various embodiments of this disclosure may be implemented or supported by one or more computer programs, and the computer programs may be formed from computer-readable program code and recorded on a computer-readable medium. In this disclosure, "application" and "program" may refer to one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in computer-readable program code. "Computer-readable program code" may include various types of computer code, including source code, object code, and executable code. "Computer-readable medium" may include various types of media accessible to a computer, such as ROM, RAM, hard disk drive (HDD), optical disk drive (CD), DVD, or other types of storage.

[0319] Furthermore, machine-readable storage media can be provided in the form of non-transitory storage media. "Non-transitory storage media" refers to tangible devices and excludes wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. On the other hand, "non-transitory storage media" does not distinguish between cases where data is semi-permanently stored on the storage medium and cases where data is temporarily stored. For example, "non-transitory storage media" may include buffers for temporarily storing data. Computer-readable recording media can be any available medium accessible by a computer and may include any volatile and non-volatile media, as well as any removable and non-removable media. Computer-readable media include media on which data can be permanently stored and media on which data can be stored and later rewritten, such as rewritable optical discs or erasable memory devices.

[0320] The methods according to embodiments of this disclosure can be provided by being included in a computer program product. The computer program product can be traded as a commodity between a seller and a buyer. The computer program product can be distributed in the form of a machine-readable storage medium (e.g., CD-ROM), or can be distributed online (e.g., downloaded or uploaded) between two user devices (e.g., smartphones) via an app store or directly. In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) is at least temporarily stored on a machine-readable storage medium (such as the memory of a manufacturer's server, an app store's server, or a relay server), or can be temporarily generated.

[0321] The above description of this disclosure is for illustrative purposes only, and those skilled in the art will understand that it can be modified in other specific forms without altering the technical spirit or essential characteristics of this disclosure. For example, appropriate results can be achieved even when the above techniques are performed in a different order than the above methods, and / or when components of the above computer system or module are coupled or combined in a different manner than the above methods, or when they are replaced or superseded by other components or equivalents. Therefore, it should be understood that the above embodiments are illustrative in all respects and not restrictive. For example, a component described as a single entity can be implemented in a distributed manner. Similarly, a component described as distributed can be implemented in a composite manner.

[0322] The scope of this disclosure is indicated by the claims described below, and all changes or modifications derived from the meaning and scope of the claims and their equivalents shall be construed as falling within the scope of this disclosure.

Claims

1. An image decoding method comprising: generating layer information (S1410) to be used in at least one layer within a filtering neural network (720) based on first state information obtained from a feature map corresponding to a first reference picture of a current picture or second state information obtained from a feature map corresponding to a second reference picture of the current picture; inputting information (710) corresponding to the current picture to the at least one layer within the filtering neural network (720) (S1420); and obtaining a picture in which in-loop filtering has been applied to the current picture by inputting the information (710) corresponding to the current picture to the at least one layer having used the layer information (S1430), wherein the state information (735) of the current picture is obtained from a feature map (725) corresponding to the current picture and has a lower resolution than a resolution of the current picture.

2. The image decoding method of claim 1, wherein, the layer information includes at least one of a kernel and an attention map applied to the at least one layer within the filtering neural network (720).

3. The image decoding method according to claim 1 or 2, wherein the first reference picture is a picture included in a first reference picture list of the current picture, the second reference picture is a picture included in a second reference picture list of the current picture, a picture order count (POC) value of the first reference picture is less than a POC value of the current picture, and the POC value of the second reference picture is greater than the POC value of the current picture.

4. The image decoding method according to claim 1 or 2, wherein the first reference picture and the second reference picture are pictures included in the second reference picture list of the current picture, and each of a POC value of the first reference picture and a POC value of the second reference picture is greater than the POC value of the current picture.

5. The image decoding method according to any one of claims 1 to 4, wherein the state information (735) of the current picture is obtained by inputting a feature map (725) corresponding to the current picture to a state encoding neural network (730) and is stored in a state information buffer (740).

6. The image decoding method according to any one of claims 1 to 5, wherein the first state information is obtained by inputting a feature map corresponding to the first reference picture to the state encoding neural network (730) and is stored in the state information buffer (740), the second state information is obtained by inputting a feature map corresponding to the second reference picture to the state encoding neural network (730) and is stored in the state information buffer (740).

7. The image decoding method according to any one of claims 1 to 6, wherein the filtering neural network (720) includes a feature extraction neural network (721) configured to extract a feature of the current picture from the information corresponding to the current picture and generate a feature map corresponding to the current picture and a regression neural network (726) used to obtain a picture in which in-loop filtering has been applied to the current picture from the feature map corresponding to the current picture.

8. The image decoding method according to any one of claims 1 to 7, wherein the information (710) corresponding to the current picture includes at least one of a reconstructed picture of the current picture, a predicted picture of the current picture, the first reference picture, the second reference picture, at least one quantization parameter for the current picture, and partition information for the current picture.

9. The image decoding method according to claim 7 or 8, wherein An image to which in-loop filtering has been applied to the current image is obtained including: based on performing an addition operation on an output image (760) obtained from a regression neural network (726) and a deblocking filtered image of the current image, an in-loop filtered image of the current image is obtained.

10. The image decoding method of any one of claims 1 to 9, wherein, The filtering neural network (720) includes at least one of a deep neural network, a convolutional neural network, and a transformer neural network.

11. The image decoding method of any one of claims 7 to 10, wherein, The feature extraction neural network (721) includes at least one down-sampling layer for generating a feature map (725) corresponding to the current image, the feature map (725) corresponding to the current image having a resolution lower than a resolution of the current image.

12. The image decoding method according to any one of claims 7 to 11, wherein The regression neural network (726) includes at least one up-sampling layer for generating an output image, the output image having a resolution higher than a resolution of the feature map (725) corresponding to the current image.

13. The image decoding method of any one of claims 5 to 12, wherein, The filtering neural network, the state encoding neural network, and the layer information neural network used to generate the layer information are trained together based on inputting information corresponding to the first reference image or information corresponding to the second reference image to the filtering neural network and inputting information corresponding to the current image to the filtering neural network.

14. An image encoding method comprising: generating layer information (S1610) to be used in at least one layer within a filtering neural network (720) based on first state information obtained from a feature map corresponding to a first reference image of a current image or second state information obtained from a feature map corresponding to a second reference image of the current image; inputting information (710) corresponding to the current image to the at least one layer within the filtering neural network (720) (S1620); and obtaining an image to which in-loop filtering has been applied to the current image by inputting the information (710) corresponding to the current image to the at least one layer that has used the layer information (S1630), wherein the state information (735) of the current image is obtained from a feature map (725) corresponding to the current image and has a resolution lower than a resolution of the current image.

15. A computer-readable recording medium having a bitstream recorded thereon, wherein, a bitstream includes prediction information used to predict a current image, the layer information to be used in at least one layer within a filtering neural network (720) is generated based on first state information obtained from a feature map corresponding to a first reference image or second state information obtained from a feature map corresponding to a second reference image based on the prediction information, information (710) corresponding to the current image is input to the at least one layer within the filtering neural network (720), an image to which in-loop filtering has been applied to the current image is obtained by inputting the information (710) corresponding to the current image to the at least one layer that has used the layer information, and the state information (735) of the current image is obtained from a feature map (725) corresponding to the current image and has a resolution lower than a resolution of the current image.