AI-based image encoding apparatus and image decoding apparatus, and image encoding and decoding methods thereof

The transformation core is generated through neural networks and applied to image residual blocks, which solves the problem of inefficient image encoding and decoding in the prior art, and achieves more efficient image processing and lower bit rate.

CN120036001APending Publication Date: 2025-05-23SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072868.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-07
Filing Date
2023-09-22
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to effectively perform image transformation and inverse transformation through artificial intelligence (AI), resulting in inefficient image encoding and decoding.

Method used

By generating transformation cores using neural networks, the residual blocks of images are applied to obtain reconstruction blocks, and the encoding and decoding of images are achieved. The method includes obtaining a transform block of the residual block from the bitstream, generating a transform core using the prediction block, adjacent pixels, and codec context information, and applying it to the residual block to obtain a reconstruction block.

Benefits of technology

The efficiency of image encoding and decoding is improved, and the transformation core is dynamically generated, and the characteristics of different image blocks are adapted to the quality of different image blocks, the bit rate is reduced, and the image quality is maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120036001A_ABST
    Figure CN120036001A_ABST
Patent Text Reader

Abstract

An artificial intelligence (AI)-based image decoding method and an apparatus for performing the AI-based image decoding method are provided. According to the AI-based image decoding method, a transform block of a residual block of a current block is obtained from a bitstream, a transform kernel of the transform block is generated by applying a prediction block of the current block, neighboring pixels of the current block, and codec context information to a neural network, the residual block is obtained by applying the generated transform kernel to the transform block, and the residual block is decoded by applying the generated transform kernel to the transform block. And reconstructing the current block by using the residual block and the prediction block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to encoding and decoding images. More specifically, the present disclosure relates to a technique for encoding and decoding images by using artificial intelligence (AI) (eg, a neural network). Background Art

[0002] In a codec such as H.264 Advanced Video Coding (AVC) or High Efficiency Video Coding (HEVC), an image is partitioned into blocks, a transform block is obtained by predicting each block and transforming it into a residual block which is the difference between the original block and the predicted block, and the transform block is quantized and entropy encoded to be sent as a bitstream.

[0003] A transform block obtained by performing entropy decoding and inverse quantization on a transmitted bit stream is inversely transformed to obtain a residual block, and the block can be reconstructed by using the residual block and a prediction block obtained by prediction.

[0004] Recently, a technology for encoding / decoding an image by using artificial intelligence (AI) has been proposed, and a method for efficiently encoding / decoding an image by performing transformation and inverse transformation by using AI (eg, a neural network) has been required. Summary of the invention

[0005] According to an embodiment of the present disclosure, an image decoding method based on artificial intelligence (AI) may include: obtaining a transform block of a residual block of a current block from a bitstream, generating a transform kernel of the transform block by applying a prediction block of the current block, neighboring pixels of the current block, and coding and decoding context information to a neural network, obtaining a residual block by applying the generated transform kernel to the transform block, and reconstructing the current block by using the residual block and the prediction block.

[0006] According to an embodiment of the present disclosure, an AI-based image decoding device may include a memory storing one or more instructions and at least one processor configured to operate according to the one or more instructions. The at least one processor may be configured to obtain a transform block of a residual block of a current block from a bitstream. The at least one processor may be configured to generate a transform kernel of a transform block by applying a prediction block of the current block, neighboring pixels of the current block, and encoding and decoding context information to a neural network. The at least one processor may be configured to obtain a residual block by applying the generated transform kernel to the transform block. The at least one processor may be configured to reconstruct the current block by using the residual block and the prediction block.

[0007] According to an embodiment of the present disclosure, an AI-based image encoding method may include: obtaining a residual block based on a prediction block of a current block and an original block of the current block, generating a transform kernel of a transform block of the residual block by applying the prediction block, adjacent pixels of the current block and coding and decoding context information to a neural network, obtaining a transform block by applying the generated transform kernel to the residual block, and generating a bitstream including the transform block.

[0008] According to an embodiment of the present disclosure, an AI-based image encoding device may include a memory storing one or more instructions and at least one processor configured to operate according to the one or more instructions. The at least one processor may be configured to obtain a residual block based on a prediction block of a current block and an original block of the current block. The at least one processor may be configured to generate a transform kernel of a transform block of a residual block by applying the prediction block, adjacent pixels of the current block, and encoding and decoding context information to a neural network. The at least one processor may be configured to obtain a transform block by applying the generated transform kernel to the residual block. The at least one processor may be configured to generate a bitstream including a transform block.

[0009] According to an embodiment of the present disclosure, an AI-based image decoding method may include: obtaining a transform feature map corresponding to a transform block of a residual block of a current block from a bitstream, generating a codec context feature map of the transform block by applying a prediction block of the current block, neighboring pixels of the current block, and codec context information to a first neural network, and reconstructing the current block by applying the transform feature map and the codec context feature map to a second neural network.

[0010] According to an embodiment of the present disclosure, an AI-based image decoding device may include a memory storing one or more instructions and at least one processor configured to operate according to the one or more instructions. The at least one processor may be configured to obtain a transform feature map corresponding to a transform block of a residual block of a current block from a bitstream. The at least one processor may be configured to generate a codec context feature map of a transform block by applying a prediction block of the current block, neighboring pixels of the current block, and codec context information to a first neural network. The at least one processor may be configured to reconstruct the current block by applying the transform feature map and the codec context feature map to a second neural network.

[0011] According to an embodiment of the present disclosure, an AI-based image encoding method may include: obtaining a residual block based on a prediction block of a current block and an original block of the current block, generating a codec context feature map of a transform block by applying the prediction block, neighboring pixels of the current block and codec context information to a first neural network, obtaining a transform feature map corresponding to the transform block by applying the codec context feature map and the residual block to a second neural network, and generating a bitstream including the transform feature map.

[0012] According to an embodiment of the present disclosure, an AI-based image encoding device may include a memory storing one or more instructions and at least one processor configured to operate according to the one or more instructions. The at least one processor may be configured to obtain a residual block based on a prediction block of a current block and an original block of the current block. The at least one processor may be configured to generate a codec context feature map of a transform block by applying the prediction block, neighboring pixels of the current block, and codec context information to a first neural network. The at least one processor may be configured to obtain a transform feature map corresponding to the transform block by applying the codec context feature map and the residual block to a second neural network. The at least one processor may be configured to generate a bitstream including a transform feature map. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a diagram showing the image encoding and decoding process.

[0014] Figure 2 is a diagram showing blocks obtained by dividing an image according to a tree structure.

[0015] Figure 3 is a diagram for describing an image encoding and decoding process based on artificial intelligence (AI) according to an embodiment of the present disclosure.

[0016] Figure 4 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0017] Figure 5 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0018] Figure 6 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0019] Figure 7 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0020] Figure 8 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0021] Fig. 9 is a flowchart of an AI-based image encoding method according to an embodiment of the present disclosure.

[0022] Fig.10 is a diagram of a configuration of an AI-based image encoding device according to an embodiment of the present disclosure.

[0023] Fig.11is a flowchart of an AI-based image decoding method according to an embodiment of the present disclosure.

[0024] Fig.12 is a diagram of a configuration of an AI-based image decoding device according to an embodiment of the present disclosure.

[0025] Fig.13 is a flowchart of an AI-based image encoding method according to an embodiment of the present disclosure.

[0026] Fig.14 is a diagram of a configuration of an AI-based image encoding device according to an embodiment of the present disclosure.

[0027] Fig.15 is a flowchart of an AI-based image decoding method according to an embodiment of the present disclosure.

[0028] Fig.16 is a diagram of a configuration of an AI-based image decoding device according to an embodiment of the present disclosure.

[0029] Fig.17 is a diagram for describing a method of training a neural network used in an AI-based image encoding method and an AI-based image decoding method according to an embodiment of the present disclosure.

[0030] Fig.18 is a diagram for describing a method of training a neural network used in an AI-based image encoding method and an AI-based image decoding method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] Throughout the disclosure, the expression "at least one of a, b, or c" means only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.

[0032] Since the present disclosure allows various changes and many examples, the specific embodiments of the present disclosure will be shown in the drawings and described in detail in the written description. However, this is not intended to limit the embodiments of the present disclosure to a specific practice mode, and it will be understood that all changes, equivalents and substitutes that do not depart from the spirit and technical scope of the present disclosure are included in the various embodiments of the present disclosure.

[0033] In the description of the embodiments of the present disclosure, when it is considered that some detailed explanations of the related art may unnecessarily obscure the essence of the present disclosure, these explanations are omitted. In addition, the numbers (e.g., first, second, etc.) used in the description of the specification are merely identifier codes for distinguishing one element from another.

[0034] Further, in the present disclosure, it will be understood that when elements are “connected” or “coupled” to each other, the elements may be directly connected or coupled to each other but may also be connected or coupled to each other through intervening elements unless otherwise specified.

[0035] In the present disclosure, for an element represented as a "device", "unit" or "module", two or more elements may be combined into one element, or one element may be divided into two or more elements according to subdivided functions. In addition, in addition to its own main function, each element described below may additionally perform part or all of the functions performed by another element, and some of the main functions of each element may be completely performed by another component.

[0036] In the present disclosure, an “image” or a “picture” may mean a still image (or a frame), a moving image including a plurality of consecutive still images, or a video.

[0037] In the present disclosure, "neural network" is a representative example of an artificial neural network model that simulates brain nerves, and is not limited to an artificial neural network model that uses a specific algorithm. A neural network may also be referred to as a deep neural network.

[0038] In the present disclosure, a "parameter" is a value used in the operation process of each layer forming a neural network, for example, a "parameter" can be used when an input value is applied to a certain operation expression. The parameter is a value set as a training result and can be updated by separate training data when necessary.

[0039] In the present disclosure, a "sample" is data assigned to a sampling location in one-dimensional (1D) or two-dimensional (2D) data (such as an image, a block, or feature data), and represents data to be processed. For example, a sample may include a pixel in a 2D image. 2D data may be referred to as a "map".

[0040] In addition, in the present disclosure, "current block" means a block to be currently processed. The current block may be a slice obtained by dividing a current image, a tile, a maximum codec unit, a codec unit, a prediction unit, or a transform unit.

[0041] Before describing an image decoding method, an image decoding device, an image encoding method, and an image encoding device according to an embodiment of the present disclosure, reference will be made to Figure 1 and Figure 2 Describe the image encoding and decoding process.

[0042] Figure 1 is a diagram showing the image encoding and decoding process.

[0043] The encoding device 110 transmits a bit stream generated by encoding an image to the decoding device 150, and the decoding device 150 reconstructs the image by receiving and decoding the bit stream.

[0044] In detail, in the encoding device 110, the prediction encoder 115 outputs a prediction block through inter-frame prediction and intra-frame prediction, and the transformer and quantizer 120 outputs a quantized transformation coefficient by transforming and quantizing a residual sample of a residual block between the prediction block and the current block. The entropy encoder 125 encodes the quantized transformation coefficient and outputs it as a bit stream.

[0045] The quantized transform coefficients are reconstructed into a residual block including residual samples in the spatial domain through the inverse quantizer and inverse transformer 130. The reconstructed block combining the prediction block and the residual block is output as a filter block through the deblocking filter 135 and the loop filter 140. The reconstructed image including the filter block can be used as a reference image for the next input image in the prediction encoder 115.

[0046] The bit stream received by the decoding device 150 is reconstructed into a residual block including residual samples in the spatial domain through the entropy decoder 155 and the inverse quantizer and inverse transformer 160. The residual block is generated, the residual block and the prediction block output from the prediction decoder 175 are combined, and the residual block is output as a filter block through the deblocking filter 165 and the loop filter 170. The reconstructed image including the filter block can be used as a reference image of the next image in the prediction decoder 175.

[0047] The loop filter 140 of the encoding device 110 performs loop filtering by using filter information input according to user input or system setting. The filter information used by the loop filter 140 is transmitted to the decoding device 150 through the entropy encoder 125. The loop filter 170 of the decoding device 150 may perform loop filtering based on the filter information input from the entropy decoder 155.

[0048] In the image encoding and decoding process, the image is hierarchically divided, and encoding and decoding are performed on the blocks obtained by dividing the image. Figure 2 Describes the blocks obtained by segmenting an image.

[0049] Figure 2 is a diagram showing blocks obtained by dividing an image according to a tree structure.

[0050] An image 200 may be divided into one or more strips or one or more tiles. A strip may include a plurality of tiles.

[0051] A slice or a tile can be a sequence of one or more largest codec units (CUs).

[0052] A maximum CU may be split into one or more CUs. A CU may be a reference block for determining a prediction mode. In other words, it may be determined whether an intra prediction mode or an inter prediction mode is applied to each CU. In the present disclosure, a maximum CU may be referred to as a maximum codec block, and a CU may be referred to as a codec block.

[0053] The size of a CU may be equal to or smaller than the size of a maximum CU. A maximum CU is a CU having a maximum size and thus may be referred to as a CU.

[0054] One or more prediction units for intra prediction or inter prediction may be determined from a CU. The size of a prediction unit may be equal to or smaller than the size of a CU.

[0055] In addition, one or more transform units for transform and quantization may be determined from the CU. The size of the transform unit may be equal to or smaller than the size of the CU. The transform unit is a reference block for transform and quantization, and the residual samples of the CU may be transformed and quantized for each transform unit in the CU.

[0056] In the present disclosure, the current block may be a slice, a tile, a maximum CU, a CU, a prediction unit, or a transform unit obtained by dividing the image 200. In addition, a lower layer block of the current block is a block obtained by dividing the current block, for example, when the current block is a maximum CU, the lower layer block may be a CU, a prediction unit, or a transform unit. In addition, an upper layer block of the current block is a block including the current block as a part, for example, when the current block is a maximum CU, the upper layer block may be a picture sequence, a picture, a slice, or a tile.

[0057] In the following, reference will be made to Figures 3 to 18 An artificial intelligence (AI)-based video decoding method, an AI-based video decoding device, an AI-based video encoding method, and an AI-based video encoding device according to an embodiment of the present disclosure are described.

[0058] Figures 3 to 5 involves a linear transformation using a transformation kernel trained via a neural network, Figures 6 to 8 A nonlinear transformation involving outputting the result obtained by performing transformation and inverse transformation through a neural network.

[0059] Figure 3 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0060] refer to Figure 3, a transformation 315 is applied to a residual block 301 of the current block. The residual block 301 represents the difference between the original block of the current block and the predicted block 303 of the current block. The predicted block 303 may be obtained by intra prediction and / or inter prediction. As part of the encoding process, a transformation 315 is performed on the residual block 301. The transform kernel generation neural network 310 is used to obtain a transform kernel for performing the transform 315 on the residual block 301. The neighboring pixels 302 (i.e., reference pixels) of the current block, the predicted block 303 of the current block, and the encoding and decoding context information 304 are input to the transform kernel generation neural network 310, and the transform kernel 311 is output from the transform kernel generation neural network 310. The transform block 320 of the residual block 301 is obtained by performing matrix multiplication on the residual block 301 and the transform kernel 311. The transform block 320 is quantized and entropy encoded, and is sent to the decoding side as a bit stream.

[0061] In the decoding process, the transform block 320 obtained from the bitstream is entropy decoded and inverse quantized, and then inverse transform 325 is performed thereon. The inverse transform kernel generation neural network 330 is used to obtain the inverse transform kernel of the inverse transform 325. The neighboring pixels 302 (i.e., reference pixels) of the current block, the prediction block 303 of the current block, and the encoding and decoding context information 304 are input to the inverse transform kernel generation neural network 330, and the inverse transform kernel 331 is output from the inverse transform kernel generation neural network 330. The residual block 335 is obtained by performing matrix multiplication on the inverse quantized residual block and the inverse transform kernel 331. The reconstructed block 345 of the current block is obtained by performing addition 340 on the residual block 335 and the prediction block 303.

[0062] pass Figure 3 The AI-based image encoding and decoding process can use transform kernels directly trained by neural networks using neighboring pixels, prediction blocks, and codec context information, instead of using fixed kernels of codec standards (e.g., discrete cosine transform (DCT) type or discrete sine transform (DST) type) that are not applicable to various block-related technologies.

[0063] The transform kernel generating neural network 310 and the inverse transform kernel generating neural network 330 may be referred to as a forward kernel generating network and a backward kernel generating network, respectively. The transform 315 and the inverse transform 325 may be referred to as a forward transform and a backward transform, respectively. The combination of the transform kernel generating neural network 310 and the inverse transform kernel generating neural network 330 may adaptively learn a convolution kernel specific to a given task, rather than providing a fixed and predetermined convolution kernel. In addition, the forward kernel generating network and the reverse kernel generating network may be implemented using a convolutional neural network, a recursive neural network, or any other type of neural network structure.

[0064] Furthermore, by using a neural network for training, the transform kernel can be trained so that the cost between accuracy and bit rate is well balanced, wherein the cost accuracy can guarantee the accuracy of the reconstructed block.

[0065] Figures 3 to 8 The coding context information used in the decoding may include a quantization parameter of the current block, a partition tree structure of the current block, a partition structure of adjacent pixels, a partition type of the current block, and a partition type of adjacent pixels.

[0066] Furthermore, the codec context information may include context about the strength of the compression level to balance bitrate and quality, and context about the current codec state to provide statistical information of the residual block.

[0067] A dense kernel can be used as both the transform kernel and the inverse transform kernel for efficient transformation in terms of rate-distortion.

[0068] In detail, at the encoding side, when the size of the residual block 301 is MxN, the transform kernel 311 outputted from the transform kernel generating neural network 310 by inputting the neighboring pixels 302 of the current block, the prediction block 303 of the current block, and the encoding and decoding context information 304 is MNxMN. The residual block 301 can be transformed into the form of a vector and rearranged in the form of MNx1 for matrix multiplication of the transform kernel 311 and the residual block 301. The transform kernel 311 of MNxMN and the residual block 301 of MNx1 are obtained by M 2 N 2 The multiplication outputs a transform block 320 in the form of a vector including transform coefficients of MNx1. The transform block 320 is quantized and entropy encoded, and is sent to the decoding side as a bit stream. At the decoding side, the transform block 320 obtained from the bit stream is entropy decoded and inverse quantized. By inputting the neighboring pixels 302 of the current block, the prediction block 303 of the current block, and the encoding and decoding context information 304, the inverse transform kernel 331 output from the inverse transform kernel generating neural network 330 is MNxMN. The residual block 335 on which the inverse transform 325 of MNx1 is performed is obtained by performing M on the inverse transform kernel 331 of MNxMN and the transform block 320 in the form of a vector including transform coefficients of MNx1. 2 N 2 The MNx1 residual block 335 is rearranged back to the form of an MxN block. The reconstructed block 345 of the MxN current block is obtained by performing an addition 340 on the MxN residual block 335 and the MxN prediction block 303.

[0069] Furthermore, separable transform kernels (eg, Kronecker kernels) may be used as transform kernels and inverse transform kernels for computationally efficient transforms.

[0070] In detail, at the encoding side, when the size of the residual block 301 is MxN, the transform kernel 311 outputted from the transform kernel generation neural network 310 by inputting the neighboring pixels 302 of the current block, the prediction block 303 of the current block, and the encoding and decoding context information 304 includes two transform kernels, i.e., an MxM left transform kernel and an NxN right transform kernel. For the transform, matrix multiplication is performed on the MxM left transform kernel, the MxN residual block 301, and the NxN right transform kernel. In this case, unlike the case of using one transform kernel, the MxM left transform kernel, the MxN right transform kernel, and the NxN left transform kernel are multiplied. 2 times multiplication and N 2 times instead of M 2 N 2 Multiplications are performed, so the scale of multiplication is relatively small. Therefore, the case of using two transform kernels is efficient in terms of calculation. Through matrix multiplication, an MxN transform block 320 is obtained. The transform block 320 is quantized and entropy encoded, and is sent to the decoding side as a bit stream. On the decoding side, the transform block 320 obtained from the bit stream is entropy decoded and inverse quantized. By inputting the adjacent pixels 302 of the current block, the prediction block 303 of the current block and the encoding and decoding context information 304, the inverse transform kernel 331 outputted from the inverse transform kernel generating neural network 330 includes two inverse transform kernels, that is, the left inverse transform kernel of MxM and the right inverse transform kernel of NxN. For inverse transform, matrix multiplication is performed on the left inverse transform kernel of MxM, the transform block 320 of MxN and the right inverse transform kernel of NxN. Through matrix multiplication, an MxN residual block 335 on which an inverse transform 325 is performed is obtained. A reconstructed block 345 of the MxN current block is obtained by performing addition 340 on the MxN residual block 335 and the MxN prediction block 303 .

[0071] Furthermore, one transform core may be used on the encoding side and two separable transform cores may be used on the decoding side.

[0072] Alternatively, two separable transform cores may be used on the encoding side and one transform core may be used on the decoding side.

[0073] refer to Figure 3 The calculation method described in the following based on the block size can be applied in the same way. Figure 4 and Figure 5 .

[0074] The following will refer to Fig.17 Describe the training Figure 3 The neural network method used in

[0075] Will refer to Figure 4 and Figure 5 A method of using a transform kernel trained by a neural network and one of a plurality of fixed transform kernels used together in related technical standards is described.

[0076] Figure 4is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0077] refer to Figure 4 , a transform 415 is applied to a residual block 401 of the current block. The residual block 401 represents the difference between the original block of the current block and the predicted block 403 of the current block. As part of the encoding process, a transform 415 is performed on the residual block 401. The transform kernel generation neural network 410 is used to obtain a transform kernel of the transform 415 of the residual block 401. The neighboring pixels 402 (i.e., reference pixels) of the current block, the predicted block 403 of the current block, and the encoding and decoding context information 404 are input to the transform kernel generation neural network 410, and the transform kernel 411 is output from the transform kernel generation neural network 410. A transform block 420 of the residual block is obtained by performing matrix multiplication on the residual block 401 and the transform kernel 411. The transform block 420 is quantized and entropy encoded, and is sent to the decoding side as a bitstream.

[0078] In the decoding process, the transform block 420 obtained from the bitstream is entropy decoded and inversely quantized, and then inversely transformed 425 is performed on it. The linear inverse transform kernel 430 is used for inverse transform 425 of the inversely quantized residual block. The linear inverse transform kernel 430 may be one of a plurality of fixed transform kernels (such as DCT type, DST type, etc.) used in the codec standard of the related art. The residual block 435 on which the inverse transform 425 is performed is obtained by performing matrix multiplication on the inversely quantized residual block and the linear inverse transform kernel 430. The reconstructed block 445 of the current block is obtained by performing addition 440 on the residual block 435 and the prediction block 403.

[0079] Figure 5 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0080] refer to Figure 5 , a transformation 515 is applied to a residual block 501 of the current block. The residual block 501 represents the difference between the original block of the current block and the predicted block 503 of the current block. As part of the encoding process, a transformation 515 is performed on the residual block 501. A linear transformation kernel 510 is used for the transformation 515 of the residual block 501. The linear transformation kernel 510 may be one of a plurality of fixed transformation kernels (such as a DCT type, a DST type, etc.) used in the codec standard of the related art. A transformation block 520 of the residual block 501 on which the transformation 515 is performed is obtained by performing matrix multiplication on the residual block 501 and the linear transformation kernel 510. The transformation block 520 is quantized and entropy encoded, and is sent to the decoding side as a bit stream.

[0081] In the decoding process, the transform block 520 obtained from the bitstream is entropy decoded and inversely quantized, and then inverse transform 525 is performed thereon. The inverse transform kernel generation neural network 530 is used to obtain the inverse transform kernel of the inverse transform 525. The neighboring pixels 502 (i.e., reference pixels) of the current block, the prediction block 503 of the current block, and the encoding and decoding context information 504 are input to the inverse transform kernel generation neural network 530, and the inverse transform kernel 531 is output from the inverse transform kernel generation neural network 530. The residual block 535 on which the inverse transform 525 is performed is obtained by performing matrix multiplication on the inversely quantized residual block and the inverse transform kernel 531. The reconstructed block 545 of the current block is obtained by performing addition 540 on the residual block 535 and the prediction block 503.

[0082] Figure 6 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0083] refer to Figure 6 , during the encoding process, a transform is applied to a residual block 601 of a current block. The residual block 601 represents the difference between the original block of the current block and the predicted block 603 of the current block. The transform neural network 615 and the codec context neural network 610 are used for the transform of the residual block 601. The neighboring pixels 602 (i.e., reference pixels) of the current block, the predicted block 603 of the current block, and the codec context information 604 are input to the codec context neural network 610, and the codec context feature map 611 is output from the codec context neural network 610. When the codec context feature map 611 and the residual block 601 are input to the transform neural network 615, a transform feature map 620 is obtained. The transform feature map 620 is quantized and entropy encoded, and is sent to the decoding side as a bit stream.

[0084] In the decoding process, the transform feature map 620 obtained from the bitstream is entropy decoded and inverse quantized. The inverse transform neural network 625 and the codec context neural network 630 are used for inverse transform. The neighboring pixels 602 (ie, reference pixels) of the current block, the prediction block 603 of the current block, and the codec context information 604 are input to the codec context neural network 630, and the codec context feature map 631 is output from the codec context neural network 630. When the inverse quantized transform feature map 620 and the codec context feature map 631 are input to the inverse transform neural network 625, an inversely transformed residual block 635 is obtained. The reconstructed block 645 of the current block is obtained by performing addition 640 on the residual block 635 and the prediction block 603.

[0085] In detail, at the encoding side, the size of the residual block 601 is MxN. By inputting the neighboring pixels 602 of the current block, the prediction block 603 of the current block, and the codec context information 604, the size of the codec context feature map 611 for transformation output from the codec context neural network 610 is M1xN1xC1. The codec context feature map 611 and the residual block 601 are input to the transform neural network 615, and the transform neural network 615 outputs a transform feature map 620 of the transform coefficient of the residual block 601, which has a size of M2xN2xC2. The transform feature map 620 is quantized and entropy encoded, and is sent to the decoding side as a bit stream. At the decoding side, the transform feature map 620 obtained from the bit stream is entropy decoded and inverse quantized. By inputting the neighboring pixels 602 of the current block, the prediction block 603 of the current block, and the codec context information 604, the codec context feature map 631 for inverse transformation output from the codec context neural network 630 is M3xN3xC3. When the inverse quantized transform feature map 620 and the encoding and decoding context feature map 631 are input to the inverse transform neural network 625, an inverse transform residual block 635 of size MxN is obtained. A reconstructed block 645 of size MxN is obtained by performing addition 640 on the residual block 635 of size MxN and the prediction block 603 of size MxN. Here, M, M1, M2, and M3 may be different and have different values, N, N1, N2, and N3 may be different and have different values, and C1, C2, and C3 may be different and have different values.

[0086] The transform feature map 620 output from the transform neural network 615 is transmitted as a bit stream, so its size needs to be limited. Therefore, the transform neural network 615 is a neural network trained to output a transform feature map 620 whose size is smaller than that of the input information to reduce the bit rate, and the inverse transform neural network 625 is a neural network trained to output a residual block 635 by reconstructing data from the input transform feature map 620.

[0087] The codec context neural network 610 for transformation may be a neural network for outputting information required for transformation from neighboring pixels 602 of the current block, a prediction block 603 of the current block, and codec context information 604 in the form of a feature map, and the codec context neural network 630 for inverse transformation may be a neural network for outputting information required for inverse transformation from neighboring pixels 602 of the current block, a prediction block 603 of the current block, and codec context information 604 in the form of a feature map.

[0088] In addition, the codec context neural network 610 for transformation can send the neighboring pixels 602 of the current block, the prediction block 603 of the current block, and part of the information in the codec context information 604 without performing any processing for input to the transformation neural network 615, and the codec context neural network 630 for inverse transformation can send the neighboring pixels 602 of the current block, the prediction block 603 of the current block, and part of the information in the codec context information 604 without performing any processing for input to the inverse transformation neural network 625.

[0089] In addition, the output of the transform neural network 615 may be a transform feature map 620 of a transform coefficient that is quantized after transformation, and the output of the inverse transform neural network 625 may be an inversely transformed residual block 635 after inverse quantization. In other words, the transform neural network 615 may be a neural network in which transformation and quantization are performed together, and the inverse transform neural network 625 may be a neural network in which inverse quantization and inverse transformation are performed together.

[0090] In detail, at the encoding side, the size of the residual block 601 is MxN, and by inputting the neighboring pixels 602 of the current block, the prediction block 603 of the current block, and the codec context information 604, the codec context feature map 611 for transformation output from the codec context neural network 610 is M1xN1xC1. The codec context feature map 611 and the residual block 601 are input to the transform neural network 615, and the transform feature map 620 of the quantized transform coefficient of the residual block 601 of M2xN2xC2 is obtained. The transform feature map 620 is entropy encoded and sent to the decoding side as a bit stream. At the decoding side, the transform feature map 620 obtained from the bit stream is entropy decoded. By inputting the neighboring pixels 602 of the current block, the prediction block 603 of the current block, and the codec context information 604, the codec context feature map 631 for inverse transformation output from the codec context neural network 630 is M3xN3xC3. When the entropy-decoded transform feature map 620 and the codec context feature map 631 are input to the inverse transform neural network 625, an inverse quantized and inverse transformed residual block 635 of size MxN is obtained. A reconstructed block 645 of size MxN is obtained by performing addition 640 on the residual block 635 of size MxN and the prediction block 603 of size MxN. Here, M, M1, M2, and M3 may be different and have different values, N, N1, N2, and N3 may be different and have different values, and C1, C2, and C3 may be different and have different values.

[0091] refer to Figure 6 The calculation method described in the following based on the block size can be applied in the same way. Figure 7 and Figure 8 .

[0092] The following will refer to Fig.18 Describe the training Figure 6 The neural network method used in

[0093] Figure 7 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0094] refer to Figure 7 , a transform is applied to a residual block 701 of the current block. The residual block 701 represents the difference between the original block of the current block and the predicted block 703 of the current block. As part of the encoding process, a transform is performed on the residual block 701. A transform neural network 715 and a codec context neural network 710 are used for the transform of the residual block 701. Neighboring pixels 702 (i.e., reference pixels) of the current block, a predicted block 703 of the current block, and codec context information 704 are input to the codec context neural network 710, and a codec context feature map 711 is output from the codec context neural network 710. The codec context feature map 711 and the residual block 701 are input to the transform neural network 715, and a transform feature map 720 of the transform coefficients of the residual block 701 is obtained. The transform feature map 720 is quantized and entropy encoded, and is sent to the decoding side as a bitstream.

[0095] In the decoding process, the transform feature map 720 obtained from the bitstream is entropy decoded and inverse quantized. The inverse transform neural network 725 and the codec context neural network 730 are used for inverse transform. The neighboring pixels 702 (ie, reference pixels) of the current block, the prediction block 703 of the current block, and the codec context information 704 are input to the codec context neural network 730, and the codec context feature map 731 is output from the codec context neural network 730. When the inverse quantized transform feature map and the codec context feature map 731 are input to the inverse transform neural network 725, a reconstructed block 745 of the current block is obtained.

[0096] The transformed feature map 720 output from the transform neural network 715 is transmitted as a bit stream, so its size needs to be limited. Therefore, the transform neural network 715 is a neural network trained to output a transform feature map 720 whose size is smaller than that of the input information to reduce the bit rate, and the inverse transform neural network 725 is a neural network trained to output a reconstructed block 745 by reconstructing data from the input transform feature map 720.

[0097] The codec context neural network 710 for transformation may be a neural network for outputting information required for transformation from neighboring pixels 702 of the current block, a prediction block 703 of the current block, and codec context information 704 in the form of a feature map, and the codec context neural network 730 for inverse transformation may be a neural network for outputting information required for inverse transformation from neighboring pixels 702 of the current block, a prediction block 703 of the current block, and codec context information 704 in the form of a feature map.

[0098] In addition, the codec context neural network 710 for transformation can send the neighboring pixels 702 of the current block, the prediction block 703 of the current block, and part of the information in the codec context information 704 without performing any processing for input to the transformation neural network 715, and the codec context neural network 730 for inverse transformation can send the neighboring pixels 702 of the current block, the prediction block 703 of the current block, and part of the information in the codec context information 704 without performing any processing for input to the inverse transformation neural network 725.

[0099] In addition, the output of the transform neural network 715 may be a transform feature map 720 of a transform coefficient that is quantized after transformation, and the output of the inverse transform neural network 725 may be a reconstructed block 745 that is inversely transformed after inverse quantization. In other words, the transform neural network 715 may be a neural network in which transformation and quantization are performed together, and the inverse transform neural network 725 may be a neural network in which inverse quantization and inverse transformation are performed together.

[0100] In detail, the residual block 701 of the current block, which is the difference between the original block of the current block and the prediction block 703 of the current block, is the transformation target in the encoding process. The transformation neural network 715 and the codec context neural network 710 are used for the transformation of the residual block 701. The neighboring pixels 702 (ie, reference pixels) of the current block, the prediction block 703 of the current block, and the codec context information 704 are input to the codec context neural network 710, and the codec context feature map 711 is output from the codec context neural network 710. The codec context feature map 711 and the residual block 701 are input to the transformation neural network 715, and the transformation feature map 720 of the quantized transformation coefficients of the residual block 701 is obtained. The transformation feature map 720 is entropy encoded and sent to the decoding side as a bit stream.

[0101] In the decoding process, the transform feature map 720 obtained from the bitstream is entropy decoded. The inverse transform neural network 725 and the codec context neural network 730 are used for inverse transform. The neighboring pixels 702 (ie, reference pixels) of the current block, the prediction block 703 of the current block, and the codec context information 704 are input to the codec context neural network 730, and the codec context feature map 731 is output from the codec context neural network 730. When the entropy decoded transform feature map and the codec context feature map 731 are input to the inverse transform neural network 725, a reconstructed block 745 of the current block is obtained.

[0102] Figure 8 is a diagram describing an AI-based image encoding and decoding process according to an embodiment of the present disclosure.

[0103] refer to Figure 8 , a transform is applied to a residual block 801 of the current block. The residual block 801 represents the difference between the original block of the current block and the predicted block 803 of the current block. As part of the encoding process, a transform is performed on the residual block 801. The transform neural network 815 and the codec context neural network 810 are used for the transform of the residual block 801. The neighboring pixels 802 (i.e., reference pixels) of the current block, the predicted block 803 of the current block, and the codec context information 804 are input to the codec context neural network 810, and the codec context feature map 811 is output from the codec context neural network 810. When the codec context feature map 811 and the residual block 801 are input to the transform neural network 815, a transform feature map 820 is obtained. The transform feature map 820 is quantized and entropy encoded, and is sent to the decoding side as a bitstream.

[0104] During the decoding process, the transform feature map 820 obtained from the bitstream is entropy decoded and inverse quantized. The inverse transform neural network 825 and the codec context neural network 830 are used for inverse transform. The neighboring pixels 802 (ie, reference pixels) of the current block, the prediction block 803 of the current block, and the codec context information 804 are input to the codec context neural network 830, and the codec context feature map 831 is output from the codec context neural network 830. The inversely quantized transform feature map and the codec context feature map 831 are input to the inverse transform neural network 825, and an extended reconstructed block 845 including a reconstructed block of the current block and reference pixels of the current block is obtained.

[0105] Obtaining the extended reconstructed block 845 including the reconstructed block of the current block and the reference pixels of the current block can assist the deblocking filtering process. In other words, the result of the deblocking filtering can be improved.

[0106] The transform feature map 820 output from the transform neural network 815 is transmitted as a bit stream, so its size needs to be limited. Therefore, the transform neural network 815 is a neural network trained to output a transform feature map 820 whose size is smaller than that of the input information to reduce the bit rate, and the inverse transform neural network 825 is a neural network trained to output an extended reconstructed block 845 including a reconstructed block of the current block and reference pixels of the current block by reconstructing data from the input transform feature map 820.

[0107] The codec context neural network 810 for transformation may be a neural network for outputting information required for transformation from neighboring pixels 802 of the current block, a prediction block 803 of the current block, and codec context information 804 in the form of a feature map, and the codec context neural network 830 for inverse transformation may be a neural network for outputting information required for inverse transformation from neighboring pixels 802 of the current block, a prediction block 803 of the current block, and codec context information 804 in the form of a feature map.

[0108] In addition, the codec context neural network 80 for transformation can send the neighboring pixels 802 of the current block, the prediction block 803 of the current block, and part of the information in the codec context information 804 without any processing for input to the transformation neural network 815, and the codec context neural network 830 for inverse transformation can send the neighboring pixels 802 of the current block, the prediction block 803 of the current block, and part of the information in the codec context information 804 without any processing for input to the inverse transformation neural network 825.

[0109] In addition, the output of the transform neural network 815 may be a transform feature map 820 of a transform coefficient that is quantized after transformation, and the output of the inverse transform neural network 825 may be an inversely transformed extended reconstructed block 845 that is inversely quantized. In other words, the transform neural network 815 may be a neural network in which transformation and quantization are performed together, and the inverse transform neural network 825 may be a neural network in which inverse quantization and inverse transformation are performed together.

[0110] In detail, the residual block 801 of the current block, which is the difference between the original block of the current block and the prediction block 803 of the current block, is the transformation target in the encoding process. The transformation neural network 815 and the codec context neural network 810 are used for the transformation of the residual block 801. The neighboring pixels 802 (ie, reference pixels) of the current block, the prediction block 803 of the current block, and the codec context information 804 are input to the codec context neural network 810, and the codec context feature map 811 is output from the codec context neural network 810. The codec context feature map 811 and the residual block 801 are input to the transformation neural network 815, and the transformation feature map 820 of the quantized transformation coefficients of the residual block 801 is obtained. The transformation feature map 820 is entropy encoded and sent to the decoding side as a bit stream.

[0111] During the decoding process, the transform feature map 820 obtained from the bitstream is entropy decoded. The inverse transform neural network 825 and the codec context neural network 830 are used for inverse transform. The neighboring pixels 802 (ie, reference pixels) of the current block, the prediction block 803 of the current block, and the codec context information 804 are input to the codec context neural network 830, and the codec context feature map 831 is output from the codec context neural network 830. The entropy-decoded transform feature map and the codec context feature map 831 are input to the inverse transform neural network 825, and an extended reconstructed block 845 including a reconstructed block of the current block and reference pixels of the current block is obtained.

[0112] Fig. 9 is a flowchart of an AI-based image encoding method according to an embodiment of the present disclosure.

[0113] refer to Fig. 9 In operation S910, the AI-based image encoding device 1000 obtains a residual block based on a predicted block of a current block and an original block of the current block. The residual block may represent a difference between the original block and the predicted block of the current block. The original block may be a portion of an image that the AI-based image encoding device 1000 intends to encode or decode, and the original block is predicted based on neighboring blocks to estimate the content of the original block. The residual block may be obtained by subtracting the predicted block from the original block to represent the difference between the predicted block and the actual content within the original block.

[0114] In operation S930, the AI-based image encoding device 1000 generates a transform kernel of a transform block of a residual block by applying a prediction block, neighboring pixels of a current block, and encoding and decoding context information to a neural network.

[0115] In operation S950, the AI-based image encoding device 1000 obtains a transformed block by applying the generated transform kernel to the residual block. The transform may be performed to reduce the amount of data required to represent the original block.

[0116] According to an embodiment of the present disclosure, the generated transform kernel may include a left transform kernel to be applied to the left side of the residual block and a right transform kernel to be applied to the right side of the residual block.

[0117] In operation S970, the AI-based image encoding device 1000 generates a bitstream including a transform block.

[0118] According to an embodiment of the present disclosure, during the image decoding process, the transformation block may be inversely transformed by a transformation kernel based on a neural network, or may be inversely transformed by one of a plurality of predetermined linear transformation kernels.

[0119] Fig.10 is a diagram of a configuration of an AI-based image encoding device according to an embodiment of the present disclosure.

[0120] refer to Fig.10 , the AI-based image encoding device 1000 may include a residual block acquirer 1010, a transform kernel generator 1020, a transformer 1030, and a generator 1040.

[0121] The residual block acquirer 1010, the transform core generator 1020, the transformer 1030, and the generator 1040 may be implemented as a processor. The residual block acquirer 1010, the transform core generator 1020, the transformer 1030, and the generator 1040 may operate according to instructions stored in a memory.

[0122] exist Fig.10 , the residual block acquirer 1010, the transform core generator 1020, the transformer 1030, and the generator 1040 are shown separately, but the residual block acquirer 1010, the transform core generator 1020, the transformer 1030, and the generator 1040 can be implemented by one processor. In this case, the residual block acquirer 1010, the transform core generator 1020, the transformer 1030, and the generator 1040 can be implemented as a dedicated processor, or can be implemented by a combination of software and a general-purpose processor (such as an application processor (AP), a central processing unit (CPU), or a graphics processing unit (GPU)). The dedicated processor may include a memory for implementing an embodiment of the present disclosure, or include a memory processor for using an external memory.

[0123] The residual block acquirer 1010, the transform core generator 1020, the transformer 1030, and the generator 1040 may be implemented as a plurality of processors. In this case, the residual block acquirer 1010, the transform core generator 1020, the transformer 1030, and the generator 1040 may be implemented as a combination of dedicated processors, or may be implemented as a combination of software and a plurality of general-purpose processors (such as an AP, a CPU, or a GPU). The processor may include an AI dedicated processor. As another example, the AI ​​dedicated processor may be configured as a chip separate from the processor.

[0124] The residual block acquirer 1010 acquires a residual block based on the prediction block of the current block and the original block of the current block.

[0125] The transform kernel generator 1020 generates a transform kernel of a transform block of a residual block by applying a prediction block, neighboring pixels of a current block, and coding context information to a neural network.

[0126] The transformer 1030 obtains a transformed block by applying the generated transform kernel to the residual block.

[0127] The generator 1040 generates a bitstream including the transformed block.

[0128] The bitstream may be transmitted to the AI-based image decoding device 1200 .

[0129] Fig.11 is a flowchart of an AI-based image decoding method according to an embodiment of the present disclosure.

[0130] refer to Fig.11 , in operation S1110, the AI-based image decoding device 1200 obtains a transform block of a residual block of a current block from a bitstream.

[0131] According to an embodiment of the present disclosure, the transformation block may be a block transformed by a transformation kernel based on a neural network or by one of a plurality of predetermined linear transformation kernels.

[0132] In operation S1130, the AI-based image decoding device 1200 generates a transform kernel of a transform block by inputting a prediction block of the current block, neighboring pixels of the current block, and encoding and decoding context information into a neural network, and by obtaining the transform kernel as an output of the neural network.

[0133] According to an embodiment of the present disclosure, the coding context information may include at least one of a quantization parameter of a current block, a partition tree structure of a current block, a partition structure of adjacent pixels, a partition type of a current block, or a partition type of adjacent pixels.

[0134] In operation S1150 , the AI-based image decoding device 1200 obtains a residual block by applying the generated transform kernel to the transform block.

[0135] According to an embodiment of the present disclosure, the generated transform kernel may include a left transform kernel to be applied to the left side of the transform block and a right transform kernel to be applied to the right side of the transform block.

[0136] In operation S1170 , the AI-based image decoding device 1200 reconstructs the current block by using the residual block and the prediction block.

[0137] Fig.12 is a diagram of a configuration of an AI-based image decoding device according to an embodiment of the present disclosure.

[0138] refer to Fig.12 , the AI-based image decoding device 1200 may include an acquirer 1210 , an inverse transform kernel generator 1220 , an inverse transformer 1230 , and a reconstructor 1240 .

[0139] The acquirer 1210, the inverse transform core generator 1220, the inverse transformer 1230, and the reconstructor 1240 may be implemented as a processor. The acquirer 1210, the inverse transform core generator 1220, the inverse transformer 1230, and the reconstructor 1240 may operate according to instructions stored in a memory.

[0140] exist Fig.12 In the embodiment of the present invention, the acquirer 1210, the inverse transform core generator 1220, the inverse transformer 1230, and the reconstructor 1240 are shown separately, but the acquirer 1210, the inverse transform core generator 1220, the inverse transformer 1230, and the reconstructor 1240 can be implemented by one processor. In this case, the acquirer 1210, the inverse transform core generator 1220, the inverse transformer 1230, and the reconstructor 1240 can be implemented as a dedicated processor, or can be implemented by a combination of software and a general-purpose processor (such as an AP, a CPU, or a GPU). The dedicated processor may include a memory for implementing the embodiments of the present disclosure, or include a memory processor for using an external memory.

[0141] The acquirer 1210, the inverse transform core generator 1220, the inverse transformer 1230, and the reconstructor 1240 may be implemented as a plurality of processors. In this case, the acquirer 1210, the inverse transform core generator 1220, the inverse transformer 1230, and the reconstructor 1240 may be implemented as a combination of dedicated processors, or may be implemented as a combination of software and a plurality of general-purpose processors (such as an AP, a CPU, or a GPU). The processor may include an AI dedicated processor. As another example, the AI ​​dedicated processor may be configured as a chip separate from the processor.

[0142] The acquirer 1210 obtains a transform block of a residual block of a current block from a bitstream.

[0143] A bitstream may be generated by and transmitted from the AI-based image encoding apparatus 1000 .

[0144] The inverse transform kernel generator 1220 generates a transform kernel of a transform block by applying a prediction block, neighboring pixels of a current block, and coding context information to a neural network.

[0145] The inverse transformer 1230 obtains a residual block by applying the generated transform kernel to the transform block.

[0146] The reconstructor 1240 reconstructs the current block by using the residual block and the prediction block.

[0147] Fig.13 is a flowchart of an AI-based image encoding method according to an embodiment of the present disclosure.

[0148] refer to Fig.13 , in operation S1310, the AI-based image encoding device 1400 obtains a residual block based on a prediction block of a current block and an original block of the current block.

[0149] In operation S1330, the AI-based image encoding device 1400 generates a coding context feature map of the transform block by applying the prediction block, neighboring pixels of the current block, and coding context information to a first neural network.

[0150] In operation S1350, the AI-based image encoding device 1400 obtains a transform feature map corresponding to the transform block by inputting the encoding and decoding context feature map and the residual block to the second neural network, and by obtaining the transform feature map as an output of the second neural network.

[0151] According to an embodiment of the present disclosure, the second neural network may output a transform feature map of the quantized transform coefficients.

[0152] In operation S1370, the AI-based image encoding device 1400 generates a bitstream including a transformation feature map.

[0153] Fig.14 is a diagram of a configuration of an AI-based image encoding device according to an embodiment of the present disclosure.

[0154] refer to Fig.14 , the AI-based image encoding device 1400 may include a residual block acquirer 1410, a coding context feature map generator 1420, a transformer 1430 and a generator 1440.

[0155] The residual block acquirer 1410, the coding context feature map generator 1420, the transformer 1430, and the generator 1440 may be implemented as a processor. The residual block acquirer 1410, the coding context feature map generator 1420, the transformer 1430, and the generator 1440 may operate according to instructions stored in a memory.

[0156] exist Fig.14 In the embodiment, the residual block acquirer 1410, the codec context feature map generator 1420, the transformer 1430, and the generator 1440 are shown separately, but the residual block acquirer 1410, the codec context feature map generator 1420, the transformer 1430, and the generator 1440 can be implemented by one processor. In this case, the residual block acquirer 1410, the codec context feature map generator 1420, the transformer 1430, and the generator 1440 can be implemented as a dedicated processor, or can be implemented by a combination of software and a general-purpose processor (such as an AP, a CPU, or a GPU). The dedicated processor may include a memory for implementing the embodiments of the present disclosure, or include a memory processor for using an external memory.

[0157] The residual block acquirer 1410, the codec context feature map generator 1420, the transformer 1430, and the generator 1440 may be implemented as a plurality of processors. In this case, the residual block acquirer 1410, the codec context feature map generator 1420, the transformer 1430, and the generator 1440 may be implemented as a combination of dedicated processors, or may be implemented as a combination of software and a plurality of general-purpose processors (such as an AP, a CPU, or a GPU). The processor may include an AI dedicated processor. As another example, the AI ​​dedicated processor may be configured as a chip separate from the processor.

[0158] The residual block acquirer 1410 acquires a residual block based on the prediction block of the current block and the original block of the current block.

[0159] The codec context feature map generator 1420 generates a codec context feature map of the transform block by applying the prediction block, neighboring pixels of the current block, and codec context information to the first neural network.

[0160] The transformer 1430 obtains a transform feature map corresponding to the transform block by applying the codec context feature map and the residual block to the second neural network.

[0161] The generator 1440 generates a bitstream including the transformed feature map.

[0162] The bitstream may be transmitted to the AI-based image decoding device 1600 .

[0163] Fig.15is a flowchart of an AI-based image decoding method according to an embodiment of the present disclosure.

[0164] refer to Fig.15 , in operation S1510, the AI-based image decoding device 1600 obtains a transform feature map corresponding to a transform block of a residual block of a current block from a bitstream.

[0165] In operation S1530, the AI-based image decoding device 1600 generates a codec context feature map of the transform block by inputting a prediction block of the current block, neighboring pixels of the current block, and codec context information into a first neural network, and by obtaining the codec context feature map as an output of the first neural network.

[0166] In operation S1550, the AI-based image decoding device 1600 reconstructs the current block by inputting the transformation feature map and the coding context feature map to the second neural network, and by obtaining the reconstructed current block as an output of the second neural network.

[0167] According to an embodiment of the present disclosure, the second neural network may output a result value obtained by performing inverse transformation after inverse quantization.

[0168] According to an embodiment of the present disclosure, reconstruction of the current block may include obtaining a residual block by applying a transform feature map and a coding context feature map to a second neural network, and reconstructing the current block by using the residual block and the prediction block.

[0169] According to an embodiment of the present disclosure, the reconstructed current block may further include neighboring pixels of the current block for deblocking filtering of the current block.

[0170] Fig.16 is a diagram of a configuration of an AI-based image decoding device according to an embodiment of the present disclosure.

[0171] refer to Fig.16 , the AI-based image decoding device 1600 may include an acquirer 1610, a coding context feature map generator 1620, an inverse transformer 1630 and a reconstructor 1640.

[0172] The acquirer 1610, the coding context feature map generator 1620, the inverse transformer 1630, and the reconstructor 1640 may be implemented as a processor. The acquirer 1610, the coding context feature map generator 1620, the inverse transformer 1630, and the reconstructor 1640 may operate according to instructions stored in a memory.

[0173] exist Fig.16In the embodiment of the present invention, the acquirer 1610, the codec context feature map generator 1620, the inverse transformer 1630 and the reconstructor 1640 are shown separately, but the acquirer 1610, the codec context feature map generator 1620, the inverse transformer 1630 and the reconstructor 1640 can be implemented by one processor. In this case, the acquirer 1610, the codec context feature map generator 1620, the inverse transformer 1630 and the reconstructor 1640 can be implemented as a dedicated processor, or can be implemented by a combination of software and a general-purpose processor (such as an AP, a CPU or a GPU). The dedicated processor may include a memory for implementing the embodiments of the present disclosure, or include a memory processor for using an external memory.

[0174] The acquirer 1610, the codec context feature map generator 1620, the inverse transformer 1630, and the reconstructor 1640 may be implemented as a plurality of processors. In this case, the acquirer 1610, the codec context feature map generator 1620, the inverse transformer 1630, and the reconstructor 1640 may be implemented as a combination of dedicated processors, or may be implemented as a combination of software and a plurality of general-purpose processors (such as an AP, a CPU, or a GPU). The processor may include an AI dedicated processor. As another example, the AI ​​dedicated processor may be configured as a chip separate from the processor.

[0175] The acquirer 1610 obtains a transform feature map corresponding to a transform block of a residual block of a current block with respect to a bitstream.

[0176] The bitstream may be generated by and transmitted from the AI-based image encoding apparatus 1400 .

[0177] The codec context feature map generator 1620 generates a codec context feature map of the transform block by applying the prediction block of the current block, the neighboring pixels of the current block, and the codec context information to the first neural network.

[0178] The inverse transformer 1630 obtains a residual block by applying the transformed feature map and the encoding and decoding context feature map to the second neural network.

[0179] The reconstructor 1640 obtains a reconstructed block by using the residual block and the prediction block.

[0180] According to an embodiment of the present disclosure, the inverse transformer 1630 may obtain a reconstructed block by inputting the transform feature map and the encoding and decoding context feature map into the second neural network. In this case, the reconstructor 1640 may be omitted in the AI-based image decoding device 1600.

[0181] According to an embodiment of the present disclosure, the inverse transformer 1630 can obtain a reconstructed block including a current block and an extended reconstructed block of adjacent pixels of the current block by inputting the transform feature map and the codec context feature map into the second neural network for deblocking filtering of the current block. In this case, the reconstructor 1640 can be omitted in the AI-based image decoding device 1600.

[0182] Fig.17 is a diagram for describing a method of training a neural network used in an AI-based image encoding method and an AI-based image decoding method according to an embodiment of the present disclosure.

[0183] refer to Fig.17 , the transform kernel generating neural network 1710 and the inverse transform kernel generating neural network 1730 can be trained by using the training original block 1700, the training residual block 1701, the training adjacent pixels 1702, the training prediction block 1703 and the training codec context information 1704.

[0184] In detail, when the training neighboring pixels 1702, the training prediction block 1703, and the training codec context information 1704 are input to the transform kernel generation neural network 1710, a training transform kernel 1711 is generated. The training transform block 1720 is obtained by performing a transform 1715 using the training residual block 1701 and the training transform kernel 1711. The training transform block 1720 is quantized and entropy encoded, and is transmitted in the form of a bit stream.

[0185] In addition, the training transform block 1720 is entropy decoded and inverse quantized. When the training neighboring pixels 1702, the training prediction block 1703, and the training codec context information 1704 are input to the inverse transform kernel generation neural network 1730, a training inverse transform kernel 1731 is generated. A training inverse transform residual block 1735 is obtained by performing an inverse transform 1725 using the training transform block 1720 and the training inverse transform kernel 1731. A training reconstructed block 1745 is obtained by performing an addition 1740 on the training inverse transform residual block 1735 and the training prediction block 1703.

[0186] exist Fig.17 During the training process, the neural network can be trained so that the training reconstructed block 1745 is as similar as possible to the training original block 1700 by comparison 1755, and the bit rate of the bit stream generated by encoding the training transformed block 1720 is minimized. In this regard, Fig.17 As shown, the first loss information 1750 and the second loss information 1760 may be used when training a neural network.

[0187] The second loss information 1760 may correspond to the difference between the training original block 1700 and the training reconstructed block 1745. According to an embodiment of the present disclosure, the difference between the training original block 1700 and the training reconstructed block 1745 may include at least one of an L1 norm value, an L2 norm value, a structural similarity (SSIM) value, a peak signal-to-noise ratio-human visual system (PSNR-HVS) value, a multi-scale SSIM (MS-SSIM) value, a variance inflation factor (VIF) value, or a video multi-method assessment fusion (VMAF) value between the training original block 1700 and the training reconstructed block 1745.

[0188] The second loss information 1760 indicates the quality of the reconstructed image including the training reconstruction block 1745 and thus may be referred to as quality loss information.

[0189] The first loss information 1750 may be calculated according to a bit rate of a bitstream generated as a result of encoding the training transform block 1720. For example, the first loss information 1750 may be calculated based on a bit rate difference between the training residual block 1701 and the training transform block 1720.

[0190] The first loss information 1750 indicates the encoding efficiency of the training transform block 1720 and thus may be referred to as compression loss information.

[0191] The transform kernel generating neural network 1710 and the inverse transform kernel generating neural network 1730 may be trained so that final loss information derived from any one or a combination of the first loss information 1750 and the second loss information 1760 is reduced or minimized.

[0192] According to an embodiment of the present disclosure, the transform kernel generating neural network 1710 and the inverse transform kernel generating neural network 1730 can reduce or minimize the final loss information while changing the values ​​of preset parameters.

[0193] According to an embodiment of the present disclosure, the final loss information may be calculated according to Equation 1 below.

[0194] [Equation 1]

[0195] Final loss information = a x first loss information + b x second loss information

[0196] In Equation 1, a and b are weights applied to the first loss information 1750 and the second loss information 1760, respectively.

[0197] According to Equation 1, it is determined that the transform kernel generating neural network 1710 and the inverse transform kernel generating neural network 1730 are trained so that the training reconstructed block 1745 becomes as similar as possible to the training original block 1700 and the size of the bitstream is minimized.

[0198] Fig.17 The transformation kernel generating neural network 1710 and the inverse transformation kernel generating neural network 1730 can correspond to Figure 3 The transformation kernel generates the neural network 310 and the inverse transformation kernel generates the neural network 330.

[0199] exist Fig.17 During the training method of the present invention, in addition to the inverse transform kernel generating neural network 1730, it is possible to train the inverse transform kernel 1731 by using the linear inverse transform kernel of the related art instead of the training inverse transform kernel 1731. Figure 4 The transformation kernel generates a neural network 410.

[0200] also, Figure 4 The transformation kernel generating neural network 410 may correspond to Fig.17 The transformation kernel generates neural network 1710.

[0201] exist Fig.17 During the training method of the present invention, in addition to the transformation kernel generating neural network 1710, it is possible to train the neural network by using a linear transformation kernel of the related art instead of the training transformation kernel 1711. Figure 5 The inverse transform kernel generates the neural network 530.

[0202] also, Figure 5 The inverse transform kernel generation neural network 530 may correspond to Fig.17 The inverse transform kernel generates the neural network 1730.

[0203] Fig.18 is a diagram describing a method of training a neural network used in an AI-based image encoding method and an AI-based image decoding method according to an embodiment of the present disclosure.

[0204] refer to Fig.18 , the codec context neural network 1810, the transform neural network 1815, the inverse transform neural network 1825 and the codec context neural network 1830 can be trained by using the training original block 1800, the training residual block 1801, the training adjacent pixels 1802, the training prediction block 1803 and the training codec context information 1804.

[0205] In detail, when the training neighboring pixels 1802, the training prediction block 1803, and the training codec context information 1804 are input to the codec context neural network 1810, a training codec context feature map 1811 is generated. By inputting the training residual block 1801 and the training codec context feature map 1811 to the transformation neural network 1815, a training transformation feature map 1820 is obtained. The training transformation feature map 1820 is quantized and entropy encoded, and is transmitted in the form of a bit stream.

[0206] In addition, the training transform feature map 1820 is entropy decoded and inverse quantized. When the training neighboring pixels 1802, the training prediction block 1803, and the training codec context information 1804 are input to the codec context neural network 1830, a training codec context feature map 1831 is generated. By applying the training transform feature map 1820 and the training codec context feature map 1831 to the inverse transform neural network 1825, a training inverse transform residual block 1835 is obtained. A training reconstruction block 1845 is obtained by performing addition 1840 on the training inverse transform residual block 1835 and the training prediction block 1803.

[0207] exist Fig.18 During the training process, the neural network can be trained so that the training reconstructed block 1845 is as similar as possible to the training original block 1800 by comparison 1855, and the bit rate of the bit stream generated by encoding the training transformed feature map 1820 is minimized. In this regard, Fig.18 As shown, the first loss information 1850 and the second loss information 1860 may be used when training a neural network.

[0208] The second loss information 1860 may correspond to the difference between the training original block 1800 and the training reconstructed block 1845. According to an embodiment of the present disclosure, the difference between the training original block 1800 and the training reconstructed block 1845 may include at least one of an L1 norm value, an L2 norm value, a structural similarity (SSIM) value, a peak signal-to-noise ratio-human visual system (PSNR-HVS) value, a multi-scale SSIM (MS-SSIM) value, a variance inflation factor (VIF) value, or a video multi-method assessment fusion (VMAF) value between the training original block 1800 and the training reconstructed block 1845.

[0209] The second loss information 1860 is related to the quality of the reconstructed image including the training reconstruction block 1845, and thus may be referred to as quality loss information.

[0210] The first loss information 1850 may be calculated according to the bit rate of the bitstream generated as a result of encoding the training transform feature map 1820. For example, the first loss information 1850 may be calculated based on the bit rate difference between the training residual block 1801 and the training transform block 1820.

[0211] The first loss information 1850 is related to the encoding efficiency of the training transformed feature map 1820, and thus can be referred to as compression loss information.

[0212] The codec context neural network 1810, the transform neural network 1815, the inverse transform neural network 1825, and the codec context neural network 1830 may be trained so that final loss information derived from any one or a combination of the first loss information 1850 and the second loss information 1860 is reduced or minimized.

[0213] According to an embodiment of the present disclosure, the encoding and decoding context neural network 1810, the transformation neural network 1815, the inverse transformation neural network 1825, and the encoding and decoding context neural network 1830 can reduce or minimize the final loss information while changing the values ​​of preset parameters.

[0214] According to an embodiment of the present disclosure, the final loss information may be calculated according to Equation 2 below.

[0215] [Equation 2]

[0216] Final loss information = a x first loss information + b x second loss information

[0217] In Equation 2, a and b are weights applied to the first loss information 1850 and the second loss information 1860, respectively.

[0218] According to Equation 2, it is determined that the codec context neural network 1810, the transform neural network 1815, the inverse transform neural network 1825 and the codec context neural network 1830 are trained so that the training reconstructed block 1845 becomes as similar as possible to the training original block 1800 and the size of the bitstream is minimized.

[0219] According to an embodiment of the present disclosure, the transformation neural network 1815 can output not only the result of the transformation coefficient, but also the quantization result. In other words, the training transformation feature map 1820 obtained from the transformation neural network 1815 can be a transformation feature map of the quantized transformation coefficient. Therefore, the training transformation feature map 1820 is entropy encoded and sent in the form of a bit stream.

[0220] In addition, the inverse transform neural network 1825 can not only perform inverse transform, but also perform inverse quantization. In other words, the training transform feature map 1820 can be entropy decoded, and the training transform feature map 1820 and the training codec context feature map 1831 can be applied to the inverse transform neural network 1825, so that the residual block 1835 of the training inverse quantization and inverse transform can be obtained.

[0221] Figure 6 The encoding and decoding context neural network 610, the transformation neural network 615, the inverse transformation neural network 625 and the encoding and decoding context neural network 630 may correspond to Fig.18The encoding and decoding context neural network 1810, the transformation neural network 1815, the inverse transformation neural network 1825 and the encoding and decoding context neural network 1830.

[0222] also, Figure 7 The encoding and decoding context neural network 710, the transformation neural network 715 and the encoding and decoding context neural network 730 may correspond to Fig.18 The encoding and decoding context neural network 1810, the transformation neural network 1815 and the encoding and decoding context neural network 1830, and Fig.18 The inverse transform neural network 1825 is different from Figure 7 The value output by the inverse transform neural network 725 can be the training reconstruction block 1845 instead of the training inverse transform residual block 1835.

[0223] also, Figure 8 The encoding and decoding context neural network 810, the transformation neural network 815 and the encoding and decoding context neural network 830 may correspond to Fig.18 The encoding and decoding context neural network 1810, the transformation neural network 1815 and the encoding and decoding context neural network 1830, and Fig.18 The inverse transform neural network 1825 is different from Figure 8 The value output by the inverse transform neural network 825 may be an extended reconstructed block including the training reconstruction block 1845 and the adjacent pixels of the training reconstruction block 1845, rather than the residual block 1835 of the training inverse transform.

[0224] According to an embodiment of the present disclosure, an AI-based image decoding method may include: obtaining a transform block of a residual block of a current block from a bitstream; generating a transform kernel of the transform block by applying a prediction block of the current block, neighboring pixels of the current block, and coding and decoding context information to a neural network; obtaining a residual block by applying the generated transform kernel to the transform block; and reconstructing the current block by using the residual block and the prediction block.

[0225] In the AI-based image decoding method according to an embodiment of the present disclosure, unlike the related technical standards that use several fixed transform kernels, a more suitable transform kernel can be used by a neural network using neighboring pixels, prediction blocks, and codec context information, and because the neighboring pixels, prediction blocks, and codec context information are used, there is no need to send additional information for determining the transform kernel, and thus the data sent will not be increased. In other words, the codec context is available on the decoding side, so when only the supplementary information required to generate a transform that is satisfactory in terms of bit rate is sent, and the neighboring pixels and prediction blocks include information related to the residual block, the bit rate can be reduced, so that the overhead sent to the decoding side for inverse transformation can be controlled.

[0226] In addition, the transform kernel generated by the neural network is highly adaptable to various characteristics of the block to be transformed, and all context information is flexibly integrated and reflected. In other words, the codec context including valuable information for the block to be transformed is taken into account, and the codec context can be considered for both the encoding and decoding sides, thereby maximizing the utility.

[0227] According to an embodiment of the present disclosure, the coding context information may include at least one of a quantization parameter of a current block, a partition tree structure of a current block, a partition structure of adjacent pixels, a partition type of a current block, or a partition type of adjacent pixels.

[0228] According to an embodiment of the present disclosure, the transformation block may be a block transformed by a transformation kernel based on a neural network or by one of a plurality of predetermined linear transformation kernels.

[0229] According to an embodiment of the present disclosure, the generated transform kernel may include a left transform kernel to be applied to the left side of the transform kernel and a right transform kernel to be applied to the right side of the transform block.

[0230] According to an embodiment of the present disclosure, an AI-based image decoding device may include: a memory storing one or more instructions; and at least one processor configured to operate according to the one or more instructions to: obtain a transform block of a residual block of a current block from a bitstream; generate a transform kernel of the transform block by applying a prediction block of the current block, neighboring pixels of the current block, and encoding and decoding context information to a neural network; obtain a residual block by applying the generated transform kernel to the transform block; and reconstruct the current block by using the residual block and the prediction block.

[0231] In the AI-based image decoding device according to the embodiment of the present disclosure, unlike the related technical standards that use several fixed transform kernels, a more suitable transform kernel can be used by a neural network using neighboring pixels, prediction blocks, and codec context information, and because the neighboring pixels, prediction blocks, and codec context information are used, there is no need to send additional information for determining the transform kernel, and thus the data sent will not be increased. In other words, the codec context is available on the decoding side, so when only the supplementary information required to generate a transform that is satisfactory in terms of bit rate is sent, and the neighboring pixels and prediction blocks include information related to the residual block, the bit rate can be reduced, so that the overhead sent to the decoding side for inverse transformation can be controlled.

[0232] According to an embodiment of the present disclosure, the coding context information may include at least one of a quantization parameter of a current block, a partition tree structure of a current block, a partition structure of adjacent pixels, a partition type of a current block, or a partition type of adjacent pixels.

[0233] According to an embodiment of the present disclosure, the transformation block may be a block transformed by a transformation kernel based on a neural network or by one of a plurality of predetermined linear transformation kernels.

[0234] According to an embodiment of the present disclosure, the generated transform kernel may include a left transform kernel to be applied to the left side of the transform kernel and a right transform kernel to be applied to the right side of the transform block.

[0235] According to an embodiment of the present disclosure, an AI-based image encoding method may include: obtaining a residual block based on a prediction block of a current block and an original block of the current block; generating a transform kernel of a transform block of the residual block by applying the prediction block, adjacent pixels of the current block, and encoding and decoding context information to a neural network; obtaining a transform block by applying the generated transform kernel to the residual block; and generating a bitstream including the transform block.

[0236] In the AI-based image encoding method according to an embodiment of the present disclosure, unlike the related technical standards that use several fixed transform kernels, a more suitable transform kernel can be used by a neural network using neighboring pixels, prediction blocks, and codec context information, and because the neighboring pixels, prediction blocks, and codec context information are used, there is no need to send additional information for determining the transform kernel, and thus the data sent will not be increased. In other words, the codec context is available on the decoding side, so when only the supplementary information required to generate a transform that is satisfactory in terms of bit rate is sent, and the neighboring pixels and prediction blocks include information related to the residual block, the bit rate can be reduced, so that the overhead sent to the decoding side for inverse transformation can be controlled.

[0237] According to an embodiment of the present disclosure, during the image decoding process, the transformation block may be inversely transformed by a transformation kernel based on a neural network, or may be inversely transformed by one of a plurality of predetermined linear transformation kernels.

[0238] According to an embodiment of the present disclosure, the generated transform kernel may include a left transform kernel to be applied to the left side of the residual block and a right transform kernel to be applied to the right side of the residual block.

[0239] According to an embodiment of the present disclosure, an AI-based image encoding device may include: a memory storing one or more instructions; and at least one processor configured to operate according to the one or more instructions to: obtain a residual block based on a prediction block of a current block and an original block of the current block; generate a transform kernel of a transform block of the residual block by applying the prediction block, neighboring pixels of the current block, and encoding and decoding context information to a neural network; obtain a transform block by applying the generated transform kernel to the residual block; and generate a bitstream including the transform block.

[0240] In the AI-based image encoding device according to the embodiment of the present disclosure, unlike the related technical standards that use several fixed transform kernels, a more suitable transform kernel can be used by a neural network using neighboring pixels, prediction blocks, and codec context information, and because the neighboring pixels, prediction blocks, and codec context information are used, there is no need to send additional information for determining the transform kernel, and thus the data sent will not be increased. In other words, the codec context is available on the decoding side, so when only the supplementary information required to generate a transform that is satisfactory in terms of bit rate is sent, and the neighboring pixels and prediction blocks include information related to the residual block, the bit rate can be reduced, so that the overhead sent to the decoding side for inverse transformation can be controlled.

[0241] According to an embodiment of the present disclosure, during the image decoding process, the transformation block may be inversely transformed by a transformation kernel based on a neural network, or may be inversely transformed by one of a plurality of predetermined linear transformation kernels.

[0242] According to an embodiment of the present disclosure, the generated transform kernel may include a left transform kernel to be applied to the left side of the residual block and a right transform kernel to be applied to the right side of the residual block.

[0243] According to an embodiment of the present disclosure, an AI-based image decoding method may include: obtaining a transform feature map corresponding to a transform block of a residual block of a current block from a bitstream; generating a codec context feature map of the transform block by applying a prediction block of the current block, neighboring pixels of the current block, and codec context information to a first neural network; and reconstructing the current block by applying the transform feature map and the codec context feature map to a second neural network.

[0244] In the AI-based image decoding method according to an embodiment of the present disclosure, a feature map of a codec context is generated by using neighboring pixels, prediction blocks and codec context information through a neural network for generating a codec context feature map, the feature map of the codec context and the transform feature map of the transform coefficient generated by the neural network are obtained, and the feature map of the codec context and the transform feature map are input to the neural network for inverse transformation to reconstruct the current block, because no additional information other than the transform feature map of the transform coefficient generated by the neural network is sent, thereby reducing the bit rate. In addition, the neighboring pixels, prediction blocks and codec context are available on the decoding side, so the overhead sent to the decoding side for inverse transformation can be controlled, and compared with the fixed transformation kernel of the related technical standard that uses less, the transformation and inverse transformation results suitable for various features of the block to be transformed can be obtained.

[0245] According to an embodiment of the present disclosure, the second neural network may output a result value obtained by performing inverse transformation after inverse quantization.

[0246] According to an embodiment of the present disclosure, reconstruction of the current block may include: obtaining a residual block by applying a transform feature map and a coding context feature map to a second neural network; and reconstructing the current block by using the residual block and the prediction block.

[0247] According to an embodiment of the present disclosure, the reconstructed current block may further include neighboring pixels of the current block for deblocking filtering of the current block.

[0248] According to an embodiment of the present disclosure, an AI-based image decoding device may include: a memory storing one or more instructions; and at least one processor configured to operate according to the one or more instructions to: obtain a transform feature map corresponding to a transform block of a residual block of a current block from a bitstream; generate a codec context feature map of the transform block by applying a prediction block of the current block, neighboring pixels of the current block, and codec context information to a first neural network; and reconstruct the current block by applying the transform feature map and the codec context feature map to a second neural network.

[0249] In an AI-based image decoding device according to an embodiment of the present disclosure, a feature map of a codec context is generated by using neighboring pixels, prediction blocks, and codec context information through a neural network for generating a codec context feature map, obtaining the feature map of the codec context and a transform feature map of the transform coefficients generated by the neural network, and inputting the feature map of the codec context and the transform feature map into a neural network for inverse transformation to reconstruct the current block, because no additional information other than the transform feature map of the transform coefficients generated by the neural network is sent, thereby reducing the bit rate. In addition, neighboring pixels, prediction blocks, and codec context are available on the decoding side, so the overhead sent to the decoding side for inverse transformation can be controlled, and compared with a fixed transform kernel of a related technical standard that uses less, transformation and inverse transformation results suitable for various features of the block to be transformed can be obtained.

[0250] According to an embodiment of the present disclosure, the second neural network may output a result value obtained by performing inverse transformation after inverse quantization.

[0251] According to an embodiment of the present disclosure, the current block can be reconstructed by applying the transform feature map and the encoding and decoding context feature map to the second neural network to obtain a residual block, and by using the residual block and the prediction block to reconstruct the current block.

[0252] According to an embodiment of the present disclosure, the reconstructed current block may further include neighboring pixels of the current block for deblocking filtering of the current block.

[0253] According to an embodiment of the present disclosure, an AI-based image encoding method may include: obtaining a residual block based on a prediction block of a current block and an original block of the current block; generating a codec context feature map of a transform block by applying the prediction block, neighboring pixels of the current block, and codec context information to a first neural network; obtaining a transform feature map corresponding to the transform block by applying the codec context feature map and the residual block to a second neural network; and generating a bitstream including the transform feature map.

[0254] In the AI-based image encoding method according to an embodiment of the present disclosure, a feature map of a codec context is generated by using neighboring pixels, prediction blocks, and codec context information through a neural network for generating a codec context feature map, the feature map of the codec context and the transform feature map of the transform coefficient generated by the neural network are obtained, and the feature map of the codec context and the transform feature map are input to the neural network for inverse transformation to reconstruct the current block, because no additional information other than the transform feature map of the transform coefficient generated by the neural network is sent, thereby reducing the bit rate. In addition, the neighboring pixels, prediction blocks, and codec context are available on the decoding side, so the overhead sent to the decoding side for inverse transformation can be controlled, and compared with the fixed transformation kernel of the related technical standard that uses less, the transformation and inverse transformation results suitable for various features of the block to be transformed can be obtained.

[0255] According to an embodiment of the present disclosure, the second neural network may output a transform feature map of the quantized transform coefficients.

[0256] According to an embodiment of the present disclosure, an AI-based image encoding device may include: a memory storing one or more instructions; and at least one processor configured to operate according to the one or more instructions to: obtain a residual block based on a prediction block of a current block and an original block of the current block; generate a codec context feature map of a transform block by applying the prediction block, neighboring pixels of the current block, and codec context information to a first neural network; obtain a transform feature map corresponding to the transform block by applying the codec context feature map and the residual block to a second neural network; and generate a bitstream including the transform feature map.

[0257] In an AI-based image encoding device according to an embodiment of the present disclosure, a feature map of a codec context is generated by using neighboring pixels, prediction blocks, and codec context information through a neural network for generating a codec context feature map, the feature map of the codec context and a transform feature map of the transform coefficient generated by the neural network are obtained, and the feature map of the codec context and the transform feature map are input to a neural network for inverse transformation to reconstruct the current block, because no additional information other than the transform feature map of the transform coefficient generated by the neural network is sent, thereby reducing the bit rate. In addition, neighboring pixels, prediction blocks, and codec context are available on the decoding side, so the overhead sent to the decoding side for inverse transformation can be controlled, and compared with a fixed transform kernel of a related technical standard that uses less, transformation and inverse transformation results suitable for various features of the block to be transformed can be obtained.

[0258] According to an embodiment of the present disclosure, the second neural network may output a transform feature map of the quantized transform coefficients.

[0259] The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory storage medium" refers only to a tangible device and does not include signals (e.g., electromagnetic waves). The term does not distinguish between a case where data is semi-permanently stored in a storage medium and a case where data is temporarily stored in a storage medium. For example, a "non-transitory storage medium" may include a buffer that temporarily stores data.

[0260] According to an embodiment of the present disclosure, the method according to various embodiments of the present disclosure in this specification may be provided by being included in a computer program product. A computer program product is a product that can be traded between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), or distributed (e.g., downloaded or uploaded) through an application store, or distributed directly or online between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable application) may be at least temporarily generated or temporarily stored in a machine-readable storage medium (such as a memory of a manufacturer's server, an application store's server, or a relay server).

Claims

1. An image decoding method based on artificial intelligence (AI), include: Obtain a transform block of a current block from a bitstream (S1110); Obtaining a transformation kernel from the neural network by inputting a prediction block of the current block, neighboring pixels of the current block, and coding context information into the neural network (S1130); Obtaining a residual block of the current block by applying the transform kernel to the transform block (S1150); and The current block is reconstructed by using the residual block and the prediction block (S1170).

2. The AI-based image decoding method according to claim 1, in, The encoding and decoding context information includes at least one of a quantization parameter of the current block, a partition tree structure of the current block, a partition structure of the neighboring pixels, a partition type of the current block, or a partition type of the neighboring pixels.

3. The AI-based image decoding method according to claim 1 or 2, in, The transformation block is a block transformed by a transformation kernel based on a neural network, or a block transformed by a linear transformation kernel among a plurality of predetermined linear transformation kernels.

4. The AI-based image decoding method according to any one of claims 1 to 3, in, The generated transform kernels include a left transform kernel to be applied to the left side of the transform block and a right transform kernel to be applied to the right side of the transform block.

5. An image decoding method based on artificial intelligence (AI), include: Obtaining a transform feature map corresponding to a transform block of a current block from a bitstream (S1510); Generating a coding context feature map of the transform block by inputting a prediction block of the current block, neighboring pixels of the current block, and coding context information into a first neural network (S1530); as well as The current block is reconstructed based on a residual block obtained from the second neural network by inputting the transformation feature map and the encoding and decoding context feature map into the second neural network (S1550).

6. The AI-based image decoding method according to claim 5, in, The second neural network outputs a result value obtained by performing inverse transformation after inverse quantization.

7. The AI-based image decoding method according to claim 5 or 6, in, The reconstruction of the current block includes: Obtaining the residual block from the second neural network by inputting the transform feature map and the codec context feature map into the second neural network; and The current block is reconstructed by using the residual block and the prediction block.

8. The AI-based image decoding method according to any one of claims 5 to 7, in, The reconstructed current block includes neighboring pixels of the current block used for deblocking filtering of the current block.

9. An image coding method based on artificial intelligence (AI), include: Obtaining a residual block based on a prediction block of a current block and an original block of the current block (S910); Obtaining a transformation kernel from the neural network by inputting the prediction block, neighboring pixels of the current block, and coding context information into the neural network (S930); Obtaining the transformed block by applying the transformed kernel to the residual block (S950); and A bitstream including the transformed block is generated (S970).

10. The AI-based image encoding method according to claim 9, in, During the image decoding process, the transform block is inversely transformed by a transform kernel based on a neural network, or is inversely transformed by one of a plurality of predetermined linear transform kernels.

11. The AI-based image coding method according to claim 9 or 10, in, The generated transform kernels include a left transform kernel to be applied to the left side of the residual block and a right transform kernel to be applied to the right side of the residual block.

12. An image coding method based on artificial intelligence (AI), include: Obtaining a residual block based on a prediction block of a current block and an original block of the current block (S1310); By inputting the prediction block, neighboring pixels of the current block and codec context information into a first neural network, generating a codec context feature map from the first neural network (S1330); By inputting the encoding and decoding context feature map and the residual block into a second neural network, a transformation feature map is obtained from the second neural network (S1350); and A bitstream including the transformed feature map is generated (S1370).

13. The AI-based image encoding method according to claim 12, in, The second neural network outputs the transform feature map of quantized transform coefficients.

14. An image decoding device (1200) based on artificial intelligence (AI), include: a memory storing one or more instructions; as well as At least one processor configured to operate according to the one or more instructions to: Obtain a transform block for a current block from a bitstream; Obtaining a transformation kernel from a neural network by inputting a prediction block of the current block, neighboring pixels of the current block, and coding context information into the neural network; Obtaining a residual block of the current block by applying the generated transform kernel to the transform block; as well as The current block is reconstructed by using the residual block and the prediction block.

15. The AI-based image decoding device according to claim 14, in, The encoding and decoding context information includes at least one of a quantization parameter of the current block, a partition tree structure of the current block, a partition structure of the neighboring pixels, a partition type of the current block, or a partition type of the neighboring pixels.