Electronic device and method of operating the same

By using deep neural networks (DNNs) to perform AI downscaling and AI upscaling on images, the problem of low decoding efficiency for high-resolution images is solved, achieving efficient image reconstruction and upscaling effects.

CN114175652BActive Publication Date: 2025-11-04SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080054848.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-25
Filing Date
2020-04-17
Publication Date
2025-11-04
Estimated Expiration
2040-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently decode and reconstruct high-resolution images, resulting in low encoding and decoding efficiency.

Method used

Deep neural networks (DNNs) are used for AI-based image downsizing and upsizing. By jointly training the first and second DNNs, the encoding and decoding processes of images are optimized, and AI data is used to reconstruct and upscale images.

Benefits of technology

It improves the efficiency of image encoding and decoding, reduces the bit rate, and achieves efficient image reconstruction and magnification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114175652B_ABST
    Figure CN114175652B_ABST
Patent Text Reader

Abstract

A decoding device is provided, including a communication interface configured to receive AI encoded data generated as a result of AI downscaling and first encoding of an original image, a processor configured to divide the AI encoded data into image data and AI data, and an input / output (I / O) device, wherein the processor is further configured to obtain a second image by performing first decoding on a first image based on the image data, wherein the first image is obtained by performing AI downscaling on the original image, and control the I / O device to transmit the second image and the AI data to an external device. In some embodiments, the external device performs AI upscaling on the second image using the AI data, and displays a resulting third image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to a decoding device for decoding a compressed image and an operating method of the decoding device, and an artificial intelligence (AI) upscaling device including a deep neural network (DNN) that upscales an image and an operating method of the AI upscaling device. BACKGROUND

[0002] An image is stored in a recording medium or transmitted through a communication channel in the form of a bitstream after being encoded by a codec that complies with a specific data compression standard, such as a Moving Picture Experts Group (MPEG) standard.

[0003] As development and supply of hardware capable of reproducing and storing high-resolution and high-definition images are increasing, the necessity of a codec capable of efficiently encoding and decoding high-resolution and high-definition images is increasing. SUMMARY

[0004] TECHNICAL SOLUTION

[0005] A decoding device for reconstructing a compressed image and transmitting a reconstructed image and data required for AI upscaling of the reconstructed image to an AI upscaling device and an operating method of the decoding device are provided.

[0006] In addition, an AI upscaling device for receiving an image and AI data from a decoding device and AI-upscaling the image by using an upscaling deep neural network (DNN) and an operating method of the AI upscaling device are provided.

[0007] Disclosed herein is a decoding device including a communication interface configured to receive artificial intelligence (AI) encoded data, wherein the AI encoded data is generated by AI downscaling of an original image followed by first encoding, a processor configured to divide the AI encoded data into image data and AI data, and an input / output (I / O) interface, wherein the processor is further configured to obtain a second image by performing first decoding on the image data, and control the I / O interface to transmit the second image and the AI data to an external device.

[0008] In some embodiments of the decoding device, the I / O interface includes a high-definition multimedia interface (HDMI), and the processor is further configured to transmit the second image and the AI data to the external device through the HDMI.

[0009] In some embodiments of the decoding device, the processor is further configured to transmit the AI data in the form of a vendor specific information frame (VSIF) packet.

[0010] In some embodiments of the decoding device, the I / O interface includes a display port (DP), and the processor is further configured to transmit the second image and the AI data to an external device through the DP.

[0011] In some embodiments of the decoding device, the AI data includes first information indicating that the second image has undergone AI upscaling.

[0012] In some embodiments of the decoding device, the AI data includes second information related to a deep neural network (DNN) used to perform AI upscaling of the second image.

[0013] In some embodiments of the decoding device, the AI data indicates one or more color channels to be applied AI upscaling.

[0014] In some embodiments of the decoding device, the AI data indicates at least one of high dynamic range (HDR) maximum illumination, HDR color gamut, HDR PQ, HDR codec, or HDR rate control.

[0015] In some embodiments of the decoding device, the AI data indicates a width resolution of the original image and a height resolution of the original image.

[0016] In some embodiments of the decoding device, the AI data indicates an output bitrate of the first encoding.

[0017] Also disclosed herein is an operating method of a decoding device, the operating method including: receiving artificial intelligence (AI) encoding data, wherein the AI encoding data is generated by AI downscaling of an original image followed by first encoding; dividing the AI encoding data into image data and AI data; obtaining a second image by performing first decoding on the image data; and transmitting the second image and the AI data to an external device through an input / output (I / O) interface.

[0018] In some embodiments of the operating method, the step of transmitting the second image and the AI data to the external device includes transmitting the second image and the AI data to the external device through a high-definition multimedia interface (HDMI).

[0019] In some embodiments of the operating method, the step of transmitting the second image and the AI data to the external device includes transmitting the AI data in the form of a vendor specific information frame (VSIF) packet.

[0020] In some embodiments of the operation method, the transmitting of the second image and the AI data to the external device includes transmitting the second image and the AI data to the external device through a display port (DP).

[0021] In some embodiments of the operation method, the AI data includes first information indicating that the second image has undergone AI upscaling.

[0022] In some embodiments of the operation method, the AI data includes second information related to a deep neural network (DNN) used to perform AI upscaling of the second image.

[0023] Also disclosed herein is an artificial intelligence (AI) upscaling device including: an input / output (I / O) interface including a high-definition multimedia interface (HDMI), wherein the I / O interface is configured to receive, through the HDMI, AI data related to AI downscaling using a first deep neural network (DNN) and a second image corresponding to a first image, wherein the first image is obtained by performing AI downscaling on an original image; a memory storing at least one instruction; and a processor configured to execute the at least one instruction stored in the memory to: obtain information about a second DNN corresponding to the first DNN based on the AI data; and perform AI upscaling of the second image by using the second DNN, wherein the I / O interface is further configured to receive the AI data in a form of a vendor specific information frame (VSIF) packet.

[0024] Advantages of the present disclosure

[0025] The decoding device according to the embodiments of the disclosure can effectively transmit AI data and a reconstructed image to an AI upscaling device via an input and output interface.

[0026] The AI upscaling device according to the embodiments of the disclosure can effectively receive AI data and a reconstructed image from a decoding device via an input and output interface. BRIEF DESCRIPTION OF DRAWINGS

[0027] The above and other aspects, features, and advantages of certain embodiments of the disclosure will be more apparent from the following description taken in conjunction with the accompanying drawings, in which:

[0028] Brief descriptions of each drawing are provided to more fully understand the drawings described in the specification.

[0029] Figure 1 is a diagram for describing an artificial intelligence (AI) encoding process and an AI decoding process according to an embodiment.

[0030] Figure 2 is a block diagram of a configuration of an AI decoding device according to an embodiment.

[0031] Figure 3 is a diagram showing a second deep neural network (DNN) for performing AI up-scaling on a second image.

[0032] Figure 4 is a diagram for describing a convolution operation by a convolution layer.

[0033] Figure 5 is a table showing a mapping relationship between a plurality of pieces of image-related information and a plurality of pieces of DNN setting information.

[0034] Figure 6 is a diagram showing a second image including a plurality of frames.

[0035] Figure 7 is a block diagram of a configuration of an AI encoding apparatus according to an embodiment.

[0036] Figure 8 is a diagram showing a first DNN for performing AI down-scaling on an original image.

[0037] Figure 9 is a diagram for describing a method of training a first DNN and a second DNN.

[0038] Figure 10 is a diagram for describing a training process of a first DNN and a second DNN by a training apparatus.

[0039] Figure 11 is a diagram of an apparatus for performing AI down-scaling on an original image and an apparatus for performing AI up-scaling on a second image.

[0040] Figure 12 is a diagram of an AI decoding system according to an embodiment of the disclosure;

[0041] Figure 13 is a diagram of a configuration of a decoding apparatus according to an embodiment of the disclosure;

[0042] Figure 14 AI data in the form of metadata according to an embodiment of the disclosure is shown;

[0043] Figure 15 is a diagram for describing a case where AI data according to an embodiment of the disclosure is received in the form of a bitstream;

[0044] Figure 16 AI codec syntax table according to an embodiment of the disclosure is shown;

[0045] Figure 17 is a block diagram of a configuration of an AI up-scaling apparatus according to an embodiment of the disclosure;

[0046] Figure 18 is a diagram illustrating an example in which a decoding device and an AI upscaling device according to an embodiment of the disclosure transmit and receive data through a high-definition multimedia interface (HDMI);

[0047] Figure 19 is a diagram of a HDMI specification (HF) vendor specific data block (VSDB) included in extended display identification data (EDID) information according to an embodiment of the disclosure;

[0048] Figure 20 is a diagram of a header structure and a content structure of a vendor specific information frame (VSIF) according to an embodiment of the disclosure;

[0049] Figure 21 is a diagram illustrating an example in which AI data is defined in a VSIF packet according to an embodiment of the disclosure;

[0050] Figure 22 is a flowchart of an operation method of a decoding device according to an embodiment of the disclosure;

[0051] Figure 23 is a flowchart of a method of transmitting a second image and AI data via an HDMI, performed by a decoding device according to an embodiment of the disclosure;

[0052] Figure 24 is a flowchart of an operation method of an AI upscaling device according to an embodiment of the disclosure;

[0053] Figure 25 is a block diagram of a configuration of a decoding device according to an embodiment of the disclosure; and

[0054] Figure 26 is a block diagram of a configuration of an AI upscaling device according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0055] Best Mode for Carrying Out the Invention

[0056] A decoding device for reconstructing a compressed image and transmitting a reconstructed image and data required for artificial intelligence (AI) upscaling of the reconstructed image to an AI upscaling device, and an operation method of the decoding device are provided.

[0057] In addition, an AI upscaling device for receiving image data and AI data from a decoding device and AI-upscaling an image by using an upscaling deep neural network (DNN), and an operation method of the AI upscaling device are provided.

[0058] Additional aspects will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art by reference to the description, or by practice of the embodiments thereof which are presented by way of illustration.

[0059] Embodiments of the present invention

[0060] Because this disclosure allows for various modifications and numerous examples, specific embodiments will be shown in the accompanying drawings and described in detail in the written description. However, this is not intended to limit this disclosure to a particular mode of practice, and it will be understood that all changes, equivalents, and substitutions without departing from the spirit and technical scope of this disclosure are included herein.

[0061] In the description of the embodiments, detailed explanations of the related technologies are omitted when it is believed that such detailed explanations might unnecessarily obscure the essence of this disclosure. Furthermore, the numbers used in the description (e.g., first, second, etc.) are merely identifier codes used to distinguish one element from another.

[0062] Throughout this disclosure, the expression "at least one of a, b, or c" indicates only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.

[0063] Furthermore, it will be understood in this specification that when elements are “connected” or “coupled” to each other, the elements may be directly connected or coupled to each other, but alternatively, unless otherwise specified, they may be connected or coupled to each other through an intermediate element between the elements.

[0064] In this specification, for elements referred to as "units" or "modules," two or more elements may be combined into one element, or one element may be divided into two or more elements according to subdivided functions. Furthermore, each element described below, in addition to its primary function, may additionally perform some or all of the functions performed by another element, and some of the primary functions of each element may be entirely performed by another component.

[0065] In addition, in this specification, "image" or "picture" may mean a still image, a moving image including multiple consecutive still images (or frames), or a video.

[0066] Furthermore, in this specification, a deep neural network (DNN) is a representative example of an artificial neural network model that simulates brain nerves, and is not limited to artificial neural network models that use a specific algorithm.

[0067] Furthermore, in this specification, "parameters" are values ​​used in the computational processing of each layer of a neural network, and may include, for example, weights used when applying input values ​​to a specific computational expression. Here, parameters may be represented in matrix form. Parameters are values ​​set as a result of training and can be updated as needed using individual training data.

[0068] Further, in the present specification, "a first DNN" indicates a DNN for AI down-scaling an image, and "a second DNN" indicates a DNN for AI up-scaling an image.

[0069] Further, in the present specification, "DNN setting information" includes information related to elements constituting a DNN. The "DNN setting information" includes the parameters described above as information related to elements constituting a DNN. The first DNN or the second DNN can be set by using the DNN setting information.

[0070] Further, in the present specification, "an original image" indicates an image to be an object of AI encoding, and "a first image" indicates an image obtained as a result of performing AI down-scaling on the original image during an AI encoding process. Further, "a second image" indicates an image obtained via first decoding during an AI decoding process, and "a third image" indicates an image obtained by performing AI up-scaling on the second image during the AI decoding process.

[0071] Further, in the present specification, "AI down-scaling" indicates a process of reducing a resolution of an image based on AI, and "first encoding" indicates an encoding process according to an image compression method based on frequency transformation. Further, "first decoding" indicates a decoding process according to an image reconstruction method based on frequency transformation, and "AI up-scaling" indicates a process of increasing a resolution of an image based on AI.

[0072] Figure 1 is a diagram for describing an AI encoding process and an AI decoding process according to an embodiment.

[0073] As described above, when the resolution of an image is significantly increased, the throughput of information for encoding and decoding the image increases, and thus, a method for improving the efficiency of encoding and decoding an image is required.

[0074] As shown in Figure 1 According to an embodiment of the disclosure, a first image 115 is obtained by performing AI down-scaling 110 on an original image 105 having a high resolution. Then, first encoding 120 and first decoding 130 are performed on the first image 115 having a relatively low resolution, and thus, the bit rate can be greatly reduced compared to when first encoding and first decoding are performed on the original image 105.

[0075] In particular, with reference to Figure 1According to an embodiment, during AI encoding processing, a first image 115 is obtained by performing AI downscaling 110 on the original image 105 and performing first encoding 120 on the first image 115. During AI decoding processing, AI encoded data, including AI data and image data, obtained as a result of AI encoding, is received, a second image 135 is obtained via first decoding 130, and a third image 145 is obtained by performing AI upscaling 140 on the second image 135.

[0076] Referring in detail to the AI ​​encoding process, when the original image 105 is received, AI downscaling 110 is performed on the original image 105 to obtain a first image 115 with a specific resolution or quality. Here, AI downscaling 110 is performed based on AI, and the AI ​​used for AI downscaling 110 needs to be jointly trained with the AI ​​used for AI upscaling 140 for the second image 135. This is because when the AI ​​used for AI downscaling 110 and the AI ​​used for AI upscaling 140 are trained separately, the difference between the original image 105, which is the object of AI encoding, and the third image 145 reconstructed through AI decoding will increase.

[0077] In embodiments of this disclosure, AI data can be used to maintain this joint relationship during AI encoding and AI decoding processes. Therefore, the AI ​​data obtained through AI encoding may include information indicating a magnification target, and during AI decoding, AI magnification 140 is performed on the second image 135 based on the magnification target verified by the AI ​​data.

[0078] The AI ​​used for AI scaling down by 110 and the AI ​​used for AI scaling up by 140 can be implemented as a DNN. (See below for further details.) Figure 9 As described, because the first DNN and the second DNN are jointly trained by sharing loss information under a specific target, the AI ​​encoding device can provide the target information used during the joint training of the first DNN and the second DNN to the AI ​​decoding device, and the AI ​​decoding device can perform AI upscaling 140 to the target resolution on the second image 135 based on the provided target information.

[0079] about Figure 1The first encoding 120 and the first decoding 130 can reduce an amount of information of the first image 115 obtained by performing the AI downscaling 110 on the original image 105 through the first encoding 120. The first encoding 120 can include a process of generating prediction data by performing prediction on the first image 115, a process of generating residual data corresponding to a difference between the first image 115 and the prediction data, a process of transforming the residual data of a spatial domain component into a frequency domain component, a process of quantizing the residual data transformed into the frequency domain component, and a process of entropy-encoding the quantized residual data. Such first encoding 120 can be performed via one of image compression methods using frequency transformation, such as MPEG-2, H.264 Advanced Video Coding (AVC), MPEG-4, High Efficiency Video Coding (HEVC), VC-1, VP8, VP9, and AOMedia Video 1 (AV1).

[0080] The second image 135 corresponding to the first image 115 can be reconstructed by performing the first decoding 130 on the image data. The first decoding 130 can include a process of generating quantized residual data by entropy-decoding the image data, a process of inverse-quantizing the quantized residual data, a process of transforming the residual data of a frequency domain component into a spatial domain component, a process of generating prediction data, and a process of reconstructing the second image 135 by using the prediction data and the residual data. Such first decoding 130 can be performed via an image reconstruction method corresponding to one of the image compression methods using frequency transformation used in the first encoding 120, such as MPEG-2, H.264 AVC, MPEG-4, HEVC, VC-1, VP8, VP9, and AV1.

[0081] The AI encoding data obtained through the AI encoding process can include image data obtained as a result of performing the first encoding 120 on the first image 115 and AI data related to the AI downscaling 110 of the original image 105. The image data can be used during the first decoding 130, and the AI data can be used during the AI upscaling 140.

[0082] The image data can be transmitted in the form of a bitstream. The image data can include data obtained based on pixel values in the first image 115, for example, residual data that is a difference between the first image 115 and prediction data of the first image 115. In addition, the image data includes information used during the first encoding 120 of the first image 115. For example, the image data can include prediction mode information, motion information, and information related to a quantization parameter used during the first encoding 120. The image data can be generated according to a rule (for example, according to a syntax) of an image compression method used during the first encoding 120 among MPEG-2, H.264 AVC, MPEG-4, HEVC, VC-1, VP8, VP9, and AV1.

[0083] The AI data is used in the AI upscaling 140 based on the second DNN. As described above, because the first DNN and the second DNN are jointly trained, the AI data includes information that enables the AI upscaling 140 to be accurately performed on the second image 135 by the second DNN. During the AI decoding process, the AI upscaling 140 can be performed on the second image 135 based on the AI data to have a target resolution and / or quality.

[0084] The AI data can be transmitted in the form of a bitstream together with the image data. Alternatively, according to an embodiment, the AI data can be transmitted separately from the image data in the form of a frame or a packet. The AI data and the image data obtained as a result of AI encoding can be transmitted through the same network or through different networks.

[0085] Figure 2 is a block diagram of a configuration of an AI decoding apparatus 200 according to an embodiment.

[0086] Referring to Figure 2 , the AI decoding apparatus 200 according to an embodiment can include a receiver 210 and an AI decoder 230. The receiver 210 can include a communication interface 212, a parser 214, and an output interface 216. The AI decoder 230 can include a first decoder 232 and an AI upscaler 234.

[0087] The receiver 210 receives and parses AI encoded data obtained as a result of AI encoding, and outputs image data and AI data distinguishably to the AI decoder 230.

[0088] Specifically, the communication interface 212 receives AI encoded data obtained as a result of AI encoding through a network. The AI encoded data obtained as a result of performing AI encoding includes image data and AI data. The image data and the AI data can be received through the same type of network or different types of network.

[0089] The parser 214 receives the AI encoded data received through the communication interface 212 and parses the AI encoded data to distinguish the image data and the AI data. For example, the parser 214 can distinguish the image data and the AI data by reading a header of the data obtained from the communication interface 212. According to an embodiment, the parser 214 transmits the image data and the AI data distinguishably to the output interface 216 via the header of the data received through the communication interface 212, and the output interface 216 transmits the distinguished image data and AI data to the first decoder 232 and the AI up-scaler 234, respectively. At this time, it can be verified that the image data included in the AI encoded data is image data generated via a specific codec (e.g., MPEG-2, H.264 AVC, MPEG-4, HEVC, VC-1, VP8, VP9, or AV1). In this case, the corresponding information can be transmitted to the first decoder 232 by the output interface 216 so that the image data is processed via the verified codec.

[0090] According to an embodiment, the AI encoded data parsed by the parser 214 can be obtained from a data storage medium including a magnetic medium (such as a hard disk, a floppy disk, or a magnetic tape), an optical recording medium (such as a CD-ROM or a DVD), or a magneto-optical medium (such as a floptical disk).

[0091] The first decoder 232 reconstructs the second image 135 corresponding to the first image 115 based on the image data. The second image 135 obtained by the first decoder 232 is provided to the AI up-scaler 234. According to an embodiment, the first decoding-related information (such as prediction mode information, motion information, quantization parameter information, etc.) included in the image data can also be provided to the AI up-scaler 234.

[0092] Upon receiving the AI data, the AI up-scaler 234 performs AI up-scaling on the second image 135 based on the AI data. According to an embodiment, the AI up-scaling can be performed by further using the first decoding-related information (such as prediction mode information, quantization parameter information, etc.) included in the image data.

[0093] The receiver 210 and the AI decoder 230 according to an embodiment are described as separate apparatuses, but can be implemented by one processor. In this case, the receiver 210 and the AI decoder 230 can be implemented by a dedicated processor or by a combination of software and a general-purpose processor such as an application processor (AP), a central processing unit (CPU), or a graphic processing unit (GPU). The dedicated processor can be implemented by including a memory for implementing the embodiments of the disclosure or by including a memory processor for using an external memory.

[0094] Further, the receiver 210 and the AI decoder 230 can be configured by a plurality of processors. In this case, the receiver 210 and the AI decoder 230 can be implemented by a combination of dedicated processors or by a combination of software and general-purpose processors such as an AP, a CPU, or a GPU. Similarly, the AI upscaler 234 and the first decoder 232 can be implemented by different processors.

[0095] The AI data provided to the AI upscaler 234 includes information that enables the second image 135 to be processed via AI upscaling. Here, the upscaling target should correspond to the downsizing of the first DNN. Accordingly, the AI data includes information for verifying the downsizing target of the first DNN.

[0096] Examples of the information included in the AI data include difference information between the resolution of the original image 105 and the resolution of the first image 115 and information related to the first image 115.

[0097] The difference information can be expressed as information on the degree of resolution conversion of the first image 115 compared to the original image 105 (e.g., resolution conversion rate information). Further, since the resolution of the first image 115 is verified by the resolution of the reconstructed second image 135 and thus the degree of resolution conversion is verified, the difference information can be expressed only as resolution information of the original image 105. Here, the resolution information can be expressed as a vertical screen size / horizontal size, or a ratio (16:9, 4:3, etc.) and a size of one axis. Further, when there is pre-set resolution information, the resolution information can be expressed in the form of an index or a flag.

[0098] The information related to the first image 115 can include information on at least one of a bit rate of image data obtained as a result of performing the first encoding on the first image 115 or a codec type used during the first encoding of the first image 115.

[0099] The AI upscaler 234 can determine an upscaling target of the second image 135 based on at least one of the difference information or the information related to the first image 115 included in the AI data. The upscaling target can indicate, for example, to what extent the resolution will be upscaled for the second image 135. When the upscaling target is determined, the AI upscaler 234 performs AI upscaling on the second image 135 through the second DNN to obtain a third image 145 corresponding to the upscaling target.

[0100] Before describing a method of performing AI upscaling on the second image 135 according to the upscaling target by the AI upscaler 234, the AI upscaling process through the second DNN will be described with reference to Figure 3 and Figure 4

[0101] ​Figure 3 is a diagram illustrating a second DNN 300 for performing AI up-scaling on the second image 135, and Figure 4 is a diagram for describing Figure 3 a convolution operation in a first convolution layer 310.

[0102] As illustrated in Figure 3 , the second image 135 is input to the first convolution layer 310. Figure 3 The 3x3x4 indicated in the first convolution layer 310 illustrated in indicates that convolution processing is performed on one input image by using four filter kernels of a size of 3x3. Four feature maps are generated by the four filter kernels as a result of the convolution processing. Each feature map indicates an intrinsic property of the second image 135. For example, each feature map can represent a vertical direction property, a horizontal direction property, or an edge property, etc. of the second image 135.

[0103] Figure 4 The convolution operation in the first convolution layer 310 will be described in detail with reference to

[0104] One feature map 450 can be generated by multiplication and addition between parameters of the filter kernel 430 of a size of 3x3 used in the first convolution layer 310 and corresponding pixel values in the second image 135. Since four filter kernels are used in the first convolution layer 310, four feature maps can be generated by the convolution operation using the four filter kernels.

[0105] Figure 4 I1 to I49 indicated in the second image 135 in indicate pixels in the second image 135, and F1 to F9 indicated in the filter kernel 430 indicate parameters of the filter kernel 430. In addition, M1 to M9 indicated in the feature map 450 indicate samples of the feature map 450.

[0106] Figure 4 In , the second image 135 includes 49 pixels, but the number of pixels is only an example, and when the second image 135 has a resolution of 4K, the second image 135 can include, for example, 3840x2160 pixels.

[0107] During the convolution operation processing, the pixel values of I1, I2, I3, I8, I9, I10, I15, I16, and I17 of the second image 135 are multiplied by F1 to F9 of the filter kernel 430, respectively, and a value of a combination (e.g., addition) of the multiplied result values can be assigned as a value of M1 of the feature map 450. When the step of the convolution operation is 2, the pixel values of I3, I4, I5, I10, I11, I12, I17, I18, and I19 of the second image 135 are multiplied by F1 to F9 of the filter kernel 430, respectively, and a value of a combination of the multiplied result values can be assigned as a value of M2 of the feature map 450.

[0108] While the filter kernel 430 is moved along the step to the last pixel of the second image 135, a convolution operation is performed between the pixel values in the second image 135 and the parameters of the filter kernel 430, and thus the feature map 450 having a certain size can be generated.

[0109] According to the disclosure, the values of the parameters of the second DNN (e.g., values of the parameters (e.g., F1 to F9 of the filter kernel 430) of the filter kernel used in the convolution layer of the second DNN) can be optimized through joint training of the first DNN and the second DNN. As described above, the AI upscaler 234 can determine an upscaling target corresponding to the downsizing target of the first DNN based on the AI data, and determine parameters corresponding to the determined upscaling target as the parameters of the filter kernel used in the convolution layer of the second DNN.

[0110] The convolution layers included in the first DNN and the second DNN can perform processing according to the convolution operation processing described with reference to Figure 4 The convolution operation processing described with reference to Figure 4 is merely an example and is not limited thereto.

[0111] Referring back to Figure 3 , the feature map output from the first convolution layer 310 can be input to the first activation layer 320.

[0112] The first activation layer 320 can impart a non-linear feature to each feature map. The first activation layer 320 can include a sigmoid function, a Tanh function, a rectified linear unit (ReLU) function, etc., but is not limited thereto.

[0113] The first activation layer 320 imparting a non-linear feature indicates changing at least one sample value of the feature map as an output of the first convolution layer 310. Here, the changing is performed by applying a non-linear feature.

[0114] The first activation layer 320 determines whether to transmit the sample values of the feature map output from the first convolution layer 310 to the second convolution layer 330. For example, some of the sample values of the feature map are activated by the first activation layer 320 and are transmitted to the second convolution layer 330, and some of the sample values are deactivated by the first activation layer 320 and are not transmitted to the second convolution layer 330. The inherent characteristics of the second image 135 represented by the feature map are emphasized by the first activation layer 320.

[0115] The feature map 325 output from the first activation layer 320 is input to the second convolution layer 330. Figure 3 One of the feature maps 325 illustrated in FIG. 4 is a feature map that is activated with respect to the first activation layer 320. Figure 4 The result of processing the feature map 450 described above is output from the second activation layer 340.

[0116] The 3x3x4 indicated in the second convolution layer 330 indicates that the convolution processing is performed on the feature map 325 by using four filter kernels having a size of 3x3. The output of the second convolution layer 330 is input to the second activation layer 340. The second activation layer 340 can impart a non-linear feature to the input data.

[0117] The feature map 345 output from the second activation layer 340 is input to the third convolution layer 350. Figure 3 The 3x3x1 indicated in the third convolution layer 350 illustrated in FIG. 4 indicates that the convolution processing is performed by using one filter kernel having a size of 3x3 to generate one output image. The third convolution layer 350 is a layer for outputting a final image and generates one output by using one filter kernel. According to an embodiment of the disclosure, the third convolution layer 350 can output the third image 145 as a result of the convolution operation.

[0118] As will be described later, there can be a plurality of pieces of DNN setting information indicating the number of filter kernels of the first convolution layer 310, the second convolution layer 330, and the third convolution layer 350 of the second DNN 300, the parameters of the filter kernels of the first convolution layer 310, the second convolution layer 330, and the third convolution layer 350 of the second DNN 300, etc., and the plurality of pieces of DNN setting information should be associated with the plurality of pieces of DNN setting information of the first DNN. The association between the plurality of pieces of DNN setting information of the second DNN and the plurality of pieces of DNN setting information of the first DNN can be achieved via the joint training of the first DNN and the second DNN.

[0119] In Figure 3In the middle, the second DNN 300 includes three convolution layers (a first convolution layer 310, a second convolution layer 330, and a third convolution layer 350) and two activation layers (a first activation layer 320 and a second activation layer 340), but this is merely an example, and the number of convolution layers and activation layers can vary according to embodiments. Further, according to embodiments, the second DNN 300 can be implemented as a recurrent neural network (RNN). In this case, the convolutional neural network (CNN) structure of the second DNN 300 according to embodiments of the disclosure is changed to an RNN structure.

[0120] According to embodiments, the AI up-scaler 234 can include at least one arithmetic logic unit (ALU) for the above-mentioned convolution operation and operation of the activation layer. The ALU can be implemented as a processor. For the convolution operation, the ALU can include a multiplier that performs multiplication between a sample value of the second image 135 or a feature map output from a previous layer and a sample value of a filter kernel, and an adder that adds the result values of the multiplication. Further, for the operation of the activation layer, the ALU can include a multiplier that multiplies an input sample value by a weight used in a predetermined sigmoid function, Tanh function, or ReLU function, and a comparator that compares the multiplication result with a certain value to determine whether to transmit the input sample value to the next layer.

[0121] Hereinafter, a method of performing AI up-scaling on the second image 135 according to a scaling target by the AI up-scaler 234 will be described.

[0122] According to embodiments, the AI up-scaler 234 can store a plurality of pieces of DNN setting information that can be set in the second DNN.

[0123] Here, the DNN setting information can include information on at least one of the number of convolution layers included in the second DNN, the number of filter kernels for each convolution layer, or the parameters of each filter kernel. The plurality of pieces of DNN setting information can respectively correspond to various scaling targets, and the second DNN can operate based on the DNN setting information corresponding to a certain scaling target. The second DNN can have different structures based on the DNN setting information. For example, the second DNN can include three convolution layers based on any one piece of DNN setting information, and can include four convolution layers based on another piece of DNN setting information.

[0124] According to embodiments, the DNN setting information can include only the parameters of the filter kernel used in the second DNN. In this case, the structure of the second DNN does not change, but only the parameters of the internal filter kernel can change based on the DNN setting information.

[0125] The AI ​​amplifier 234 can obtain DNN setting information from multiple DNN setting information for performing AI amplification on the second image 135. Each of the multiple DNN setting information used at this time is information for obtaining a third image 145 with a predetermined resolution and / or predetermined quality, and is jointly trained with the first DNN.

[0126] For example, one of the multiple DNN setting information may include information for obtaining a third image 145 with a resolution twice that of the second image 135 (e.g., a third image 145 with a resolution of 4K (4096×2160) twice that of the second image 135 with a resolution of 2K (2048×1080)), and another DNN setting information may include information for obtaining a third image 145 with a resolution four times that of the second image 135 (e.g., a third image 145 with a resolution of 8K (8192×4320) four times that of the second image 135 with a resolution of 2K (2048×1080)).

[0127] Each of the multiple DNN setting information is related to Figure 7 The DNN setting information of the first DNN of the AI ​​encoding device 600 is jointly obtained, and the AI ​​amplifier 234 obtains one of the plurality of DNN setting information according to an amplification ratio corresponding to the reduction ratio of the DNN setting information of the first DNN. In this respect, the AI ​​amplifier 234 can verify the information of the first DNN. In order for the AI ​​amplifier 234 to verify the information of the first DNN, the AI ​​decoding device 200 according to the embodiment receives AI data including the information of the first DNN from the AI ​​encoding device 600.

[0128] In other words, the AI ​​amplifier 234 can use information received from the AI ​​encoding device 600 to verify the information targeted as the DNN setup information of the first DNN used to obtain the first image 115, and obtain the DNN setup information of the second DNN jointly trained with the DNN setup information of the first DNN.

[0129] When DNN setting information for performing AI upscaling on the second image 135 is obtained from multiple DNN setting information, the input data can be processed based on the second DNN that operates according to the obtained DNN setting information.

[0130] For example, when any DNN setting information is obtained, Figure 3 The number of filter kernels included in each of the first convolutional layer 310, the second convolutional layer 330, and the third convolutional layer 350 of the second DNN 300, as well as the parameters of the filter kernels, are set to values ​​included in the obtained DNN setup information.

[0131] Specifically, in Figure 3 the parameters of the 3x3 filter kernel used in any one of the convolution layers of the second DNN are set to {1, 1, 1, 1, 1, 1, 1, 1, 1}, and when the DNN setting information is subsequently changed, the parameters are replaced with {2, 2, 2, 2, 2, 2, 2, 2, 2} included as parameters in the changed DNN setting information.

[0132] The AI up-scaler 234 can obtain DNN setting information for AI up-scaling from the plurality of pieces of DNN setting information based on information included in the AI data, and now the AI data for obtaining the DNN setting information will be described.

[0133] According to an embodiment, the AI up-scaler 234 can obtain DNN setting information for AI up-scaling from the plurality of pieces of DNN setting information based on difference information included in the AI data. For example, when it is verified based on the difference information that the resolution of the original image 105 (e.g., 4K (4096x2160)) is twice as high as the resolution of the first image 115 (e.g., 2K (2048x1080)), the AI up-scaler 234 can obtain DNN setting information for increasing the resolution of the second image 135 by two times.

[0134] According to another embodiment, the AI up-scaler 234 can obtain DNN setting information for AI up-scaling of the second image 135 from the plurality of pieces of DNN setting information based on information related to the first image 115 included in the AI data. The AI up-scaler 234 can pre-determine a mapping relationship between the image-related information and the DNN setting information, and obtain DNN setting information mapped to the information related to the first image 115.

[0135] Figure 5 is a table showing a mapping relationship between a plurality of pieces of image-related information and a plurality of pieces of DNN setting information.

[0136] By according to the embodiment of Figure 5 , it will be determined that the AI encoding and AI decoding processes according to the embodiment of the disclosure consider not only a change in resolution. As shown in Figure 5 , resolution (such as standard definition (SD), high definition (HD), or full HD), bit rate (such as 10 Mbps, 15 Mbps, or 20 Mbps), and codec information (such as AV1, H.264, or HEVC) can be considered individually or collectively to select DNN setting information. In consideration of such resolution, bit stream, and codec information, it is considered that training of each element should be performed in conjunction with encoding and decoding processes during an AI training process (see Figure 9 ).

[0137] Thus, when a plurality of pieces of DNN setting information is provided according to the training based on the image-related information including the codec type as illustrated in Figure 5 When a plurality of pieces of DNN setting information is provided according to the training based on the image-related information including the codec type as illustrated in

[0138] In other words, the AI upscaler 234 can use the DNN setting information according to the image-related information by matching the image-related information on the left side of the table with the DNN setting information on the right side of the table. Figure 5

[0139] As illustrated in Figure 5 When it is verified from the information related to the first image 115 that the resolution of the first image 115 is SD, the bit rate of the image data obtained as a result of performing the first encoding on the first image 115 is 10 Mbps, and the first encoding is performed via the AV1 codec on the first image 115, the AI upscaler 234 can use the A DNN setting information among the plurality of pieces of DNN setting information.

[0140] Further, when it is verified from the information related to the first image 115 that the resolution of the first image 115 is HD, the bit rate of the image data obtained as a result of performing the first encoding is 15 Mbps, and the first encoding is performed via the H.264 codec, the AI upscaler 234 can use the B DNN setting information among the plurality of pieces of DNN setting information.

[0141] ​Also, when it is verified from the information related to the first image 115 that the resolution of the first image 115 is full HD, the bit rate of the image data obtained as a result of performing the first encoding is 20 Mbps, and the first encoding is performed via the HEVC codec, the AI up-scaler 234 can use the C DNN setting information among the plurality of pieces of DNN setting information, and when it is verified that the resolution of the first image 115 is full HD, the bit rate of the image data obtained as a result of performing the first encoding is 15 Mbps, and the first encoding is performed via the HEVC codec, the AI up-scaler 234 can use the D DNN setting information among the plurality of pieces of DNN setting information. One of the C DNN setting information and the D DNN setting information is selected based on whether the bit rate of the image data obtained as a result of performing the first encoding on the first image 115 is 20 Mbps or 15 Mbps. Different bit rates of the image data obtained when the first encoding is performed on the first image 115 of the same resolution via the same codec indicate different qualities of the reconstructed image. Accordingly, the first DNN and the second DNN can be jointly trained based on a certain image quality, and thus the AI up-scaler 234 can obtain the DNN setting information according to the bit rate of the image data indicating the quality of the second image 135.

[0142] According to another embodiment, the AI up-scaler 234 can obtain the DNN setting information for performing AI up-scaling on the second image 135 from among the plurality of pieces of DNN setting information, considering both the information related to the first image 115 included in the AI data and the information provided from the first decoder 232 (prediction mode information, motion information, quantization parameter information, etc.). For example, the AI up-scaler 234 can receive the quantization parameter information used during the first encoding process of the first image 115 from the first decoder 232, verify the bit rate of the image data obtained as a result of encoding of the first image 115 from the AI data, and obtain the DNN setting information corresponding to the quantization parameter information and the bit rate. Even when the bit rate is the same, the quality of the reconstructed image can vary according to the complexity of the image. The bit rate is a value representing the entire first image 115 on which the first encoding is performed, and even within the first image 115, the quality of each frame can vary. Accordingly, when the prediction mode information, the motion information, and / or the quantization parameter available for each frame from the first decoder 232 are considered together, compared to when only the AI data is used, more suitable DNN setting information for the second image 135 can be obtained.

[0143] Also, according to an embodiment, the AI data can include an identifier of mutually agreed DNN setting information. The identifier of the DNN setting information is information for distinguishing a pair of DNN setting information to be jointly trained between the first DNN and the second DNN, so that AI upscaling is performed on the second image 135 to an upscaling target corresponding to the downsizing target of the first DNN. The AI upscaler 234 can perform AI upscaling on the second image 135 by using the DNN setting information corresponding to the identifier of the DNN setting information included in the AI data, after obtaining the identifier of the DNN setting information included in the AI data. For example, an identifier indicating each of the plurality of DNN setting information settable in the first DNN and an identifier indicating each of the plurality of DNN setting information settable in the second DNN can be designated in advance. In this case, the same identifier can be designated for a pair of DNN setting information settable in each of the first DNN and the second DNN. The AI data can include an identifier of the DNN setting information set in the first DNN for AI downsizing of the original image 105. The AI upscaler 234 receiving the AI data can perform AI upscaling on the second image 135 by using the DNN setting information indicated by the identifier included in the AI data among the plurality of DNN setting information.

[0144] Also, according to an embodiment, the AI data can include DNN setting information. The AI upscaler 234 can perform AI upscaling on the second image 135 by using the DNN setting information, after obtaining the DNN setting information included in the AI data.

[0145] According to an embodiment, when a plurality of pieces of information constituting the DNN setting information (e.g., the number of convolution layers, the number of filter kernels for each convolution layer, the parameters of each filter kernel, etc.) are stored in the form of a lookup table, the AI upscaler 234 can obtain the DNN setting information by combining some values selected from the values in the lookup table based on the information included in the AI data, and perform AI upscaling on the second image 135 by using the obtained DNN setting information.

[0146] According to an embodiment, when the structure of the DNN corresponding to the upscaling target is determined, the AI upscaler 234 can obtain DNN setting information, e.g., the parameters of the filter kernel, corresponding to the determined structure of the DNN.

[0147] The AI upscaler 234 obtains the DNN setting information of the second DNN through the AI data including information related to the first DNN, and performs AI upscaling on the second image 135 through the second DNN set based on the obtained DNN setting information, and in this case, memory usage and throughput can be reduced compared to when the characteristics of the second image 135 are directly analyzed for upscaling.

[0148] According to an embodiment, when the second image 135 includes a plurality of frames, the AI up-scaler 234 can independently obtain DNN setting information for a certain number of frames, or can obtain common DNN setting information for all frames.

[0149] Figure 6 is a diagram illustrating the second image 135 including a plurality of frames.

[0150] As illustrated in Figure 6 , the second image 135 can include frame t0 to frame tn. For example, the second image 135 includes frame t0, …, frame ta, …, frame tb, …, frame tn.

[0151] According to an embodiment, the AI up-scaler 234 can obtain DNN setting information of a second DNN through AI data, and perform AI up-scaling on frame t0 to frame tn based on the obtained DNN setting information. In other words, frame t0 to frame tn can be processed through AI up-scaling based on common DNN setting information.

[0152] According to another embodiment, the AI up-scaler 234 can perform AI up-scaling on some of frame t0 to frame tn (e.g., frame t0 to frame ta) by using "A" DNN setting information obtained from AI data, and perform AI up-scaling on frame ta+1 to frame tb by using "B" DNN setting information obtained from AI data. In addition, the AI up-scaler 234 can perform AI up-scaling on frame tb+1 to frame tn by using "C" DNN setting information obtained from AI data. In other words, the AI up-scaler 234 can independently obtain DNN setting information for each group including a certain number of frames among a plurality of frames, and perform AI up-scaling on frames included in each group by using the independently obtained DNN setting information.

[0153] According to another embodiment, the AI up-scaler 234 can independently obtain DNN setting information for each frame forming the second image 135. In other words, when the second image 135 includes three frames, the AI up-scaler 234 can perform AI up-scaling on a first frame by using DNN setting information obtained with respect to the first frame, perform AI up-scaling on a second frame by using DNN setting information obtained with respect to the second frame, and perform AI up-scaling on a third frame by using DNN setting information obtained with respect to the third frame. According to a method of obtaining DNN setting information based on information (prediction mode information, motion information, quantization parameter information, etc.) provided from the first decoder 232 and information related to the first image 115 included in the above-described AI data, DNN setting information can be independently obtained for each frame included in the second image 135. This is because mode information, quantization parameter information, etc. can be independently determined for each frame included in the second image 135.

[0154] According to another embodiment, the AI data can include information about which frame the DNN setting information is valid for, wherein the DNN setting information is obtained based on the AI data. For example, when the AI data includes information indicating that the DNN setting information is valid until frame ta, the AI upscaler 234 performs AI upscaling on frames t0 to ta by using the DNN setting information obtained based on the AI data. Also, when another piece of AI data includes information indicating that the DNN setting information is valid until frame tn, the AI upscaler 234 performs AI upscaling on frames ta+1 to tn by using the DNN setting information obtained based on the other piece of AI data.

[0155] Hereinafter, a description will be given of an AI encoding apparatus 600 for performing AI encoding on an original image 105. Figure 7 The AI encoding apparatus 600 will be described in detail with reference to FIG. 6.

[0156] Figure 7 is a block diagram of a configuration of the AI encoding apparatus 600 according to an embodiment.

[0157] Referring to FIG. 6, Figure 7 The AI encoding apparatus 600 can include an AI encoder 610 and a transmitter 630. The AI encoder 610 can include an AI downscaler 612 and a first encoder 614. The transmitter 630 can include a data processor 632 and a communication interface 634.

[0158] In Figure 7 , the AI encoder 610 and the transmitter 630 are illustrated as independent apparatuses, but the AI encoder 610 and the transmitter 630 can be implemented by one processor. In this case, the AI encoder 610 and the transmitter 630 can be implemented by a dedicated processor or by a combination of software and a general-purpose processor such as an AP, a CPU, or a graphics processor GPU. The dedicated processor can be implemented by including a memory for implementing the embodiments of the disclosure or by including a memory processor for using an external memory.

[0159] Also, the AI encoder 610 and the transmitter 630 can be composed of a plurality of processors. In this case, the AI encoder 610 and the transmitter 630 can be implemented by a combination of dedicated processors or by a combination of software and a plurality of general-purpose processors such as an AP, a CPU, or a GPU. The AI downscaler 612 and the first encoder 614 can be implemented by different processors.

[0160] The AI encoder 610 performs AI downscaling on the original image 105 and first encoding on the first image 115, and transmits AI data and image data to the transmitter 630. The transmitter 630 transmits the AI data and the image data to the AI decoding apparatus 200.

[0161] The image data includes data obtained as a result of performing the first encoding on the first image 115. The image data can include data obtained based on pixel values in the first image 115, for example, residual data that is a difference between the first image 115 and prediction data of the first image 115. Also, the image data includes information used during the first encoding process of the first image 115. For example, the image data can include prediction mode information, motion information, quantization parameter information, etc. used to perform the first encoding on the first image 115.

[0162] The AI data includes information that enables AI upscaling of the second image 135 to an upscaling target corresponding to the downsizing target of the first DNN. According to an embodiment, the AI data can include difference information between the original image 105 and the first image 115. Also, the AI data can include information related to the first image 115. The information related to the first image 115 can include information on at least one of a resolution of the first image 115, a bit rate of image data obtained as a result of performing the first encoding on the first image 115, and a codec type used during the first encoding of the first image 115.

[0163] According to an embodiment, the AI data can include an identifier of mutually agreed DNN setting information, such that AI upscaling of the second image 135 to an upscaling target corresponding to the downsizing target of the first DNN is performed.

[0164] Also, according to an embodiment, the AI data can include DNN setting information that can be set in the second DNN.

[0165] The AI downsizer 612 can obtain the first image 115 obtained by performing AI downsizing on the original image 105 via the first DNN. The AI downsizer 612 can determine a downsizing target of the original image 105 based on a predetermined criterion.

[0166] To obtain the first image 115 matching the downsizing target, the AI downsizer 612 can store a plurality of pieces of DNN setting information that can be set in the first DNN. The AI downsizer 612 obtains DNN setting information corresponding to the downsizing target from among the plurality of pieces of DNN setting information, and performs AI downsizing on the original image 105 via the first DNN set in the obtained DNN setting information.

[0167] Each of the plurality of pieces of DNN setting information can be trainable to obtain the first image 115 of a predetermined resolution and / or a predetermined quality. For example, any one of the plurality of pieces of DNN setting information can include information for obtaining the first image 115 of which the resolution is half of the resolution of the original image 105 (e.g., the first image 115 of 2K (2048x1080) which is half of 4K (4096x2160) of the original image 105), and another piece of DNN setting information can include information for obtaining the first image 115 of which the resolution is a quarter of the resolution of the original image 105 (e.g., the first image 115 of 2K (2048x1080) which is a quarter of 8K (8192x4320) of the original image 105).

[0168] According to an embodiment, when the plurality of pieces of information constituting the DNN setting information (e.g., the number of convolution layers, the number of filter kernels for each convolution layer, the parameters of each filter kernel, etc.) are stored in the form of a lookup table, the AI downscaler 612 can obtain the DNN setting information by combining some values selected from the values in the lookup table based on the downscaling target, and perform AI downscaling on the original image 105 by using the obtained DNN setting information.

[0169] According to an embodiment, the AI downscaler 612 can determine the structure of the DNN corresponding to the downscaling target, and obtain the DNN setting information corresponding to the determined structure of the DNN, for example, obtain the parameters of the filter kernel.

[0170] As the first DNN and the second DNN are jointly trained, the plurality of pieces of DNN setting information for performing AI downscaling on the original image 105 can have optimized values. Here, each piece of DNN setting information includes at least one of the number of convolution layers included in the first DNN, the number of filter kernels for each convolution layer, or the parameters of each filter kernel.

[0171] The AI downscaler 612 can set the first DNN with the DNN setting information for performing AI downscaling on the original image 105 to obtain the first image 115 of a certain resolution and / or a certain quality through the first DNN. When the DNN setting information for performing AI downscaling on the original image 105 is obtained from the plurality of pieces of DNN setting information, each layer in the first DNN can process input data based on information included in the DNN setting information.

[0172] Hereinafter, a method of determining a downscaling target performed by the AI downscaler 612 will be described. The downscaling target can indicate, for example, how much the resolution is reduced from the original image 105 to obtain the first image 115.

[0173] According to an embodiment, the AI downscaler 612 can determine the downscaling target based on at least one of a compression ratio (e.g., a resolution difference between the original image 105 and the first image 115, a target bit rate, etc.), a compression quality (e.g., a type of bit rate), compression history information, or a type of the original image 105.

[0174] For example, the AI downscaler 612 can determine the downscaling target based on a compression ratio, a compression quality, etc., which are preset or input from a user.

[0175] As another example, the AI downscaler 612 can determine the downscaling target by using compression history information stored in the AI encoding device 600. For example, according to the compression history information usable by the AI encoding device 600, an encoding quality, a compression ratio, etc., preferred by a user can be determined, and the downscaling target can be determined according to the encoding quality determined based on the compression history information. For example, the resolution, quality, etc., of the first image 115 can be determined based on an encoding quality most frequently used according to the compression history information.

[0176] As another example, the AI downscaler 612 can determine the downscaling target based on an encoding quality more frequently used than a certain threshold (e.g., an average quality of an encoding quality more frequently used than a certain threshold) according to the compression history information.

[0177] As another example, the AI downscaler 612 can determine the downscaling target based on a resolution, a type (e.g., a file format), etc., of the original image 105.

[0178] According to an embodiment, when the original image 105 includes a plurality of frames, the AI downscaler 612 can independently determine the downscaling target for a certain number of frames, or can determine the downscaling target for all the frames.

[0179] According to an embodiment, the AI downscaler 612 can divide the frames included in the original image 105 into a certain number of groups, and independently determine the downscaling target for each group. The same or different downscaling target can be determined for each group. According to each group, the number of frames included in the group can be the same or different.

[0180] According to another embodiment, the AI downscaler 612 can independently determine the downscaling target for each frame included in the original image 105. The same or different downscaling target can be determined for each frame.

[0181] Hereinafter, an example of a structure of the first DNN 700 based on which AI downscaling is performed will be described.

[0182] Figure 8 is a diagram illustrating the first DNN 700 for performing AI downscaling on the original image 105.

[0183] As in the above-described example, the first DNN 700 can include a first input layer 710, a first feature extraction layer 720, a first downscaling layer 730, a first quality enhancement layer 740, and a first output layer 750.Figure 8 As illustrated in FIG. 7, the original image 105 is input to the first convolution layer 710. The first convolution layer 710 performs a convolution process on the original image 105 by using 32 filter kernels of a size of 5x5. 32 feature maps generated as a result of the convolution process are input to the first activation layer 720. The first activation layer 720 can impart non-linear features to the 32 feature maps.

[0184] The first activation layer 720 determines whether to transmit the sample values of the feature maps output from the first convolution layer 710 to the second convolution layer 730. For example, some of the sample values of the feature maps are activated by the first activation layer 720 and are transmitted to the second convolution layer 730, and some of the sample values are deactivated by the first activation layer 720 and are not transmitted to the second convolution layer 730. Information represented by the feature maps output from the first convolution layer 710 is emphasized by the first activation layer 720.

[0185] The output 725 of the first activation layer 720 is input to the second convolution layer 730. The second convolution layer 730 performs a convolution process on the input data by using 32 filter kernels of a size of 5x5. 32 feature maps output as a result of the convolution process are input to the second activation layer 740, and the second activation layer 740 can impart non-linear features to the 32 feature maps.

[0186] The output 745 of the second activation layer 740 is input to the third convolution layer 750. The third convolution layer 750 performs a convolution process on the input data by using one filter kernel of a size of 5x5. As a result of the convolution process, one image can be output from the third convolution layer 750. The third convolution layer 750 generates one output by using the one filter kernel as a layer for outputting a final image. According to an embodiment of the disclosure, the third convolution layer 750 can output the first image 115 as a result of the convolution operation.

[0187] There can be a plurality of pieces of DNN setting information indicating the number of filter kernels of the first convolution layer 710, the second convolution layer 730, and the third convolution layer 750 of the first DNN 700, parameters of each of the filter kernels of the first convolution layer 710, the second convolution layer 730, and the third convolution layer 750 of the first DNN 700, etc., and the plurality of pieces of DNN setting information can be associated with the plurality of pieces of DNN setting information of the second DNN. The association between the plurality of pieces of DNN setting information of the first DNN and the plurality of pieces of DNN setting information of the second DNN can be achieved via joint training of the first DNN and the second DNN.

[0188] In Figure 8In the middle, the first DNN 700 includes three convolution layers (a first convolution layer 710, a second convolution layer 730, and a third convolution layer 750) and two activation layers (a first activation layer 720 and a second activation layer 740), but this is merely an example, and the number of convolution layers and activation layers can vary according to embodiments. Also, according to embodiments, the first DNN 700 can be implemented as an RNN. In this case, the CNN structure of the first DNN 700 according to embodiments of the disclosure is changed to an RNN structure.

[0189] According to embodiments, the AI downscaler 612 can include at least one ALU for the operations of the above-described convolution operation and activation layer. The ALU can be implemented as a processor. For the convolution operation, the ALU can include a multiplier that performs multiplication between a sample value of the original image 105 or a feature map output from a previous layer and a sample value of a filter kernel, and an adder that adds the result values of the multiplication. Also, for the operation of the activation layer, the ALU can include a multiplier that multiplies an input sample value by a weight used in a predetermined sigmoid function, Tanh function, or ReLU function, and a comparator that compares the multiplication result with a certain value to determine whether to transmit the input sample value to the next layer.

[0190] Referring back to Figure 7 Upon receiving the first image 115 from the AI downscaler 612, the first encoder 614 can reduce the amount of information of the first image 115 by performing first encoding on the first image 115. Image data corresponding to the first image 115 can be obtained as a result of the first encoding performed by the first encoder 614.

[0191] The data processor 632 processes at least one of the AI data or the image data to be transmitted in a certain form. For example, when the AI data and the image data are to be transmitted in the form of a bitstream, the data processor 632 can process the AI data to be represented in the form of a bitstream, and transmit the image data and the AI data in the form of one bitstream through the communication interface 634. As another example, the data processor 632 can process the AI data to be represented in the form of a bitstream, and transmit each of a bitstream corresponding to the AI data and a bitstream corresponding to the image data through the communication interface 634. As another example, the data processor 632 can process the AI data to be represented in the form of a frame or a packet, and transmit the image data in the form of a bitstream and the AI data in the form of a frame or a packet through the communication interface 634.

[0192] The communication interface 634 transmits AI encoded data obtained as a result of performing AI encoding through a network. The AI encoded data obtained as a result of performing AI encoding includes image data and AI data. The image data and the AI data can be transmitted through the same type of network or different types of networks.

[0193] According to an embodiment, the AI encoded data obtained as a result of processing by the data processor 632 can be stored in a data storage medium including a magnetic medium such as a hard disk, a floppy disk, or a magnetic tape, an optical recording medium such as a CD-ROM or a DVD, or a magneto-optical medium such as a floptical disk.

[0194] Hereinafter, a method of jointly training the first DNN 700 and the second DNN 300 will be described with reference to Figure 9

[0195] Figure 9 is a diagram for describing a method of training the first DNN 700 and the second DNN 300.

[0196] In an embodiment, the original image 105 AI encoded through the AI encoding process is reconstructed into a third image 145 via the AI decoding process, and in order to maintain similarity between the original image 105 and the third image 145 obtained as a result of AI decoding, a correlation between the AI encoding process and the AI decoding process is required. In other words, information lost in the AI encoding process needs to be reconstructed during the AI decoding process, and in this regard, the first DNN 700 and the second DNN 300 need to be jointly trained.

[0197] In order to perform accurate AI decoding, ultimately, quality loss information 830 corresponding to a result of comparing the third training image 804 and the original training image 801 as illustrated in FIG. 8 needs to be reduced. Accordingly, the quality loss information 830 is used to train both the first DNN 700 and the second DNN 300. Figure 9

[0198] First, a training process as illustrated in FIG. 8 will be described. Figure 9

[0199] In Figure 9 , the original training image 801 is an image to be AI down-scaled, and the first training image 802 is an image obtained by performing AI down-scaling on the original training image 801. Further, the third training image 804 is an image obtained by performing AI up-scaling on the first training image 802.

[0200] ​​​The original training image 801 includes a still image or a moving image including a plurality of frames. According to an embodiment, the original training image 801 can include a luminance image extracted from a still image or a moving image including a plurality of frames. Also, according to an embodiment, the original training image 801 can include a patch image extracted from a still image or a moving image including a plurality of frames. When the original training image 801 includes a plurality of frames, the first training image 802, the second training image, and the third training image 804 also each include a plurality of frames. When the plurality of frames of the original training image 801 are sequentially input to the first DNN 700, the plurality of frames of the first training image 802, the second training image, and the third training image 804 can be sequentially obtained through the first DNN 700 and the second DNN 300.

[0201] For joint training of the first DNN 700 and the second DNN 300, the original training image 801 is input to the first DNN 700. The original training image 801 input to the first DNN 700 is output as the first training image 802 via AI downscaling, and the first training image 802 is input to the second DNN 300. The third training image 804 is output as a result of performing AI upscaling on the first training image 802.

[0202] Referring to Figure 9 , the first training image 802 is input to the second DNN 850, and according to an embodiment, a second training image obtained when first encoding and first decoding are performed on the first training image 802 can be input to the second DNN 300. To input the second training image to the second DNN 300, any one codec of MPEG-2, H.264, MPEG-4, HEVC, VC-1, VP8, VP9, and AV1 can be used. In particular, any one codec of MPEG-2, H.264, MPEG-4, HEVC, VC-1, VP8, VP9, and AV1 can be used to perform first encoding on the first training image 802 and perform first decoding on image data corresponding to the first training image 802.

[0203] Referring to Figure 9 Separately from the first training image 802 output through the first DNN 700, a reduced training image 803 obtained by performing conventional downscaling on the original training image 801 is obtained. Here, the conventional downscaling can include at least one of bilinear scaling, bicubic scaling, lanczos scaling, or ladder scaling.

[0204] To prevent the structural features of the first image 115 from greatly deviating from the structural features of the original image 105, the reduced training image 803 is obtained to preserve the structural features of the original training image 801.

[0205] Before performing the training, the first DNN 700 and the second DNN 300 can be set to predetermined DNN setting information. When the training is performed, the structure loss information 810, the complexity loss information 820, and the quality loss information 830 can be determined.

[0206] The structure loss information 810 can be determined based on a result of comparing the reduced training image 803 and the first training image 802. For example, the structure loss information 810 can correspond to a difference between structure information of the reduced training image 803 and structure information of the first training image 802. The structure information can include various features that can be extracted from an image, such as brightness, contrast, a histogram, etc. of the image. The structure loss information 810 indicates how much structure information of the original training image 801 is maintained in the first training image 802. When the structure loss information 810 is low, the structure information of the first training image 802 is similar to the structure information of the original training image 801.

[0207] The complexity loss information 820 can be determined based on a spatial complexity of the first training image 802. For example, a total variance value of the first training image 802 can be used as the spatial complexity. The complexity loss information 820 is related to a bit rate of image data obtained by performing the first encoding on the first training image 802. It is defined that when the complexity loss information 820 is low, the bit rate of the image data is low.

[0208] The quality loss information 830 can be determined based on a result of comparing the original training image 801 and the third training image 804. The quality loss information 830 can include at least one of an L1 norm value, an L2 norm value, a structural similarity (SSIM) value, a peak signal-to-noise ratio-human visual system (PSNR-HVS) value, a multi-scale SSIM (MS-SSIM) value, a variance of information fidelity (VIF) value, or a video multi-method assessment fusion (VMAF) value regarding a difference between the original training image 801 and the third training image 804. The quality loss information 830 indicates how similar the third training image 804 is to the original training image 801. When the quality loss information 830 is low, the third training image 804 is more similar to the original training image 801.

[0209] Referring to Figure 9 The structure loss information 810, the complexity loss information 820, and the quality loss information 830 are used to train the first DNN 700, and the quality loss information 830 is used to train the second DNN 300. In other words, the quality loss information 830 is used to train both the first DNN 700 and the second DNN 300.

[0210] The first DNN 700 can update the parameters such that the final loss information determined based on the structural loss information 810, the complexity loss information 820, and the quality loss information 830 is reduced or minimized. Also, the second DNN 300 can update the parameters such that the quality loss information 830 is reduced or minimized.

[0211] The final loss information for training the first DNN 700 and the second DNN 300 can be determined as Equation 1 below.

[0212] [Equation 1]

[0213] LossDS = a x structural loss information + b x complexity loss information + c x quality loss information

[0214] LossUS = d x quality loss information

[0215] In Equation 1, LossDS indicates the final loss information to be reduced or minimized to train the first DNN 700, and LossUS indicates the final loss information to be reduced or minimized to train the second DNN 300. Also, a, b, c, and d can be predetermined specific weights.

[0216] In other words, the first DNN 700 updates the parameters in a direction in which LossDS of Equation 1 is reduced, and the second DNN 300 updates the parameters in a direction in which LossUS is reduced. When the parameters of the first DNN 700 are updated according to LossDS derived during training, the first training image 802 obtained based on the updated parameters becomes different from the previous first training image 802 obtained based on the un-updated parameters, and thus, the third training image 804 also becomes different from the previous third training image 804. When the third training image 804 becomes different from the previous third training image 804, the quality loss information 830 is also re-determined, and the second DNN 300 updates the parameters accordingly. When the quality loss information 830 is re-determined, LossDS is also re-determined, and the first DNN 700 updates the parameters according to the re-determined LossDS. In other words, the update of the parameters of the first DNN 700 causes the update of the parameters of the second DNN 300, and the update of the parameters of the second DNN 300 causes the update of the parameters of the first DNN 700. In other words, because the first DNN 700 and the second DNN 300 are jointly trained by sharing the quality loss information 830, the parameters of the first DNN 700 and the parameters of the second DNN 300 can be jointly optimized.

[0217] Referring to Equation 1, it is verified that LossUS is determined from the quality loss information 830, but this is only an example, and LossUS can be determined based on at least one of the structure loss information 810 and the complexity loss information 820 and the quality loss information 830.

[0218] In the above, it has been described that the AI up-scaler 234 of the AI decoding device 200 and the AI down-scaler 612 of the AI encoding device 600 store a plurality of pieces of DNN setting information, and now a method of training each piece of DNN setting information stored in the AI up-scaler 234 and the AI down-scaler 612 will be described.

[0219] As described with reference to Equation 1, the first DNN 700 updates the parameters in consideration of the similarity between the structure information of the first training image 802 and the structure information of the original training image 801 (structure loss information 810), the bit rate of the image data obtained as a result of performing the first encoding on the first training image 802 (complexity loss information 820), and the difference between the third training image 804 and the original training image 801 (quality loss information 830).

[0220] Specifically, the parameters of the first DNN 700 can be updated so that the first training image 802 having similar structure information to the original training image 801 is obtained and the image data having a small bit rate is obtained when the first encoding is performed on the first training image 802, and at this time, the second DNN 300 performing the AI up-scaling on the first training image 802 obtains the third training image 804 similar to the original training image 801.

[0221] The direction in which the parameters of the first DNN 700 are optimized can vary by adjusting the weights a, b, and c of Equation 1. For example, when the weight b is determined to be high, the parameters of the first DNN 700 can be updated by giving priority to the low bit rate of the third training image 804 over the high quality. Also, when the weight c is determined to be high, the parameters of the first DNN 700 can be updated by giving priority to the high quality of the third training image 804 over the high bit rate or maintaining the structure information of the original training image 801.

[0222] Also, the direction in which the parameters of the first DNN 700 are optimized can vary according to the type of codec used to perform the first encoding on the first training image 802. This is because the second training image to be input to the second DNN 300 can vary according to the type of codec.

[0223] In other words, the parameters of the first DNN 700 and the parameters of the second DNN 300 can be jointly updated based on the weights a, b, and c and the type of codec used to perform the first encoding on the first training image 802. Accordingly, when the first DNN 700 and the second DNN 300 are trained after the weights a, b, and c are each determined to be a certain value and the type of codec is determined to be a certain type, the parameters of the first DNN 700 and the parameters of the second DNN 300 that are associated with and optimized for each other can be determined.

[0224] Further, when the first DNN 700 and the second DNN 300 are trained after the weights a, b, and c and the type of codec are changed, the parameters of the first DNN 700 and the parameters of the second DNN 300 that are associated with and optimized for each other can be determined. In other words, when the first DNN 700 and the second DNN 300 are trained while the values of the weights a, b, and c and the type of codec are changed, a plurality of pieces of DNN setting information that are jointly trained with each other can be determined in the first DNN 700 and the second DNN 300.

[0225] As described above with reference to Figure 5 The plurality of pieces of DNN setting information of the first DNN 700 and the second DNN 300 can be mapped to the information related to the first image, as described above with reference to FIG. 7. In order to set such a mapping relationship, the first training image 802 output from the first DNN 700 can be first encoded according to a certain bit rate via a certain codec according to a certain bit rate, and a second training image obtained by first decoding a bitstream obtained as a result of performing the first encoding can be input to the second DNN 300. In other words, by training the first DNN 700 and the second DNN 300 after setting an environment so that the first training image 802 of a certain resolution is first encoded according to a certain bit rate via a certain codec, a pair of DNN setting information that is mapped to the resolution of the first training image 802, the type of codec used to perform the first encoding on the first training image 802, and the bit rate of the bitstream obtained as a result of performing the first encoding on the first training image 802 can be determined. By differently changing the resolution of the first training image 802, the type of codec used to perform the first encoding on the first training image 802, and the bit rate of the bitstream obtained as a result of the first encoding of the first training image 802, a mapping relationship between the plurality of pieces of DNN setting information of the first DNN 700 and the second DNN 300 and the plurality of pieces of information related to the first image can be determined.

[0226] Figure 10 is a diagram for describing a training process of the training apparatus 1000 on the first DNN 700 and the second DNN.

[0227] Referring to Figure 9The training of the described first DNN 700 and second DNN 300 can be performed by the training device 1000. The training device 1000 includes the first DNN 700 and the second DNN 300. The training device 1000 can be, for example, the AI encoding device 600 or a separate server. The DNN setting information of the second DNN 300 obtained as a result of the training is stored in the AI decoding device 200.

[0228] Referring to Figure 10 In operations S840 and S845, the training device 1000 initially sets the DNN setting information of the first DNN 700 and the second DNN 300. Accordingly, the first DNN 700 and the second DNN 300 can operate according to the predetermined DNN setting information. The DNN setting information can include information on at least one of the number of convolution layers included in the first DNN 700 and the second DNN 300, the number of filter kernels for each convolution layer, the size of the filter kernel for each convolution layer, or the parameters of each filter kernel.

[0229] In operation S850, the training device 1000 inputs the original training image 801 into the first DNN 700. The original training image 801 can include at least one frame included in a still image or a moving image.

[0230] In operation S855, the first DNN 700 processes the original training image 801 according to the initially set DNN setting information and outputs the first training image 802 obtained by performing AI down-scaling on the original training image 801. S855 is originated from the first DNN 700. In Figure 10 In the embodiment, the first training image 802 output from the first DNN 700 is directly input to the second DNN 300, but the first training image 802 output from the first DNN 700 can be input to the second DNN 300 by the training device 1000. In addition, the training device 1000 can perform first encoding and first decoding on the first training image 802 via a specific codec and then input the second training image to the second DNN 300.

[0231] In operation S860, the second DNN 300 processes the first training image 802 or the second training image according to the initially set DNN setting information and outputs the third training image 804 obtained by performing AI up-scaling on the first training image 802 or the second training image.

[0232] In operation S865, the training device 1000 calculates the complexity loss information 820 based on the first training image 802.

[0233] The training device 1000 calculates the structure loss information 810 by comparing the reduced training image 803 with the first training image 802 at operation S870.

[0234] The training device 1000 calculates the quality loss information 830 by comparing the original training image 801 with the third training image 804 at operation S875.

[0235] The initially set DNN setting information is updated via a backpropagation process based on the final loss information at operation S880. The training device 1000 can calculate the final loss information for training the first DNN 700 based on the complexity loss information 820, the structure loss information 810, and the quality loss information 830.

[0236] The second DNN 300 updates the initially set DNN setting information via a backpropagation process based on the quality loss information 830 or the final loss information at operation S885. The training device 1000 can calculate the final loss information for training the second DNN 300 based on the quality loss information 830.

[0237] Then, the training device 1000, the first DNN 700, and the second DNN 300 can repeat operations S850 to S885 until the final loss information is minimized to update the DNN setting information. At this time, during each repetition, the first DNN 700 and the second DNN 300 operate according to the DNN setting information updated in the previous operation.

[0238] Table 1 below shows the effect when the original image 105 according to the embodiment of the disclosure is performed AI encoding and AI decoding and when the original image 105 is performed encoding and decoding via HEVC.

[0239]

Table 1

[0240]

[0241] As shown in Table 1, although the subjective image quality is higher when AI encoding and AI decoding are performed on the content including 300 frames of 8K resolution according to the embodiment of the disclosure than when encoding and decoding are performed via HEVC, the bit rate is reduced by at least 50%.

[0242] Figure 11 is a diagram of the device 20 for performing AI downscaling on the original image 105 and the device 40 for performing AI upscaling on the second image 135.

[0243] Device 20 receives the original image 105 and provides image data 25 and AI data 30 to device 40 using an AI reducer 1124 and a transform-based encoder 1126. According to an embodiment, image data 25 corresponds to... Figure 1 Image data, and AI data 30 corresponding to Figure 1 AI data. Furthermore, according to an embodiment, the transform-based encoder 1126 corresponds to... Figure 7 The first encoder 614, and the AI ​​reducer 1124 corresponding to Figure 7 AI Shrinker 612.

[0244] Device 40 receives AI data 30 and image data 25, and obtains a third image 145 by using a transform-based decoder 1146 and an AI amplifier 1144. According to an embodiment, the transform-based decoder 1146 corresponds to... Figure 2 The first decoder 232, and the AI ​​amplifier 1144 corresponding to Figure 2 AI amplifier 234.

[0245] According to an embodiment, device 20 includes a CPU, a memory, and a computer program including instructions. The computer program is stored in the memory. According to an embodiment, device 20 executes according to the CPU's execution of the computer program, referring to... Figure 11 The described functions. According to the embodiment, reference will be made to... Figure 11 The described functions are executed by a dedicated hardware chip and / or CPU.

[0246] According to an embodiment, device 40 includes a CPU, a memory, and a computer program including instructions. The computer program is stored in the memory. According to an embodiment, device 40 executes according to the CPU's execution of the computer program, as described above. Figure 11 The described functions. According to the embodiment, reference will be made to... Figure 11 The described functions are executed by a dedicated hardware chip and / or CPU.

[0247] exist Figure 11 In this configuration, the controller 1122 receives at least one input value 10. According to an embodiment, the at least one input value 10 may include at least one of the following: the target resolution difference between the AI ​​reducer 1124 and the AI ​​amplifier 1144; the bit rate of the image data 25; the bit rate type of the image data 25 (e.g., variable bit rate type, constant bit rate type, or average bit rate type); or the codec type of the transform-based encoder 1126. The at least one input value 10 may include a value pre-stored in the device 20 or a value input by the user.

[0248] The configuration controller 1122 controls the operation of the AI downscaler 1124 and the transform-based encoder 1126 based on the received input value 10. According to an embodiment, the configuration controller 1122 obtains DNN setting information for the AI downscaler 1124 according to the received input value 10, and sets the AI downscaler 1124 with the obtained DNN setting information. According to an embodiment, the configuration controller 1122 can transmit the received input value 10 to the AI downscaler 1124, and the AI downscaler 1124 can obtain DNN setting information for performing AI downscaling on the original image 105 based on the received input value 10. According to an embodiment, the configuration controller 1122 can provide additional information (e.g., color format (luminance component, chrominance component, red component, green component, or blue component) information to which AI downscaling is applied and tone mapping information for high dynamic range (HDR)) to the AI downscaler 1124 together with the input value 10, and the AI downscaler 1124 can obtain the DNN setting information in consideration of the input value 10 and the additional information. According to an embodiment, the configuration controller 1122 transmits at least a part of the received input value 10 to the transform-based encoder 1126, and the transform-based encoder 1126 performs first encoding on the first image 115 through a bit rate of a specific value, a specific type of bit rate, and a specific codec.

[0249] The AI downscaler 1124 receives the original image 105 and performs the operation described in at least one of Figure 1 、 Figure 7 、 Figure 8 、 Figure 9 or Figure 10 to obtain the first image 115.

[0250] According to an embodiment, the AI data 30 is provided to the device 40. The AI data 30 can include at least one of resolution difference information between the original image 105 and the first image 115 or information related to the first image 115. The resolution difference information can be determined based on a target resolution difference of the input value 10, and the information related to the first image 115 can be determined based on at least one of a target bit rate, a bit rate type, or a codec type. According to an embodiment, the AI data 30 can include a parameter used during AI upscaling. The AI data 30 can be provided to the device 40 from the AI downscaler 1124.

[0251] The image data 25 is obtained as the original image 105 is processed by the transform-based encoder 1126, and is transmitted to the device 40. The transform-based encoder 1126 can process the first image 115 according to MPEG-2, H.264 AVC, MPEG-4, HEVC, VC-1, VP8, VP9, or VA1.

[0252] The configuration controller 1142 controls the operation of the AI up-scaler 1144 based on the AI data 30. According to an embodiment, the configuration controller 1142 obtains DNN setting information for the AI up-scaler 1144 according to the received AI data 30, and sets the AI up-scaler 1144 with the obtained DNN setting information. According to an embodiment, the configuration controller 1142 can transmit the received AI data 30 to the AI up-scaler 1144, and the AI up-scaler 1144 can obtain DNN setting information for performing AI up-scaling on the second image 135 based on the AI data 30. According to an embodiment, the configuration controller 1142 can provide additional information (e.g., color format (luminance component, chrominance component, red component, green component, or blue component) information to which AI up-scaling is applied, and tone mapping information of HDR) to the AI up-scaler 1144 together with the AI data 30, and the AI up-scaler 1144 can obtain the DNN setting information in consideration of the AI data 30 and the additional information. According to an embodiment, the AI up-scaler 1144 can receive the AI data 30 from the configuration controller 1142, receive at least one of prediction mode information, motion information, or quantization parameter information from the transform-based decoder 1146, and obtain DNN setting information based on at least one of the prediction mode information, the motion information, and the quantization parameter information and the AI data 30.

[0253] The transform-based decoder 1146 can process the image data 25 to reconstruct the second image 135. The transform-based decoder 1146 can process the image data 25 according to MPEG-2, H.264 AVC, MPEG-4, HEVC, VC-1, VP8, VP9, or AV1.

[0254] The AI up-scaler 1144 can obtain a third image 145 by performing AI up-scaling on the second image 135 provided from the transform-based decoder 1146 based on the set DNN setting information.

[0255] The AI up-scaler 1144 can include a first DNN, and the AI up-scaler 1144 can include a second DNN, and according to an embodiment, the DNN setting information for the first DNN and the second DNN is trained according to the training method described with reference to Figure 9 and Figure 10

[0256] In addition, the AI decoding apparatus 200 shown in Figure 2 may receive broadcast (e.g., terrestrial broadcast, cable broadcast, or satellite broadcast) data or receive streaming content, perform AI decoding on the received broadcast data or streaming content, and display or externally output the AI-decoded image. However, when a dedicated media streaming hub (e.g., Firestick​TM or Chromecast TM ) receiving streaming content, the streaming hub can perform first decoding, and a separate device connected to the streaming hub can perform AI upscaling on a second image on which the first decoding is performed.

[0257] Accordingly, there is a need for a device for performing first decoding 130 and AI upscaling 140, respectively, and a method for connecting these devices to each other and transmitting and receiving AI data required for AI upscaling.

[0258] Figure 12 is a diagram of an AI decoding system 1001 according to an embodiment of the disclosure.

[0259] The AI decoding system 1001 according to an embodiment of the disclosure can include a decoding device 1100 and an AI upscaling device 1200.

[0260] The decoding device 1100 according to an embodiment of the disclosure can be a device that receives encoded data or an encoded signal from an external source, an external server, or an external device and decodes the encoded data or the encoded signal. The decoding device 1100 according to an embodiment of the disclosure can be implemented in the form of a set-top box or a dongle. However, embodiments are not limited thereto, and the decoding device 1100 can be implemented as any electronic device capable of receiving multimedia data from an external device.

[0261] The decoding device 1100 according to an embodiment of the disclosure can receive AI encoded data and perform first decoding based on the AI encoded data. The AI encoded data is data generated as a result of AI downscaling and first encoding of an original image, and can include image data and AI data. The decoding device 1100 can reconstruct a second image corresponding to the first image via first decoding of the image data. Here, the first image can be an image obtained by performing AI downscaling on the original image.

[0262] The first decoding according to an embodiment can include a process of generating quantized residual data by performing entropy decoding on image data, a process of dequantizing the quantized residual data, a process of generating prediction data, and a process of reconstructing a second image by using the prediction data and the residual data. The above-described first decoding can be performed via an image reconstruction method corresponding to one of image compression methods using frequency transformation, such as MPEG-2, H.264, MPEG-4, HEVC, VC-1, VP8, VP9, and AV1, which is used when first encoding is performed on the first image on which AI downscaling is performed.

[0263] The decoding device 1100 according to an embodiment of the disclosure can transmit the reconstructed second image and AI data included in the AI-encoded data to the AI upscaling device 1200. Here, the decoding device 1100 can transmit the second image and the AI data to the AI upscaling device 1200 via an input and output interface.

[0264] In addition, the decoding device 1100 can further transmit first decoding-related information such as mode information and quantization parameter information included in the image data to the AI upscaling device 1200 through the input and output interface.

[0265] For example, the decoding device 1100 and the AI upscaling device 1200 can be connected to each other via an HDMI cable or a display port (DP) cable, and the decoding device 1100 can transmit the second image and the AI data to the AI upscaling device 1200 through the HDMI or the DP.

[0266] The AI upscaling device 1200 according to an embodiment of the disclosure can perform AI upscaling on the second image by using the AI data received from the decoding device 1100. For example, the third image can be generated by performing AI upscaling on the second image via a second DNN.

[0267] In addition, the AI upscaling device 1200 according to an embodiment of the disclosure can be implemented as an electronic device including a display. For example, the AI upscaling device 1200 can be implemented as any one of various electronic devices such as a television (TV), a mobile phone, a tablet personal computer (PC), a digital camera, a camcorder, a laptop computer, a desktop computer, a compatible computer monitor, a video projector, a digital broadcasting terminal, a personal digital assistant (PDA), a portable multimedia player (PMP), and a navigation device.

[0268] When the AI upscaling device 1200 according to an embodiment of the disclosure includes a display, the AI upscaling device 1200 can display the second image or the third image on the display.

[0269] Figure 13 is a diagram of a configuration of the decoding device 1100 according to an embodiment of the disclosure.

[0270] Referring to Figure 13 , the decoding device 1100 according to an embodiment of the disclosure can include a receiver 1110, a first decoder 1120, and an input / output (I / O) device 1130.

[0271] The receiver 1110 according to an embodiment of the disclosure can receive AI encoded data generated as a result of AI encoding. The receiver 1110 can include a communication interface 1111, a parser 1112, and an output interface 1113. The receiver 1110 receives and parses AI encoded data generated as a result of AI encoding, divides the AI encoded data into image data and AI data, and outputs the image data to the first decoder 1120 and the AI data to the I / O interface 1130.

[0272] Specifically, the communication interface 1111 receives AI encoded data generated as a result of AI encoding data via a network. The AI encoded data generated as a result of AI encoding includes image data and AI data.

[0273] The AI data according to an embodiment of the disclosure can be received by being included in a video file together with the image data. When the AI data is included in the video file, the AI data can be included in the metadata of the header of the video file.

[0274] Alternatively, when the image data on which AI encoding is performed is received as a segment divided in a preset time unit, the AI data can be included in the metadata of the segment.

[0275] Alternatively, the AI data can be encoded and received by being included in a bitstream. Alternatively, the AI data can be received as a separate file.

[0276] Figure 13 A case in which the AI data is received in the form of metadata is illustrated.

[0277] The AI encoded data can be divided into image data and AI data. For example, the parser 1112 receives the AI encoded data through the communication interface 1111 and parses the AI encoded data to divide the AI encoded data into image data and AI data. For example, data received via a network can be configured in an MP4 file format, which conforms to the ISO base media file format standard widely used for storing or transmitting multimedia data. The MP4 file format includes a plurality of boxes, and each box can include type information indicating what data is contained and size information indicating the size of the box. Here, data received in the MP4 file format can include a media data box in which actual media data including image data is stored and a metadata box in which metadata related to the media is stored. By parsing the box type in the received data, it is determined whether the data is image data or AI data. For example, the parser 1112 distinguishes the image data and the AI data by identifying the box type of the data in the MP4 file format received via the communication interface 1111, and transmits the image data and the AI data to the output interface 1113, and the output interface 1113 transmits the image data and the AI data to the first decoder 1120 and the I / O interface 1130, respectively.

[0278] Here, the image data included in the AI encoded data can be identified as image data generated via a specific codec (for example, MPEG-2, H.264, MPEG-4, HEVC, VC-1, VP8, VP9, or AV1). In this case, the corresponding information can be transmitted to the first decoder 1120 through the output interface 1113 so that the image data is processed in the identified codec.

[0279] The AI encoded data according to an embodiment of the disclosure can be obtained from a data storage medium including a hard disk or the like, and the decoding device 1100 according to an embodiment of the disclosure can obtain the AI encoded data from the data storage medium through an input and output interface such as a universal serial bus (USB) port or the like.

[0280] In addition, the AI encoded data received from the communication interface 1111 can be stored in the memory, and the parser 1112 can parse the AI encoded data obtained from the memory. However, embodiments of the disclosure are not limited thereto.

[0281] The first decoder 1120 reconstructs a second image corresponding to the first image based on the image data. The second image generated by the first decoder 1120 is transmitted to the I / O interface 1130. According to an embodiment of the disclosure, first decoding-related information (such as mode information and quantization parameter information) included in the image data can be further transmitted to the I / O interface 1130.

[0282] I / O interface 1130 can receive AI data from output interface 1113.

[0283] I / O interface 1130 can send data to or receive data from external devices via input and output interfaces. For example, I / O interface 1130 can send and receive video data, audio data, and additional data. Optionally, I / O interface 1130 can request commands from or receive commands from external devices and send response messages regarding the commands. However, embodiments of this disclosure are not limited thereto.

[0284] Return to reference Figure 13 The I / O interface 1130 can send the second image received from the first decoder 1120 and the AI ​​data received from the output interface 1113 to the AI ​​amplification device 1200.

[0285] For example, I / O interface 1130 may include HDMI, and transmit the second image and AI data to AI amplification device 1200 via HDMI.

[0286] Optionally, the I / O interface 1130 may include a DisplayPort and transmit the second image and AI data to the AI ​​magnification device 1200 via the DisplayPort.

[0287] In the following text, reference will be made to Figure 14 Describe in detail the data structure of AI data in the form of metadata.

[0288] Figure 14 The data structure of AI data in the form of metadata according to an embodiment of this disclosure is shown.

[0289] AI data according to embodiments of this disclosure can be included in the metadata of the header of a video file or in the metadata of a segment. For example, when using the MP4 file format described above, the video file or segment may include a media data frame and a metadata frame, wherein the media data frame includes the actual media data, and the metadata frame includes media-related metadata. References can be sent in the metadata frame. Figure 14 AI data in the form of metadata.

[0290] Reference Figure 14 The AI ​​data according to the embodiment may include elements such as ai_codec_info 1300, ai_codec_applied_channel_info 1302, target_bitrate_info 1304, res_info 1306, ai_codec_DNN_info 1312 and ai_codec_supplementary_info 1314. Figure 14The arrangement of elements shown in the middle is merely an example, and one of ordinary skill in the art can change the arrangement of elements.

[0291] According to embodiments of the disclosure, ai_codec_info 1300 indicates whether AI upscaling is applied to a low-resolution image such as the second image 135. When the ai_codec_info 1300 indicates that AI upscaling is applied to the second image 135 reconstructed from the image data, the data structure of the AI data includes an element for obtaining upscaling DNN information for AI upscaling.

[0292] The ai_codec_applied_channel_info 1302 is channel information indicating a color channel to which AI upscaling is applied. An image can be represented in an RGB format, a YUV format, a YCbCr format, etc., and the color channel to which AI upscaling is needed can be indicated among YCbCr color channels, RGB color channels, or YUV color channels according to the type of frame.

[0293] The target_bitrate_info 1304 is information indicating a bitrate of image data obtained as a result of first encoding by the first encoder 614. The AI upscaler 234 can obtain AI upscaling DNN information suitable for the quality of the second image 135 according to the target_bitrate_info 1304.

[0294] The res_info 1306 indicates resolution information related to the resolution of a high-resolution image to which AI upscaling is performed, such as the third image 145. The res_info 1306 can include pic_width_org_luma 1308 and pic_height_org_luma 1310. The pic_width_org_luma 1308 and the pic_height_org_luma 1310 indicate the width and the height of the high-resolution image, respectively, and are high-resolution image width information and high-resolution image height information, respectively. The AI upscaler 234 can determine an AI upscaling ratio according to the resolution of the high-resolution image determined from the pic_width_org_luma 1308 and the pic_height_org_luma 1310 and the resolution of the low-resolution image reconstructed by the first decoder 232.

[0295] The ai_codec_DNN_info 1312 is information indicating mutually agreed AI upscaling DNN information used for AI upscaling of the second image 135. The AI upscaler 234 can determine the AI upscaling DNN information among a plurality of pieces of DNN setting information pre-stored according to the ai_codec_applied_channel_info 1302, the target_bitrate_info 1304, and the res_info 1306. Also, the AI upscaler 234 can determine the AI upscaling DNN information among the plurality of pieces of DNN setting information pre-stored by additionally considering other characteristics (type, maximum luminance, color gamut, etc.) of the image and codec information of encoding.

[0296] The DNN information indicating the AI upscaling DNN can be represented by an identifier indicating one of the plurality of pieces of DNN setting information pre-stored in the AI upscaler 234 as described above, or can include information on at least one of the number of convolution layers included in the DNN, the number of filter kernels for each convolution layer, or the parameters of each filter kernel.

[0297] The ai_codec_supplementary_info 1314 indicates supplementary information on AI upscaling. The ai_codec_supplementary_info 1314 can include information required to determine the AI upscaling DNN information applied to the video. The ai_codec_supplementary_info 1314 can include information on the genre, HDR maximum luminance, HDR color gamut, HDR PQ, codec, and rate control type.

[0298] In addition, when the AI data is included in the metadata of the segment, the AI data can further include dependent_ai_condition_info indicating dependency information.

[0299] The dependent_ai_condition_info indicates whether the current segment inherits the AI data of the previous segment.

[0300] For example, when the dependent_ai_condition_info indicates that the current segment inherits the AI data of the previous segment, the metadata of the current segment does not include the AI data corresponding to the above-described ai_codec_info 1300 to ai_codec_supplementary_info 1314. Instead, the AI data of the current segment is determined to be the same as the AI data of the previous segment.

[0301] Further, when the dependent_ai_condition_info indicates that the current segment does not inherit the AI data of the previous segment, the metadata of the current segment includes the AI data. Accordingly, the AI data related to the media data of the current segment can be obtained.

[0302] In addition, the receiver 1110 and the first decoder 1120 according to the embodiments of the disclosure are described as separate devices, but can be implemented via one processor. In this case, the receiver 1110 and the first decoder 1120 can be implemented via a separate dedicated processor, or can be implemented via a combination of software (S / W) and a general-purpose processor such as an application processor (AP), a central processing unit (CPU), or a graphic processor (GPU). In addition, the dedicated processor can be implemented by including a memory for implementing the embodiments of the disclosure or by including a memory processor for using an external memory.

[0303] In addition, the receiver 1110 and the first decoder 1120 can be implemented via one or more processors. In this case, the receiver 1110 and the first decoder 1120 can be implemented via a combination of dedicated processors, or can be implemented via a combination of S / W and a plurality of general-purpose processors such as an AP, a CPU, or a GPU.

[0304] Figure 15 is a diagram for describing a case in which AI data according to an embodiment of the disclosure is received in the form of a bitstream by being included in image data.

[0305] Because the above has been described with reference to Figure 13 the configuration of the communication interface 1111, the output interface 1113, the first decoder 1120, and the I / O interface 1130 will not be provided with the same description. Figure 15 With reference to

[0306] , the communication interface 1111 according to the embodiments of the disclosure can receive a bitstream in which image data and AI data are encoded together. Here, the AI data can be included in the bitstream in the form of a supplemental enhancement information (SEI) message, which is information capable of additionally enhancing the function of a codec for the first encoding and the first decoding. The SEI message can be transmitted in units of frames. Figure 15

[0307] ​When the AI coded data is received in the form of a bitstream in which the image data and the AI data are coded together, the image data and the AI data cannot be distinguished from each other. Accordingly, the communication interface 1111 transmits the AI coded data to the output interface 1113 in the form of a bitstream, and the output interface 1113 transmits the AI coded data to the first decoder 1120 in the form of a bitstream.

[0308] The first decoder 1120 reconstructs a second image corresponding to the first image based on the image data included in the bitstream received from the output interface 1113, and transmits the second image to the I / O interface 1130.

[0309] In addition, the first decoder 1120 separates a payload of an SEI message including the AI data from the bitstream, and transmits the payload to the I / O interface 1130.

[0310] The I / O interface 1130 can transmit the second image and the payload (e.g., AI data) of the SEI message received from the first decoder 1120 to the AI upscaling device 1200. In some embodiments, the AI upscaling device 1200 generates a third image. The third image can be displayed by the AI upscaling device 1200 or provided to a display device.

[0311] Here, as shown in Figure 16 , the AI data can be included in the SEI message in the form of high-level syntax.

[0312] Figure 16 An AI codec syntax table according to an embodiment of the present application is illustrated.

[0313] Referring to Figure 16 , the AI codec syntax table can include an AI codec main syntax table (ai_codec_usage_main). The AI codec main syntax table includes elements related to AI upscaling DNN information used for AI upscaling of a second image reconstructed from image data. The AI codec main syntax table can include AI data applied to AI upscaling of all frames in a video file.

[0314] According to the AI codec main syntax table Figure 16 , syntax elements such as ai_codec_info, ai_codec_applied_channel_info, target_bitrate, pic_width_org_luma, pic_height_org_luma, ai_codec_DNN_info, and ai_codec_supplementary_info_flag are parsed.

[0315] ai_codec_info corresponds to Figure 14 ai_codec_info 1300 of the second image, and indicates whether AI upscaling is allowed. When ai_codec_info indicates that AI upscaling is allowed (if(ai_codec_info)), syntax elements required to determine AI upscaling DNN information are parsed.

[0316] ai_codec_applied_channel_info is channel information corresponding to Figure 14 ai_codec_applied_channel_info 1302 of the second image. target_bitrate is target bitrate information corresponding to Figure 14 target_bitrate_info 1304. pic_width_org_luma and pic_height_org_luma are high resolution image width information and high resolution image height information, respectively, corresponding to Figure 14 pic_width_org_luma 1308 and pic_height_org_luma 1310. ai_codec_DNN_info is DNN information corresponding to Figure 14 ai_codec_DNN_info 1312.

[0317] ai_codec_supplementary_info_flag is a supplementary information flag indicating whether Figure 14 ai_codec_supplementary_info 1314 is included in the syntax table. When ai_codec_supplementary_info_flag indicates that supplementary information for AI upscaling is not parsed, additional supplementary information is not obtained. However, when ai_codec_supplementary_info_flag indicates that supplementary information for AI upscaling is parsed (if(ai_codec_supplementary_info_flag)), additional supplementary information is obtained.

[0318] The obtained additional supplemental information can include ai_codec_DNNstruct_info, genre_info, hdr_max_luminance, hdr_color_gamut, hdr_pq_type, and rate_control_type. The ai_codec_DNNstruct_info is information indicating a structure and parameters of new DNN setting information suitable for an image, which is separate from DNN setting information pre-stored in the AI up-scaler. For example, information about at least one of the number of convolution layers, the number of filter kernels of each convolution layer, or parameters of each filter kernel.

[0319] The genre_info indicates a genre of content of the image data, the hdr_max_luminance indicates a high dynamic range (HDR) maximum luminance applied to a high resolution image, the hdr_color_gamut indicates a HDR color gamut applied to the high resolution image, the hdr_pq_type indicates a HDR perceptual quantizer (PQ) information applied to the high resolution image, and the rate_control_type indicates a rate control type applied to the image data obtained as a result of the first encoding. According to an embodiment of the present application, a specific syntax element can be parsed from among syntax elements corresponding to the supplemental information.

[0320] In addition, the AI codec syntax table according to an embodiment of the present disclosure can include an AI codec frame syntax table (ai_codec_usage_frame) including AI data applied to a current frame.

[0321] Figure 17 is a block diagram of a configuration of an AI up-scaler 1200 according to an embodiment of the present disclosure.

[0322] Referring to Figure 17 , the AI up-scaler 1200 can include an input / output (I / O) interface 1210 and an AI up-scaler 1230.

[0323] The I / O interface 1210 can receive the second image and the AI data from the decoding device 1100. Here, the I / O interface 1210 can include an HDMI, a DP, or the like.

[0324] When the decoding device 1100 and the AI up-scaler 1200 according to an embodiment of the present disclosure are connected to each other via an HDMI cable, the I / O interface 1210 can receive the second image and the AI data through the HDMI.

[0325] Optionally, when the decoding device 1100 and the AI upscaling device 1200 according to the embodiments of the disclosure are connected to each other via a DP cable, the I / O interface 1210 can receive the second image and the AI data through the DP. However, embodiments of the disclosure are not limited thereto, and the second image and the AI data can be received via any one of various input and output interfaces. In addition, the I / O interface 1210 can receive the second image and the AI data via another manner of input and output interface.

[0326] The I / O interface 1210 can transmit the second image and the AI data to the AI upscaler 1230. When the AI data is transmitted, the AI upscaler 1230 according to the embodiments of the disclosure can determine a scaling target of the second image based on at least one of the difference information or the first image-related information included in the AI data. For example, based on the AI data described with reference to Figure 14 and Figure 16 , the scaling target of the second image can be determined.

[0327] The scaling target can indicate, for example, to what extent the second image is to be scaled. When the scaling target is determined, the AI upscaler 1230 can perform AI scaling on the second image via a second DNN to generate a third image corresponding to the scaling target. Because the method of performing AI scaling on the second image via the second DNN has been described in detail with reference to Figures 3 to 6 , a detailed description thereof will be omitted.

[0328] The AI upscaler 1230 according to the embodiments of the disclosure can obtain new DNN setting information based on the AI data rather than the DNN setting information pre-stored in the AI upscaling device 1200, and perform AI scaling on the second image by setting the second DNN with the obtained new DNN setting information.

[0329] In addition, when the I / O interface 1210 receives only the second image without receiving the AI data, the I / O interface 1210 can transmit the second image to the AI upscaler 1230. The AI upscaler 1230 can generate a fourth image by performing AI scaling on the second image according to a pre-set method without using the AI data. Here, the fourth image can have lower image quality than the third image on which AI scaling is performed by using the AI data.

[0330] Figure 18 is a diagram showing an example in which the decoding device 1100 and the AI upscaling device 1200 according to the embodiments of the disclosure transmit and receive data through HDMI.

[0331] The I / O interface 1130 of the decoding device 1100 and the I / O interface 1210 of the AI amplification device 1200 can be connected to each other via an HDMI cable. When the I / O interface 1130 of the decoding device 1100 and the I / O interface 1210 of the AI amplification device 1200 are connected to each other via the HDMI cable, pairing of four channels providing a TMDS data lane and a TMDS clock lane can be performed. The TMDS lane includes three data transmission lanes, and can be used to transmit video data, audio data, and additional data. Here, a packet structure is used to transmit the audio data and the additional data through the TMDS data lane

[0332] In addition, the I / O interface 1130 of the decoding device 1100 and the I / O interface 1210 of the AI amplification device 1200 can provide a display data channel (DDC). The DDC is a protocol standard defined by the Video Electronics Standards Association (VESA) for transmitting digital information between a computer graphics adapter and a monitor (e.g., a computer display device). The DDC is used for configuration and status information exchange between one source device (e.g., a decoding device) and one sink device (e.g., an AI amplification device). In some embodiments, the I / O interface 1210 is included in a display device such as a TV, a mobile phone, a tablet, etc.

[0333] Referring to Figure 18 , the I / O interface 1130 of the decoding device 1100 can include an HDMI transmitter 1610, a VSIF constructor 1620, and an extended display identification data (EDID) obtainer 1630. In addition, the I / O interface 1210 of the AI amplification device 1200 can include an HDMI receiver 1640 and an EDID memory 1650.

[0334] The EDID memory 1650 of the AI amplification device 1200 according to an embodiment of the disclosure can include EDID information. The EDID information is a data structure including various types of information about the AI amplification device 1200, and can be transmitted to the decoding device 1100 via the DDC.

[0335] The EDID information according to an embodiment of the disclosure can include information about an AI amplification capability of the AI amplification device 1200. For example, the EDID information can include information about whether the AI amplification device 1200 is capable of performing AI amplification. This will be described in detail with reference to Figure 19 Detailed Description.

[0336] Figure 19 is a diagram of a HDMI specification (HF) vendor specific data block (VSDB) included in the EDID information according to an embodiment of the disclosure.

[0337] The EDID information can include an EDID extension block including supplemental information. The EDID extension block can include the HF-VSDB 1710. The HF-VSDB 1710 is a data block in which vendor-specific data can be defined, and HDMI-specific data can be defined by using the HF-VSDB 1710.

[0338] The HF-VSDB 1710 according to an embodiment can include a reserved field 1720 and a reserved field 1730. Information about an AI upscaling capability of the AI upscaling device 1200 can be described by using at least one of the reserved field 1720 and the reserved field 1730 of the HF-VSDB 1710. For example, when the AI upscaling device is capable of performing AI upscaling by using a 1-bit reserved field, a bit value of the reserved field can be set to 1, and when the AI upscaling device is incapable of performing AI upscaling, the bit value of the reserved field can be set to 0. Alternatively, when the AI upscaling device is capable of performing AI upscaling, the bit value of the reserved field can be set to 0, and when the AI upscaling device is incapable of performing AI upscaling, the bit value of the reserved field can be set to 1.

[0339] Referring back to Figure 18 , the EDID obtainer 1630 of the decoding device 1100 can receive EDID information of the AI upscaling device 1200 through the DDC. The EDID information according to an embodiment of the disclosure can be transmitted as a HF-VSDB, and the EDID obtainer 1630 can obtain information about an AI upscaling capability of the AI upscaling device 1200 by using a reserved field value of the HF-VSDB.

[0340] The EDID obtainer 1630 can determine whether to transmit AI data to the AI upscaling device 1200 based on the information about the AI upscaling capability of the AI upscaling device 1200. For example, when the AI upscaling device 1200 is capable of performing AI upscaling, the EDID obtainer 1630 can operate such that the VSIF constructor 1620 constructs AI data in the form of a VSIF packet. On the other hand, when the AI upscaling device 1200 is incapable of performing AI upscaling, the EDID obtainer 1630 can operate such that the VSIF constructor 1620 does not construct AI data in the form of a VSIF packet.

[0341] The VSIF constructor 1620 can construct AI data transmitted from the first decoder 1120 or the output interface 1113 in the form of a VSIF packet. The VSIF packet will be described with reference to Figure 20 .

[0342] Figure 20 is a diagram of a header structure and a content structure of a VSIF according to an embodiment of the disclosure.

[0343] Referring toFigure 20 The VSIF packet includes a VSIF packet header 1810 and a VSIF packet content 1820. The VSIF packet header 1810 can include 3 bytes, in which a first byte HB0 is a value indicating a packet type, and a value of the VSIF packet is expressed as 0x81, a second byte HB1 indicates version information, and lower 6 bits of a third byte HB2 indicate a length of the VSIF packet content 1820 in bytes.

[0344] The VSIF constructor 1620 according to an embodiment of the disclosure can construct AI data in the form of a VSIF packet. For example, the VSIF constructor 1620 can generate a VSIF packet so that the VSIF packet includes AI data. The VSIF constructor 1620 can generate the VSIF packet content 1820 so that AI data described with reference to Figure 14 and 16 described in a reserved field value 1830 of a fifth packet byte PB5 and a reserved field value 1840 of an Nv-th packet byte PB(Nv) included in the VSIF packet content 1820. Alternatively, the VSIF packet content 1820 can be generated so that AI data is described in a reserved field value of an Nv+k-th packet byte, where k is an integer from 1 to n.

[0345] The VSIF constructor 1620 can determine a packet byte for describing AI data according to an amount of AI data. When the amount of AI data is small, AI data can be described by using only the reserved field value 1830 of the fifth packet byte PB5. On the other hand, when the amount of AI data is large, AI data can be described by using the reserved field value 1830 of the fifth packet byte PB5 and the reserved field value 1840 of the Nv-th packet byte PB(Nv). Alternatively, AI data can be described by using the reserved field value 1840 of the Nv-th packet byte PB(Nv) and a reserved field value of an Nv+k-th packet byte. However, embodiments of the disclosure are not limited thereto, and AI data can be constructed in the form of a VSIF packet via any one of various methods.

[0346] Figure 21 is a diagram illustrating an example of defining AI data in a VSIF packet according to an embodiment of the disclosure.

[0347] With reference to Figure 21Referring to FIG. 19, according to an embodiment of the disclosure, the VSIF constructor 1620 can describe AI data by using only the reserved field value of the fifth packet byte PB5.

[0348] In addition, ai_codec_org_width can be defined by using at least one of bits 4 to 5 of the fifth packet byte PB5, and ai_codec_org_height can be defined by using at least one of bits 6 to 7 of the fifth packet byte PB5. ai_codec_org_width represents the width of the original image 105 while representing the width of the third image 145. In addition, ai_codec_org_height represents the height of the original image 105 while representing the height of the third image 145. ai_codec_org_height and ai_codec_org_width are used to determine the size of the upscaling target.

[0349] Referring to FIG. 19, according to an embodiment of the disclosure, the VSIF constructor 1620 can describe AI data by using only the reserved field value of the fifth packet byte PB5. Figure 21 ai_codec_available_info can be defined by using at least one of bits 0 to 3 of the Nv-th packet byte PB(Nv).

[0350] In addition, ai_codec_DNN_info can be defined by using at least one of bits 4 to 7 of the Nv-th packet byte PB(Nv).

[0351]

[0352] ​In addition, ai_codec_org_width can be defined by using at least one of bits 0 to 3 of the Nv+1th packet byte PB(Nv+1), and ai_codec_org_height can be defined by using at least one of bits 4 to 7 of the Nv+1th packet byte PB(Nv+1).

[0353] In addition, bitrate_info can be defined by using at least one of bits 0 to 3 of the Nv+2th packet byte PB(Nv+2). bitrate_info is information indicating a degree of quality of a reconstructed second image.

[0354] In addition, ai_codec_applied_channel_info can be defined by using at least one of bits 4 to 7 of the Nv+2th packet byte PB(Nv+2). ai_codec_applied_channel_info is channel information indicating a color channel requiring AI upscaling. Depending on the type of frame, the color channel requiring AI upscaling can be indicated in a YCbCr color channel, an RGB color channel, or a YUV color channel.

[0355] In addition, ai_codec_supplementary_info can be defined by using at least one of bits of the remaining packet bytes (for example, bits included in the Nv+4th packet byte PB(Nv+4) to the Nv+nth packet byte PB(Nv+n)). ai_codec_supplementary_info indicates supplementary information for AI upscaling. The supplementary information can include structure and parameters of new DNN setting information suitable for a current image, a type, a color range, an HDR maximum luminance, an HDR color gamut, HDR PQ information, codec information, and a rate control (RC) type.

[0356] However, Figure 21 The structure of the VSIF packet shown in FIG. 1 is merely an example, and thus is not limited thereto. The positions or sizes of the fields in which AI data is defined in the VSIF packet of FIG. 1 can be changed as necessary, and the AI data described with reference to FIGS. 2 to 5 can be further included in the VSIF packet. Figure 19 The AI data defined in the VSIF packet of FIG. 1 can be further included in the VSIF packet. Figure 14 and 16 The AI data described with reference to FIGS. 2 to 5 can be further included in the VSIF packet.

[0357] Referring back to FIG. 1, Figure 18According to embodiments of the disclosure, the VSIF constructor 1620 can generate a VSIF packet corresponding to each of the plurality of frames. For example, when AI data is received once for the plurality of frames, the VSIF constructor 1620 can generate a VSIF packet corresponding to each of the plurality of frames by using the AI data received once. For example, the VSIF packets corresponding to the plurality of frames can be generated based on the same AI data.

[0358] On the other hand, when AI data is received multiple times for the plurality of frames, the VSIF constructor 1620 can generate a new VSIF packet by using the newly received AI data.

[0359] The VSIF constructor 1620 can transmit the generated VSIF packet to the HDMI transmitter 1610, and the HDMI transmitter 1610 can transmit the VSIF packet to the AI amplification device 1200 through the TMDS channel.

[0360] In addition, the HDMI transmitter 1610 can transmit the second image received from the first decoder 1120 to the AI amplification device 1200 through the TMDS channel.

[0361] The HDMI receiver 1640 of the AI amplification device 1200 can receive AI data and a second image configured in the form of a VSIF packet through the TMDS channel.

[0362] The HDMI receiver 1640 of the AI amplification device 1200 according to embodiments of the disclosure can determine whether AI data is included in a VSIF packet by searching for the VSIF packet after checking header information of an HDMI packet.

[0363] For example, the HDMI receiver 1640 can determine whether the received HDMI packet is a VSIF packet by determining whether a first byte HB0 indicating a packet type among the header information of the received HDMI packet is 0x81. In addition, when it is determined that the HDMI packet is a VSIF packet, the HDMI receiver 1640 can determine whether AI data is included in the VSIF packet content. For example, when the value of a bit is set, the HDMI receiver 1640 can obtain AI data by using the value of the bit included in the Nvth packet byte PB(Nv) to the Nv+nth packet byte PB(Nv+n) included in the VSIF packet content. For example, the HDMI receiver 1640 can obtain ai_codec_available_info by using at least one of bits 0 to 3 of the Nvth packet byte PB(Nv) of the VSIF packet content, and obtain ai_codec_DNN_info by using at least one of bits 4 to 7 of the Nvth packet byte PB(Nv).

[0364] In addition, the HDMI receiver 1640 can obtain ai_codec_org_width by using at least one of bits 0 to 3 of the Nv+1 block byte PB (Nv+1), and obtain ai_codec_org_height by using at least one of bits 4 to 7 of the Nv+1 block byte PB (Nv+1).

[0365] In addition, the HDMI receiver 1640 can obtain bitrate_info by using at least one of bits 0 to 3 of the Nv+2 block byte PB (Nv+2), and ai_codec_applied_channel_info by using at least one of bits 4 to 7 of the Nv+2 block byte PB (Nv+2).

[0366] In addition, the HDMI receiver 1640 can obtain ai_codec_supplementary_info by using at least one bit of the remaining packet bytes (e.g., bits included in the Nv+4th packet byte PB(Nv+4) to the Nv+nth packet byte PB(Nv+n)).

[0367] The HDMI receiver 1640 can provide AI data obtained from VSIF packet content to the AI ​​amplifier 1230, and also provide a second image to the AI ​​amplifier 1230.

[0368] Upon receiving the second image and AI data from the HDMI receiver 1640, the AI ​​amplifier 1230 according to embodiments of the present disclosure may determine a magnification target for the second image based on at least one of difference information or first image-related information included in the AI ​​data. The magnification target may indicate, for example, the extent to which the second image will be magnified. When the magnification target is determined, the AI ​​amplifier 1230 may perform AI magnification on the second image via a second DNN to generate a third image corresponding to the magnification target. Because reference has already been made... Figures 3 to 6 The method for performing AI upscaling on the second image via a second DNN is described in detail, so its detailed description will be omitted.

[0369] In addition, Figures 18 to 21 In this embodiment, the decoding device 1100 and the AI ​​amplification device 1200 are connected to each other via an HDMI cable, but the embodiment is not limited to this. Furthermore, according to embodiments of this disclosure, the decoding device 1100 and the AI ​​amplification device 1200 may be connected via a DP cable. When the decoding device 1100 and the AI ​​amplification device 1200 are connected to each other via a DP cable, the decoding device 1100 may send the second image and AI data to the AI ​​amplification device 1200 via DP in a manner similar to HDMI.

[0370] In addition, the decoding device 1100 according to an embodiment of the disclosure can transmit the second image and the AI data to the AI upscaling device 1200 via an input and output interface other than HDMI or DP.

[0371] In addition, the decoding device 1100 according to an embodiment of the disclosure can transmit the second image and the AI data to the AI upscaling device 1200 via different interfaces. For example, the second image can be transmitted via HDMI, and the AI data can be transmitted via DP. Alternatively, the second image can be transmitted via DP, and the AI data can be transmitted via HDMI.

[0372] Figure 22 is a flowchart of an operation method of the decoding device 1100 according to an embodiment of the disclosure.

[0373] Referring to Figure 22 At operation S2010, the decoding device 1100 according to an embodiment of the disclosure can receive AI encoded data.

[0374] For example, the decoding device 1100 receives AI encoded data generated as a result of AI encoding via a network. The AI encoded data is data generated as a result of AI downscaling and first encoding of an original image, and can include image data and AI data.

[0375] Here, the AI data according to an embodiment of the disclosure can be received by being included in a video file together with the image data. When the AI data is included in the video file, the AI data can be included in metadata of a header of the video file. Alternatively, when the image data on which AI encoding is performed is received as a segment divided in a preset time unit, the AI data can be included in metadata of the segment. Alternatively, the AI data can be encoded and received by being included in a bitstream, or can be received as a file separate from the image data. However, embodiments of the disclosure are not limited thereto.

[0376] At operation S2020, the decoding device 1100 can divide the AI encoded data into image data and AI data.

[0377] When the AI data according to an embodiment of the disclosure is received in the form of metadata of a header of a video file or metadata of a segment, the decoding device 1100 can parse the AI encoded data and divide the AI encoded data into image data and AI data. For example, the decoding device 1100 can read box type data received through a network to determine whether the data is image data or AI data.

[0378] When the AI data according to the embodiment of the disclosure is received in the form of a bitstream, the decoding device 1100 can receive a bitstream in which the image data and the AI data are encoded together. Here, the AI data can be inserted in the form of an SEI message. The decoding device 1100 can distinguish the payload including the image data and the SEI message of the AI data from the bitstream.

[0379] In operation S2030, the decoding device 1100 according to the embodiment of the disclosure can decode the second image based on the image data.

[0380] In operation S2040, the decoding device 1100 according to the embodiment of the disclosure can transmit the second image and the AI data to an external device through an input and output interface.

[0381] The external device according to the embodiment of the disclosure includes an AI upscaling device 1200.

[0382] For example, the decoding device 1100 can transmit the second image and the AI data to the external device via HDMI or DP. When the AI data is transmitted via HDMI, the decoding device 1100 can transmit the AI data in the form of a VSIF packet.

[0383] Further, the transmitted AI data includes information that causes the second image to be AI-upscaled. For example, the AI data can include information indicating whether AI upscaling is applied to the second image, information about a DNN used to upscale the second image, etc.

[0384] Figure 23 is a flowchart of a method of transmitting the second image and the AI data via HDMI, which is performed by the decoding device 1100 according to the embodiment of the disclosure.

[0385] Referring to Figure 23 In operation S2110, the decoding device 1100 according to the embodiment of the disclosure can be connected to the AI upscaling device 1200 via an HDMI cable.

[0386] In operation S2120, the decoding device 1100 can transmit an EDID information request to the AI upscaling device 1200 via DDC. In response to the EDID information request of the decoding device 1100, the AI upscaling device 1200 can transmit EDID information stored in an EDID memory to the decoding device 1100 via DDC (S2130). Here, the EDID information can include HF-VSDB, and the HF-VSDB can include information about AI upscaling capability of the AI upscaling device 1200.

[0387] The decoding device 1100 can obtain information about the AI upscaling capability of the AI upscaling device 1200 by receiving the EDID information (e.g., HF-VSDB).

[0388] The decoding device 1100 can determine whether to transmit AI data to the AI upscaling device 1200 based on the information about the AI upscaling capability of the AI upscaling device 1200, at operation S2150. For example, when the AI upscaling device 1200 is incapable of performing AI upscaling, the decoding device 1100 can not construct AI data in the form of a VSIF packet, but can transmit only the second image to the AI upscaling device 1200 via a TMDS channel, at operation S2160.

[0389] When the AI upscaling device 1200 is capable of performing AI upscaling, the decoding device 1100 can operate to construct AI data in the form of a VSIFD packet, at operation S2170.

[0390] The decoding device 1100 can define AI data by using a value of a reserved field included in the VSIF packet. Because the method of defining AI data in the VSIF packet has been described with reference to Figure 20 and 21 The method of defining AI data in the VSIF packet has been described in detail, and thus a description thereof will not be provided again.

[0391] The decoding device 1100 can transmit the second image and the AI data constructed in the VSIF packet to the AI upscaling device 1200 via a TMDS channel, at operation S2180.

[0392] Figure 24 is a flowchart of an operation method of an AI upscaling device 1200 according to an embodiment of the disclosure.

[0393] Referring to Figure 24 The AI upscaling device 1200 according to an embodiment of the disclosure can receive a second image and AI data via an input and output interface, at operation S2210.

[0394] For example, when connected to the decoding device 1100 via an HDMI cable, the AI upscaling device 1200 can receive a second image and AI data via HDMI. Alternatively, when connected to the decoding device 1100 via a DP cable, the AI upscaling device 1200 can receive a second image and AI data via DP. However, embodiments of the disclosure are not limited thereto, and can receive a second image and AI data via any one of various input and output interfaces. Furthermore, the AI upscaling device 1200 can receive a second image and AI data via another manner of input and output interface.

[0395] The AI upscaling device 1200 can determine whether to perform AI upscaling on a second image based on whether AI data is received through the input and output interface. When AI data is not received, a second image can be output without performing AI upscaling on the second image.

[0396] The AI up-scaling device 1200 according to an embodiment of the disclosure can receive the HDMI packet from the decoding device 1100 and search for the VSIF packet by recognizing the header information of the HDMI packet. When the VSIF packet is found, the AI up-scaling device 1200 can determine whether the AI data is included in the VSIF packet.

[0397] The AI data can include information indicating whether AI up-scaling is applied to the second image, information about a DNN used to upscale the second image, and the like.

[0398] The AI up-scaling device 1200 can determine whether to perform AI up-scaling on the second image based on the information indicating whether AI up-scaling is applied to the second image.

[0399] Further, the AI up-scaling device 1200 can obtain information about a DNN used to perform up-scaling on the second image based on the AI data, at operation S2220, and generate a third image by performing AI up-scaling on the second image using the DNN determined according to the obtained information, at operation S2230.

[0400] Figure 25 is a block diagram of a configuration of a decoding device 2300 according to an embodiment of the disclosure.

[0401] Figure 25 The decoding device 2300 of Figure 12 is an example of the decoding device 1100.

[0402] Referring to Figure 25 , the decoding device 2300 according to an embodiment of the disclosure can include a communication interface 2310, a processor 2320, a memory 2330, and an input / output (I / O) interface 2340.

[0403] Figure 25 The communication interface 2310 of Figure 13 and Figure 15 may correspond to the communication interface 1111 of Figure 25 The I / O interface 2340 of Figure 13 and Figure 15 may correspond to the I / O interface 1130 of Figure 13 and Figure 15 . Thus, the same description about Figure 25 as described with reference to

[0404] The communication interface 2310 according to an embodiment of the disclosure can transmit and receive data or a signal to and from an external device (e.g., a server) under the control of the processor 2320. The processor 2320 can transmit and receive content to and from an external device connected via the communication interface 2310. According to the performance and structure of the decoding device 2300, the communication interface 2310 can include one of a wireless local area network (LAN) 2311 (e.g., Wi-Fi), Bluetooth 2312, and a wired Ethernet 2313. Alternatively, the communication interface 2310 can include a combination of the wireless LAN 2311, the Bluetooth 2312, and the wired Ethernet 2313.

[0405] The communication interface 2310 according to an embodiment of the disclosure can receive AI encoded data generated as a result of AI encoding. The AI encoded data is data generated as a result of AI downscaling and first encoding of an original image, and can include image data and AI data.

[0406] The processor 2320 according to an embodiment of the disclosure can control the decoding device 2300 as a whole. The processor 2320 according to an embodiment of the disclosure can execute one or more programs stored in the memory 2330.

[0407] The memory 2330 according to an embodiment of the disclosure can store various types of data, programs, or applications for driving and controlling the decoding device 2300. Furthermore, the memory 2330 can store AI encoded data received according to an embodiment of the disclosure or a second image obtained via first decoding. The program stored in the memory 2330 can include one or more instructions. The program (one or more instructions) or the application stored in the memory 2330 can be executed by the processor 2320.

[0408] The processor 2320 according to an embodiment of the disclosure can include a CPU 2321, a GPU 2323, and a video processing unit (VPU) 2322. Alternatively, according to an embodiment of the disclosure, the CPU 2321 can include the GPU 2323 or the VPU 2322. Alternatively, the CPU 2321 can be implemented in the form of a system on chip (SoC) in which at least one of the GPU 2323 or the VPU 2322 is integrated. Alternatively, the GPU 2323 and the VPU 2322 can be integrated.

[0409] The processor 2320 can perform functions of controlling the overall operation of the decoding device 2300 and signal flow between internal components of the decoding device 2300 and processing data. The processor 2320 can control the communication interface 2310 and the I / O interface 2340. The GPU 2323 can perform graphics processing and can generate a screen including various objects such as icons, images, and text. The VPU 2322 can perform processing on image data or video data received by the decoding device 2300 and perform various image processing such as decoding (e.g., first decoding), scaling, noise filtering, frame rate conversion, resolution conversion, etc. on the image data or video data.

[0410] The processor 2320 according to an embodiment of the disclosure can perform at least one of the operations of the parser 1112, the output interface 1113, and the first decoder 1120 described with reference to FIGS. 11A and 11B, or can control at least one of the operations to be performed. Figure 13 The processor 2320 according to an embodiment of the disclosure can perform at least one of the operations of the parser 1112, the output interface 1113, and the first decoder 1120 described with reference to FIGS. 11A and 11B, or can control at least one of the operations to be performed. Figure 15 The processor 2320 according to an embodiment of the disclosure can perform at least one of the operations of the parser 1112, the output interface 1113, and the first decoder 1120 described with reference to FIGS. 11A and 11B, or can control at least one of the operations to be performed. Figure 18 The processor 2320 according to an embodiment of the disclosure can perform at least one of the operations of the VSIF constructor 1620 and the EDID obtainer 1630 described with reference to FIGS. 16A and 16B, or can control at least one of the operations to be performed.

[0411] For example, the processor 2320 can divide AI encoded data received by the communication interface 2310 into image data and AI data. When the AI data is received in the form of metadata of a header of a video file or metadata of a segment, the processor 2320 can parse the AI encoded data and divide the AI encoded data into image data and AI data. For example, when AI encoded data according to an embodiment of the disclosure is configured in the form of an MP4 file, the processor 2320 can parse received data of a box type configured in the form of an MP4 file to determine whether the data is image data or AI data.

[0412] Further, when the AI data is received in the form of a bitstream, the AI data can be included in the bitstream in the form of an SEI message, and the processor 2320 can distinguish a payload of the SEI message including the image data and the AI data from the bitstream.

[0413] The processor 2320 can control the GPU 2323 or the VPU 2322 to reconstruct a second image corresponding to the first image based on the image data.

[0414] The processor 2320 can transmit the second image and the AI data to an external device via the I / O interface 2340. The I / O interface 2340 can transmit or receive video, audio, and supplementary information to or from the outside of the decoding device 2300 under the control of the processor 2320. The I / O interface 2340 can include an HDMI 2341, a DP 2342, and a USB 2343. It will be apparent to those of ordinary skill in the art that the configuration and operation of the I / O interface 2340 will be implemented in various ways according to embodiments of the disclosure.

[0415] For example, when the decoding device 2300 and the AI upscaling device are connected to each other via the HDMI 2341, the I / O interface 2340 can receive EDID information of the AI upscaling device via DDC. Also, the processor 2320 can parse the EDID information of the AI upscaling device received via DDC. The EDID information can include information on AI upscaling capability of the AI upscaling device.

[0416] The processor 2320 can determine whether to construct the AI data in the form of a VSIF packet based on the information on the AI upscaling capability of the AI upscaling device. For example, when the AI upscaling device is capable of performing AI upscaling, the processor 2320 can control to construct the AI data in the form of a VSIF packet, and when the AI upscaling device is not capable of performing AI upscaling, the processor 2320 can control not to perform the operation of constructing the AI data in the form of a VSIF packet.

[0417] The I / O interface 2340 can construct the AI data in the form of a VSIF packet under the control of the processor 2320, and transmit the AI data constructed in the VSIF packet and the second image to the AI upscaling device via a TMDS channel.

[0418] The AI data according to an embodiment of the disclosure includes information that causes the second image to be AI-upscaled. For example, the AI data can include information indicating whether AI upscaling is applied to the second image, information on a DNN used to upscale the second image, or the like.

[0419] Figure 26 is a block diagram of a configuration of an AI upscaling device 2400 according to an embodiment of the disclosure.

[0420] Figure 26 The AI upscaling device 2400 of Figure 12 is an example of the AI upscaling device 1200.

[0421] Referring to Figure 26 The AI upscaling device 2400 according to an embodiment of the disclosure can include an input / output (I / O) interface 2410, a processor 2420, a memory 2430, and a display 2440.

[0422] The I / O interface 2410 of the AI upscaling device 2400 can correspond to the I / O interface 1210 of the decoding device 1200. Figure 26 Figure 17 Therefore, the same description about the decoding device 1200 described with reference to the description of the decoding device 1200 will not be provided again. Figure 17 Figure 26 The I / O interface 2410 of the AI upscaling device 2400 according to an embodiment of the disclosure can receive or transmit video, audio, and supplementary information from or to the outside of the AI upscaling device 2400 under the control of the processor 2420. The I / O interface 2340 can include an HDMI 2411, a DP 2412, and a USB 2413. It will be apparent to those of ordinary skill in the art that the configuration and operation of the I / O interface 2410 will be implemented in various ways according to an embodiment of the disclosure.

[0423] For example, when the decoding device and the AI upscaling device 2400 are connected to each other via the HDMI 2411, the I / O interface 2410 can transmit EDID information of the AI upscaling device 2400 to the decoding device when an EDID information reading request is received via a DDC. Also, the I / O interface 2410 can receive a second image and AI data configured in the form of a VSIF packet via a TMDS channel.

[0424] Alternatively, the I / O interface 2410 can receive a second image and AI data via a DP. However, embodiments of the disclosure are not limited thereto, and the second image and the AI data can be received via any one of various input and output interfaces. Alternatively, the I / O interface 2410 can receive a second image and AI data via another way of input and output interface.

[0425] The processor 2420 according to an embodiment of the disclosure can control the AI upscaling device 2400 as a whole. The processor 2420 according to an embodiment of the disclosure can execute one or more programs stored in the memory 2430.

[0426] The memory 2430 according to an embodiment of the disclosure can store various types of data, programs, or applications for driving and controlling the AI upscaling device 2400. For example, the memory 2430 can store EDID information of the AI upscaling device 2400. The EDID information can include various types of information about the AI upscaling device 2400, and in particular, can include information about AI upscaling capabilities of the AI upscaling device 2400. The program stored in the memory 2430 can include one or more instructions. The program (one or more instructions) or the application stored in the memory 2430 can be executed by the processor 2420.

[0427] The memory 2430 according to an embodiment of the disclosure can store various types of data, programs, or applications for driving and controlling the AI upscaling device 2400. For example, the memory 2430 can store EDID information of the AI upscaling device 2400. The EDID information can include various types of information about the AI upscaling device 2400, and in particular, can include information about AI upscaling capabilities of the AI upscaling device 2400. The program stored in the memory 2430 can include one or more instructions. The program (one or more instructions) or the application stored in the memory 2430 can be executed by the processor 2420.

[0428] ​The processor 2420 according to an embodiment of the disclosure can include a CPU 2421, a GPU 2423, and a VPU 2422. Alternatively, the CPU 2421 can include the GPU 2423 or the VPU 2422 according to an embodiment of the disclosure. Alternatively, the CPU 2421 can be implemented in the form of a SoC that integrates at least one of the GPU 2423 or the VPU 2422. Alternatively, the GPU 2423 and the VPU 2422 can be integrated. Alternatively, the processor 2420 can further include a neural processing unit (NPU).

[0429] The processor 2420 can perform a function of controlling overall operations of the AI up-scaling device 2400 and a signal flow between internal components of the AI up-scaling device 2400 and processing data. The processor 2420 can control the I / O interface 2410 and the display 2440. The GPU 2423 can perform a graphic processing, and can generate a screen including various objects such as icons, images, and texts. The VPU 2422 can perform a processing on image data or video data received by the AI up-scaling device 2400, and perform various image processing such as decoding (e.g., first decoding), scaling, noise filtering, frame rate conversion, resolution conversion, etc. on the image data or the video data. The processor 2420 according to an embodiment of the disclosure can perform at least one operation of the AI upscaler 1230 described above with reference to FIG. 1, or can control to perform at least one operation. Figure 17

[0430] For example, the processor 2420 can perform AI upscaling on the second image based on whether AI data is received through the I / O interface 2410.

[0431] The processor 2420 can search for a VSIF packet by recognizing header information of an HDMI packet received by the I / O interface 2410. When the VSIF packet is found, the processor 2420 can determine whether AI data is included in the VSIF packet. The AI data can include information indicating whether AI upscaling is applied to the second image, information about a DNN used to upscale the second image, etc.

[0432] In addition, the processor 2420 can determine whether to perform AI upscaling on the second image based on the information indicating whether to apply AI upscaling to the second image.

[0433] The processor 2420 can obtain information about a DNN used to upscale the second image based on the AI data, and generate a third image by performing AI upscaling on the second image using the DNN determined according to the obtained information. The processor 2420 can control the NPU to perform AI upscaling on the second image using the determined DNN.

[0434] ​When the I / O interface 2410 does not receive the AI data, the processor 2420 can generate a fourth image by performing AI upscaling on the second image according to a preset method without using the AI data. Here, the fourth image can have a lower image quality than the third image on which AI upscaling is performed by using the AI data.

[0435] The display 2440 generates a driving signal by converting an image signal, a data signal, an OSD signal, or a control signal processed by the processor 2420. The display 2440 can be implemented as a plasma display panel (PDP), a liquid crystal display (LCD), an organic light emitting diode (OLED), a flexible display, or the like, or can be implemented as a three-dimensional (3D) display. Furthermore, in addition to an output device, the display 2440 can be configured as a touch screen to function as an input device. The display 2440 can display the third image or the fourth image.

[0436] In addition, in the block diagrams of the decoding device 2300 and the AI upscaling device 2400 shown in FIGS. 21A and 21B, the components of the block diagrams can be integrated or omitted, or other components can be added to the block diagrams. In other words, two or more components can be integrated into one component, or one component can be divided into two or more components. Further, the functions performed in each block are used to describe the embodiments of the disclosure, and a specific operation or device does not limit the scope of the disclosure. Figure 25 Figure 26 The decoding device according to the embodiments of the disclosure can efficiently transmit AI data and a reconstructed image to the AI upscaling device via an input and output interface.

[0437] The AI upscaling device according to the embodiments of the disclosure can efficiently receive AI data and a reconstructed image from the decoding device via an input and output interface.

[0438] The AI upscaling device according to the embodiments of the disclosure can efficiently receive AI data and a reconstructed image from the decoding device via an input and output interface.

[0439] In addition, the above-described embodiments of the disclosure can be written as computer executable programs or instructions that can be stored in a medium.

[0440] ​The medium can continuously store computer-executable programs or instructions, or temporarily store computer-executable programs or instructions for execution or download. In addition, the medium can be any one of various recording media or storage media combined with a single piece of hardware or multiple pieces of hardware, and the medium is not limited to a medium directly connected to a computer system, but can be distributed over a network. Examples of the medium include magnetic media (such as a hard disk, a floppy disk, and a magnetic tape), optical recording media (such as a CD-ROM and a DVD), magneto-optical media (such as an optical floppy disk), and ROM, RAM, and flash memory configured to store program instructions. Other examples of the medium include recording media and storage media managed by an application store that publishes applications, or by a website, a server, or the like that provides or publishes other various types of software.

[0441] In addition, a model related to the above-described DNN can be implemented via a software module. When the DNN model is implemented via a software module (e.g., a program module including instructions), the DNN model can be stored in a computer-readable recording medium.

[0442] In addition, the DNN model can be integrated in the form of a hardware chip to be a part of the above-described AI amplification device 1200. For example, the DNN model can be manufactured in the form of a dedicated hardware chip for AI, or can be manufactured as a part of an existing general-purpose processor (e.g., a CPU or an application processor) or a graphics-dedicated processor (e.g., a GPU).

[0443] In addition, the DNN model can be provided in the form of downloadable software. A computer program product can include a product (e.g., a downloadable application) in the form of a software program electronically published by a manufacturer or an electronic market. For electronic publication, at least a part of the software program can be stored on a storage medium or can be temporarily generated. In this case, the storage medium can be a server of the manufacturer or the electronic market, or a storage medium of a relay server. Although one or more embodiments of the present disclosure have been described with reference to the accompanying drawings, it will be understood by those of ordinary skill in the art that various changes in form and details can be made therein without departing from the spirit and scope as defined by the following claims.

Claims

1. An electronic device comprising: a communication interface configured to receive artificial intelligence (AI) encoded data, wherein the AI encoded data is generated by AI down-scaling of an original image followed by encoding; one or more processors configured to: obtain AI data related to AI down-scaling of the original image to a first image from the received AI encoded data, wherein the AI data includes an identifier indicating first neural network (NN) setting information used for the AI down-scaling; obtain image data corresponding to an encoding result on the first image from the received AI encoded data; obtain a second image by decoding the obtained image data; and an input / output (I / O) interface, wherein the one or more processors are further configured to: control the I / O interface to receive extended display identification data (EDID) information from an external device, determine whether to transmit the AI data to the external device based on information on AI up-scaling capability of the external device included in the EDID information, when the external device is incapable of performing AI up-scaling, control the I / O interface to transmit the second image to the external device, and when the external device is capable of performing AI up-scaling, control the I / O interface to transmit the second image and the AI data to the external device, wherein the identifier is used to select second NN setting information from a plurality of second NN setting information for an up-scaling NN, wherein the external device includes the up-scaling NN, and wherein the first image is obtained by a down-scaling NN configured with first NN setting information selected from a plurality of first NN setting information for AI down-scaling. 2.The electronic device of claim 1, wherein, the I / O interface includes a high-definition multimedia interface (HDMI), and the one or more processors are further configured to transmit the second image and the AI data to the external device through the HDMI. 3.The electronic device of claim 2, wherein, the one or more processors are further configured to transmit the AI data in a vendor specific information frame (VSIF) packet. 4.The electronic device of claim 1, wherein the I / O interface includes a display port (DP), and the one or more processors are further configured to transmit the second image and the AI data to the external device through the DP. 5.The electronic device of claim 1, wherein the AI data includes first information indicating that the second image has been AI down-scaled. 6.The electronic device of claim 1, wherein the AI data indicates one or more color channels to which AI up-scaling is to be applied. 7.The electronic device of claim 1, wherein the AI data indicates at least one of high dynamic range (HDR) maximum brightness, HDR color gamut, HDR PQ, codec, or rate control type. 8.The electronic device of claim 1, wherein the AI data indicates a width resolution of the original image and a height resolution of the original image. 9.An operating method of an electronic device, the operating method comprising: receiving artificial intelligence (AI) encoded data, the AI encoded data being generated by AI down-scaling of an original image followed by encoding; obtaining AI data related to AI downscaling of the original image to a first image from the received AI encoded data, wherein the AI data includes an identifier indicating first neural network (NN) setting information for AI downscaling; obtaining image data corresponding to an encoding result with respect to the first image from the received AI encoded data; obtaining a second image by decoding the obtained image data; receiving extended display identification data (EDID) information from an external device through an input / output (I / O) interface; determining whether to transmit the AI data to the external device based on information about AI upscaling capability of the external device included in the EDID information, when the external device is incapable of performing AI upscaling, transmitting the second image to the external device through the I / O interface, and when the external device is capable of performing AI upscaling, transmitting the second image and the AI data to the external device through the I / O interface, wherein the identifier is used to select second NN setting information from among a plurality of second NN setting information for an upscaling NN included in the external device, and wherein the first image is obtained through a downscaling NN configured with first NN setting information selected from among a plurality of first NN setting information for AI downscaling.

10. The operating method of claim 9, wherein, The step of transmitting the second image and the AI data to the external device includes transmitting the second image and the AI data to the external device through a high-definition multimedia interface (HDMI).

11. The operating method of claim 10, wherein, The step of transmitting the second image and the AI data to the external device includes transmitting the AI data in the form of a vendor specific information frame (VSIF) packet.

12. The operating method of claim 9, wherein, The step of transmitting the second image and the AI data to the external device includes transmitting the second image and the AI data to the external device through a display port (DP).

13. The operating method of claim 9, wherein, The AI data includes first information indicating that the second image has been AI downscaled. The AI data includes first information indicating that the second image has been AI downscaled.