Image providing apparatus using artificial intelligence and method thereby, and display apparatus using artificial intelligence and method thereby

AI-based image processing using jointly trained neural networks for downscaling and upscaling addresses the challenge of high-resolution video encoding and decoding, achieving efficient low-bitrate processing and cross-device compatibility.

KR102996222B1Active Publication Date: 2026-07-27SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2021-08-13
Publication Date
2026-07-27

AI Technical Summary

Technical Problem

The increasing demand for high-resolution/high-quality video encoding and decoding poses challenges in preventing storage capacity saturation and maintaining compatibility between devices with and without AI functions.

Method used

An AI-based image providing device and display device utilize neural networks for downscaling and upscaling images, with jointly trained neural networks to maintain compatibility and reduce bit rate.

Benefits of technology

The solution enables efficient image processing at a low bit rate, preventing storage capacity saturation and ensuring compatibility across devices with and without AI capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure R1020210107626_ABST
    Figure R1020210107626_ABST
Patent Text Reader

Abstract

An electronic device that provides an image using AI is disclosed, comprising a processor that executes one or more instructions stored in the electronic device, wherein the processor obtains a first image by performing AI downscaling on an original image through a neural network for downscaling, obtains first image data by encoding the first image, and if the display device does not support an AI upscaling function, obtains a second image by decoding the first image data, obtains a third image by performing AI upscaling on the second image through a neural network for upscaling, and provides the second image data obtained through encoding of the third image to the display device.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to the field of image processing. More specifically, the present disclosure relates to an apparatus and method for processing and displaying images based on AI. Background Technology

[0002] The video is encoded by a codec that follows a specified data compression standard, such as the MPEG (Moving Picture Expert Group) standard, and then stored in the form of a bitstream on a recording medium or transmitted through a communication channel.

[0003] With the development and widespread adoption of hardware capable of playing and storing high-resolution / high-quality video, the need for codecs capable of effectively encoding and decoding high-resolution / high-quality video is increasing. The problem to be solved

[0004] An AI-based image providing device and a method based thereon according to one embodiment, and an AI-based display device and a method based thereon, have as technical challenges the encoding and decoding of images based on AI to prevent storage capacity saturation and achieve a low bit rate.

[0005] In addition, the AI-based image providing device and method according to one embodiment, and the AI-based display device and method according to the same have a technical objective of maintaining compatibility between devices by having the image providing device process an image based on AI and provide it to a display device for a display device that does not support AI functions. means of solving the problem

[0006] An electronic device that provides an image using artificial intelligence (AI) according to one embodiment includes a processor that executes one or more instructions stored in the electronic device, and the processor may obtain a first image by performing AI downscaling on an original image through a neural network for downscaling, obtain first image data by encoding the first image, and if the display device does not support an AI upscaling function, obtain a second image by decoding the first image data, obtain a third image by performing AI upscaling on the second image through a neural network for upscaling, and provide the second image data obtained through encoding the third image to the display device.

[0007] The electronic device may further include a camera for acquiring the original image.

[0008] The processor can provide the first image data to the display device if the display device supports an AI upscale function.

[0009] The processor can obtain the first image by AI downscaling the original image through a neural network for downscaling to which predetermined first neural network setting information is applied, select a second neural network setting information corresponding to the supported resolution of the display device among a plurality of second neural network setting information, and obtain the third image by AI upscaling the second image through a neural network for upscaling to which the selected second neural network setting information is applied.

[0010] The first neural network setting information that is predetermined above, and the second neural network setting information corresponding to the first neural network setting information among the plurality of second neural network setting information, are obtained through joint training of the downscale neural network and the upscale neural network, and the second neural network setting information other than the second neural network setting information corresponding to the first neural network setting information among the plurality of second neural network setting information can be obtained through individual training of the upscale neural network after the joint training.

[0011] The processor can obtain the first image by AI downscaling the original image through a neural network for downscaling to which predetermined first neural network setting information is applied, and obtain the third image by AI upscaling the second image n times (n is a natural number) through the neural network for upscaling to which predetermined second neural network setting information is applied, taking into account the scaling ratio of the neural network for upscaling to which predetermined second neural network setting information is applied and the supported resolution of the display device.

[0012] The processor can obtain the first image by selecting a first neural network setting information for AI downscaling among a plurality of first neural network setting information, and performing AI downscaling on the original image through a downscaling neural network to which the selected first neural network setting information is applied, and then obtain the third image by selecting a second neural network setting information corresponding to the selected first neural network setting information among a plurality of second neural network setting information, and performing AI upscaling on the second image n times (n is a natural number) through the upscaling neural network by considering the scaling ratio of the upscaling neural network to which the selected second neural network setting information is applied and the supported resolution of the display device.

[0013] As the original frames constituting the original image are sequentially AI downscaled through the downscaled neural network, the first frames constituting the first image are obtained, and the processor can sequentially delete the original frames for which AI downscaled is completed.

[0014] As the second frames constituting the second image are sequentially AI upscaled through the neural network for upscaling, the third frames constituting the third image are obtained, and when the third frames constituting the group of pictures (GOP) are obtained through the AI ​​upscaling, the processor can encode the third frames constituting the GOP and delete the encoded third frames constituting the GOP.

[0015] The processor receives performance information from the display device and can determine from the performance information whether the AI ​​upscale function is supported and the supported resolution of the display device.

[0016] An electronic device for displaying an image using artificial intelligence (AI) according to one embodiment includes: a display; and a processor for executing one or more instructions stored in the electronic device. The processor transmits information that the electronic device supports an AI upscaling function to an image providing device, receives image data corresponding to a first encoding result of a first image and AI data related to AI downscaling from an original image to the first image from the image providing device, decodes the image data to obtain a second image, selects neural network setting information for AI upscaling among a plurality of neural network setting information based on the AI ​​data, performs AI upscaling on the second image through an upscaling NN to which the selected neural network setting information is applied to obtain a third image, and provides the third image to the display.

[0017] A method for providing an image by an electronic device using artificial intelligence (AI) according to one embodiment may include: a step of obtaining a first image by performing AI downscaling on an original image through a neural network for downscaling; a step of obtaining first image data by encoding the first image; a step of obtaining a second image by decoding the first image data if the display device does not support an AI upscaling function; a step of obtaining a third image by AI upscaling the second image through a neural network for upscaling; and a step of providing the second image data obtained through encoding the third image to the display device.

[0018] A method for displaying an image using an electronic device utilizing artificial intelligence (AI) according to one embodiment may include: transmitting information to an image providing device that the electronic device supports an AI upscaling function; receiving from the image providing device image data corresponding to a first encoding result of a first image and AI data related to AI downscaling from an original image to the first image; decoding the image data to obtain a second image; selecting neural network setting information for AI upscaling among a plurality of neural network setting information based on the AI ​​data; performing AI upscaling on the second image through an upscaling NN to which the selected neural network setting information is applied to obtain a third image; and providing the third image to the display. Effects of the invention

[0019] An AI-based image providing device and method according to one embodiment, and an AI-based display device and method according to the same, can process images at a low bit rate through AI-based image encoding and decoding, and can prevent the storage capacity of the device from becoming saturated.

[0020] In addition, an AI-based image providing device and a method based thereon according to one embodiment, and an AI-based display device and a method based thereon, can maintain compatibility between devices by having the image providing device process an image based on AI and provide it to a display device for a display device that does not support AI functions. Brief explanation of the drawing

[0021] A brief description of each drawing is provided to help to better understand the drawings cited in this specification. FIG. 1 is a diagram illustrating an AI encoding process and an AI decoding process according to one embodiment. FIG. 2 is a block diagram showing the configuration of an AI decoding device according to one embodiment. FIG. 3 is an exemplary diagram showing a second neural network for AI upscaling of a second image. Figure 4 is a diagram illustrating a convolution operation by a convolution layer. Figure 5 is an exemplary diagram showing the mapping relationship between various image-related information and various neural network configuration information. FIG. 6 is a drawing illustrating a second image composed of multiple frames. FIG. 7 is a block diagram showing the configuration of an AI encoding device according to one embodiment. Figure 8 is an exemplary diagram showing a first neural network for AI downscaling of an original image. FIG. 9 is a diagram showing the structure of AI encoding data according to one embodiment. FIG. 10 is a diagram showing the structure of AI encoding data according to another embodiment. Figure 11 is a diagram illustrating a method for training a first neural network and a second neural network. FIG. 12 is a diagram illustrating the training process of the first neural network and the second neural network by a training device. FIG. 13 is a block diagram illustrating the configuration of an image providing device according to one embodiment. Figure 14 is a diagram illustrating the processing process of the original frames of the original image. Figure 15 is a flowchart illustrating the process of AI downscaling the original frames of the original video. Figure 16 is a flowchart illustrating the process of AI upscaling the second frames of the second image. FIG. 17 is a diagram illustrating first neural network configuration information applicable to a first neural network for AI downscaling, and second neural network configuration information applicable to a second neural network for AI upscaling. FIG. 18 is a diagram illustrating a plurality of first neural network configuration information applicable to a first neural network for AI downscaling, and a plurality of second neural network configuration information applicable to a second neural network for AI upscaling. FIG. 19 is a diagram illustrating first neural network configuration information applicable to a first neural network for AI downscaling, and a plurality of second neural network configuration information applicable to a second neural network for AI upscaling. FIG. 20 is a diagram illustrating a plurality of first neural network configuration information applicable to a first neural network for AI downscaling, and a plurality of second neural network configuration information applicable to a second neural network for AI upscaling. FIG. 21 is an exemplary drawing showing a second neural network according to one embodiment. FIG. 22 is a diagram illustrating a method for individually training a second neural network after the training in conjunction with the first neural network has been completed. FIG. 23 is a diagram illustrating the individual training process of a second neural network by a training device. FIG. 24 is a flowchart for explaining a method of providing an image according to one embodiment. Specific details for implementing the invention

[0022] The present disclosure is capable of various modifications and may have various embodiments, and specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the embodiments of the present disclosure, and it should be understood that the present disclosure includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the various embodiments.

[0023] In describing the embodiments, if it is determined that a detailed description of related prior art could unnecessarily obscure the gist of the present disclosure, such detailed description is omitted. Additionally, numbers used in the description of the specification (e.g., first, second, etc.) are merely identifiers to distinguish one component from another.

[0024] In addition, when a component is described in this specification as being "connected" or "connected" to another component, it should be understood that the component may be directly connected to or directly connected to the other component, but unless otherwise specifically stated, it may also be connected or connected through another component in between.

[0025] In addition, components expressed as '~part (unit)', 'module', etc. in this specification may consist of two or more components combined into a single component, or a single component may be divided into two or more components according to more detailed functions. Furthermore, each component described below may additionally perform some or all of the functions performed by other components in addition to the primary function it is responsible for, and it goes without saying that some of the primary functions performed by each component may be exclusively performed by other components.

[0026] Additionally, in this specification, 'image' or 'picture' may refer to a still image, a video composed of a plurality of consecutive still images (or frames), or a video.

[0027] Additionally, in this specification, 'neural network (NN)' is a representative example of an artificial neural network model that mimics brain neurons and is not limited to an artificial neural network model using a specific algorithm. The neural network may also be referred to as a deep neural network (DNN).

[0028] Additionally, in this specification, 'parameter' refers to a value used in the computation process of each layer constituting a neural network, and may include, for example, weights used when applying input values ​​to a predetermined calculation formula. Parameters may be expressed in the form of a matrix. Parameters are values ​​set as a result of training and may be updated through separate training data as needed.

[0029] In addition, in this specification, 'first neural network' refers to a neural network used for AI downscaling of an image, and 'second neural network' refers to a neural network used for AI upscaling of an image.

[0030] Additionally, in this specification, 'neural network (NN) configuration information' refers to information related to elements constituting a neural network and includes the aforementioned parameters. A neural network may be configured using the neural network configuration information. The neural network configuration information may also be referenced as deep neural network (DNN) configuration information.

[0031] Additionally, in this specification, 'original image' refers to an image that is the subject of AI encoding, and 'first image' refers to an image obtained as a result of AI downscaling of the original image during the AI ​​encoding process. Additionally, 'second image' refers to an image obtained through the first decoding during the AI ​​decoding process, and 'third image' refers to an image obtained by AI upscaling the second image during the AI ​​decoding process.

[0032] Additionally, in this specification, 'AI downscale' refers to a process of reducing the resolution of an image based on AI, and 'first encoding' refers to encoding processing by a frequency conversion-based image compression method. Additionally, 'first decoding' refers to decoding processing by a frequency conversion-based image restoration method, and 'AI upscale' refers to a process of increasing the resolution of an image based on AI. In this specification, 'first encoding' and 'first decoding' may be referred to as 'encoding' and 'decoding', respectively.

[0033] FIG. 1 is a diagram illustrating an artificial intelligence (AI) encoding process and an AI decoding process according to one embodiment.

[0034] As mentioned above, as image resolution increases rapidly, the amount of information processing required for encoding and decoding increases; accordingly, measures are needed to improve the efficiency of image encoding and decoding.

[0035] As illustrated in FIG. 1, according to one embodiment of the present disclosure, a first image (115) is obtained by AI downscaling (110) an original image (105) with high resolution. Then, since a first encoding (120) and a first decoding (130) are performed on the first image (115) with relatively low resolution, the bitrate can be significantly reduced compared to the case where the first encoding (120) and the first decoding (130) are performed on the original image (105).

[0036] Referring to FIG. 1, the description is as follows: In one embodiment, during the AI ​​encoding process, an original image (105) is AI downscaled (110) to obtain a first image (115), and the first image (115) is first encoded (120). In the AI ​​decoding process, AI encoded data including AI data and image data obtained as a result of AI encoding is received, a second image (135) is obtained through first decoding (130), and a third image (145) is obtained by AI upscaled (140) the second image (135).

[0037] Looking at the AI ​​encoding process in more detail, when an original video (105) is input, the original video (105) is AI downscaled (110) to obtain a first video (115) of a predetermined resolution and / or predetermined quality. At this time, the AI ​​downscaled (110) is performed based on AI, and the AI ​​for the AI ​​downscaled (110) must be jointly trained with the AI ​​for the AI ​​upscaled (140) of the second video (135). This is because if the AI ​​for the AI ​​downscaled (110) and the AI ​​for the AI ​​upscaled (140) are trained separately, the difference between the original video (105) that is the subject of AI encoding and the third video (145) restored through AI decoding becomes large.

[0038] In an embodiment of the present disclosure, AI data may be used to maintain this relationship during the AI ​​encoding process and the AI ​​decoding process. Accordingly, the AI ​​data obtained through the AI ​​encoding process must include information indicating an upscale target, and during the AI ​​decoding process, the second image (135) must be AI upscaled (140) according to the upscale target identified based on the AI ​​data.

[0039] AI for AI downscaling (110) and AI for AI upscaling (140) can be implemented as neural networks. As described below with reference to FIG. 11, since the first neural network and the second neural network are trained in conjunction by sharing loss information under a predetermined target, the AI ​​encoding device provides the target information used when the first neural network and the second neural network are trained in conjunction to the AI ​​decoder, and the AI ​​decoder can AI upscale (140) the second image (135) to the target quality and / or resolution based on the provided target information.

[0040] To explain in detail the first encoding (120) and the first decoding (130) illustrated in FIG. 1, the first image (115) that has been AI downscaled (110) from the original image (105) can have its amount of information reduced through the first encoding (120). The first encoding (120) may include a process of predicting the first image (115) to generate prediction data, a process of generating residual data corresponding to the difference between the first image (115) and the prediction data, a process of transforming the residual data, which is a spatial domain component, into a frequency domain component, a process of quantizing the residual data transformed into a frequency domain component, and a process of entropy encoding the quantized residual data. This first encoding process (120) can be implemented through one of the video compression methods using frequency conversion, such as MPEG-2, H.264 AVC (Advanced Video Coding), MPEG-4, HEVC (High Efficiency Video Coding), VC-1, VP8, VP9, ​​and AV1 (AOMedia Video 1).

[0041] A second image (135) corresponding to a first image (115) can be restored through a first decoding (130) of image data. The first decoding (130) may include a process of generating quantized residual data by entropy decoding image data, a process of inversely quantizing the quantized residual data, a process of converting residual data of frequency domain components into spatial domain components, a process of generating prediction data, and a process of restoring the second image (135) using the prediction data and residual data. Such a first decoding (130) process can be implemented through an image restoration method corresponding to one of the image compression methods that utilize the frequency conversion used in the first encoding (120) process.

[0042] The AI ​​encoded data obtained through the AI ​​encoding process may include image data obtained as a result of the first encoding (120) of the first image (115) and AI data related to the AI ​​downscale (110) of the original image (105). The image data may be used in the first decoding (130) process, and the AI ​​data may be used in the AI ​​upscale (140) process.

[0043] The image data may be transmitted in the form of a bitstream. The image data may include data obtained based on pixel values ​​within the first image (115), for example, residual data which is the difference between the first image (115) and the prediction data of the first image (115). Additionally, the image data includes information used in the first encoding (120) process of the first image (115). For example, the image data may include prediction mode information, motion information, and quantization parameter-related information used in the first encoding (120) used to first encode (120) the first image (115). The image data may be generated according to the rules, for example, the syntax, of the image compression method used in the first encoding (120) process among image compression methods using frequency conversion, such as MPEG-2, H.264 AVC, MPEG-4, HEVC, VC-1, VP8, VP9, ​​and AV1.

[0044] AI data is used for AI upscaling (140) based on the second neural network. As described above, since the first neural network and the second neural network are trained in conjunction, the AI ​​data contains information that enables accurate AI upscaling (140) of the second image (135) through the second neural network. In the AI ​​decoding process, the second image (135) can be AI upscaled (140) to a target resolution and / or quality based on the AI ​​data.

[0045] AI data can be transmitted together with video data in the form of a bitstream. Depending on the embodiment, AI data may be transmitted separately from video data in the form of frames or packets. Alternatively, depending on the embodiment, AI data may be transmitted embedded within the video data. Video data and AI data may be transmitted through the same network or through different networks.

[0046] FIG. 2 is a block diagram showing the configuration of an AI decoding device (200) according to one embodiment.

[0047] Referring to FIG. 2, an AI decoding device (200) according to one embodiment includes a receiving unit (210) and an AI decoding unit (230). The AI ​​decoding unit (230) may include a parsing unit (232), a first decoding unit (234), an AI upscaling unit (236), and an AI setting unit (238).

[0048] In FIG. 2, the receiver (210) and the AI ​​decoder (230) are depicted as separate devices, but the receiver (210) and the AI ​​decoder (230) can be implemented through a single processor. In this case, the receiver (210) and the AI ​​decoder (230) may be implemented as a dedicated processor, or they may be implemented through a combination of a general-purpose processor such as an application processor (AP), a central processing unit (CPU), or a graphic processing unit (GPU) and software. Additionally, in the case of a dedicated processor, it may include memory for implementing the embodiment of the present disclosure or a memory processing unit for using external memory.

[0049] The receiver (210) and the AI ​​decoding unit (230) may be composed of multiple processors. In this case, they may be implemented by a combination of dedicated processors, or by a combination of multiple general-purpose processors such as AP, CPU, and GPU, and S / W. In one embodiment, the receiver (210) is implemented as a first processor, the first decoding unit (234) is implemented as a second processor different from the first processor, and the parsing unit (232), the AI ​​upscaling unit (236), and the AI ​​setting unit (238) may be implemented as a third processor different from the first processor and the second processor.

[0050] The receiver (210) receives AI encoded data obtained as a result of AI encoding. For example, the AI ​​encoded data may be a video file having a file format such as mp4, mov, etc.

[0051] The receiver (210) can receive AI encoded data transmitted through the network. The receiver (210) outputs the AI ​​encoded data to the AI ​​decoder (230).

[0052] In one embodiment, the AI ​​encoded data may be obtained from a data storage medium including a magnetic medium such as a hard disk, a floppy disk and a magnetic tape, an optical recording medium such as a CD-ROM and a DVD, a magneto-optical medium such as a floptical disk, etc.

[0053] The parsing unit (232) parses the AI ​​encoded data and transmits the image data generated as a first encoding result of the first image (115) to the first decoding unit (234), and transmits the AI ​​data to the AI ​​setting unit (238).

[0054] In one embodiment, the parsing unit (232) can parse video data and AI data that are separated and included within the AI ​​encoded data. The parsing unit (232) can distinguish between the AI ​​data and video data included within the AI ​​encoded data by reading the header within the AI ​​encoded data. In one example, the AI ​​data may be included in the Vendor Specific InfoFrame (VSIF) within the HDMI stream. The structure of the AI ​​encoded data containing the separated AI data and video data will be described later with reference to FIG. 9.

[0055] In another embodiment, the parsing unit (232) parses image data from the AI ​​encoded data, extracts AI data from the image data, transmits the AI ​​data to the AI ​​setting unit (238), and transmits the remaining image data to the first decoding unit (234). That is, the AI ​​data may be included in the image data; for example, the AI ​​data may be included in the SEI (Supplemental enhancement information), which is an additional information area of ​​the bitstream corresponding to the image data. The structure of the AI ​​encoded data including the image data containing the AI ​​data will be described later with reference to FIG. 10.

[0056] In another embodiment, the parsing unit (232) may divide a bitstream corresponding to image data into a bitstream to be processed by the first decoding unit (234) and a bitstream corresponding to AI data, and output each divided bitstream to the first decoding unit (234) and the AI ​​setting unit (238).

[0057] The parsing unit (232) may determine that the video data included in the AI ​​encoded data is video data obtained through a predetermined codec (e.g., MPEG-2, H.264, MPEG-4, HEVC, VC-1, VP8, VP9, ​​or AV1). In this case, the information may be transmitted to the first decoding unit (234) so ​​that the video data can be processed by the confirmed codec.

[0058] The first decoding unit (234) restores a second image (135) corresponding to the first image (115) based on image data received from the parsing unit (232). The second image (135) obtained by the first decoding unit (234) is provided to the AI ​​upscaling unit (236).

[0059] According to an embodiment, first decoding-related information, such as prediction mode information, motion information, and quantization parameter information, may be provided from the first decoding unit (234) to the AI ​​setting unit (238). The first decoding-related information may be used to obtain neural network setting information.

[0060] The AI ​​data provided to the AI ​​setting unit (238) includes information that enables AI upscaling of the second image (135). At this time, the upscaling target of the second image (135) must correspond to the downscaling target of the first neural network. Accordingly, the AI ​​data must include information that can identify the downscaling target of the first neural network.

[0061] Specifically, examples of the information included in the AI ​​data include information on the difference between the resolution of the original image (105) and the resolution of the first image (115), and information related to the first image (115).

[0062] The difference information can be expressed as information regarding the degree of resolution conversion of the first image (115) relative to the original image (105) (e.g., resolution conversion rate information). Also, since the resolution of the first image (115) can be determined through the resolution of the restored second image (135) and the degree of resolution conversion can be verified through this, the difference information may be expressed only by the resolution information of the original image (105). Here, the resolution information may be expressed as horizontal and vertical screen sizes, or as a ratio (16:9, 4:3, etc.) and a size of one axis. Additionally, if there is pre-set resolution information, it may be expressed in the form of an index or a flag.

[0063] Information related to the first image (115) may include information on at least one of the resolution of the first image (115), the bitrate of the image data obtained as a result of the first encoding of the first image (115), and the type of codec used during the first encoding of the first image (115).

[0064] The AI ​​setting unit (238) can determine an upscale target for the second image (135) based on at least one of the difference information included in the AI ​​data and information related to the first image (115). The upscale target may indicate, for example, to what resolution the second image (135) should be upscaled.

[0065] When the AI ​​upscaler (236) determines the upscaler target, it AI upscales the second image (135) through the second neural network to obtain the third image (145) corresponding to the upscaler target.

[0066] Before explaining how the AI ​​setting unit (238) determines an upscale target based on AI data, the AI ​​upscale process through the second neural network is explained with reference to FIG. 3 and FIG. 4.

[0067] FIG. 3 is an exemplary drawing showing a second neural network (300) for AI upscaling of a second image (135), and FIG. 4 shows a convolution operation in the first convolution layer (310) shown in FIG. 3.

[0068] As illustrated in FIG. 3, the second image (135) is input into the first convolution layer (310). The 3x3x4 shown in FIG. 3 illustrates convolution processing on one input image using four filter kernels of size 3 x 3. As a result of the convolution processing, four feature maps are generated by the four filter kernels. Each feature map represents unique characteristics of the second image (135). For example, each feature map may represent vertical characteristics, horizontal characteristics, or edge characteristics of the second image (135).

[0069] Referring to FIG. 4, the convolution operation in the first convolution layer (310) will be described in detail.

[0070] A single feature map (450) can be generated through multiplication and addition operations between the parameters of a filter kernel (430) having a size of 3 x 3 used in the first convolution layer (310) and the corresponding pixel values ​​in the second image (135). Since four filter kernels are used in the first convolution layer (310), four feature maps can be generated through a convolution operation process using four filter kernels.

[0071] In FIG. 4, I1 to I49 displayed in the second image (135) represent pixels of the second image (135), and F1 to F9 displayed in the filter kernel (430) represent parameters of the filter kernel (430). Additionally, M1 to M9 displayed in the feature map (450) represent samples of the feature map (450).

[0072] FIG. 4 illustrates that the second image (135) contains 49 pixels, but this is merely an example, and if the second image (135) has a resolution of 4K, for example, it may contain 3840 x 2160 pixels.

[0073] In the convolution operation, each of the pixel values ​​of I1, I2, I3, I8, I9, I10, I15, I16, and I17 of the second image (135) and each of F1, F2, F3, F4, F5, F6, F7, F8, and F9 of the filter kernel (430) are multiplied, and the result of the multiplication operation combined (e.g., addition operation) can be assigned as the value of M1 of the feature map (450). If the stride of the convolution operation is 2, the multiplication operation is performed on each of the pixel values ​​of I3, I4, I5, I10, I11, I12, I17, I18, and I19 of the second image (135) and each of F1, F2, F3, F4, F5, F6, F7, F8, and F9 of the filter kernel (430), and the combined result of the multiplication operation can be assigned as the value of M2 of the feature map (450).

[0074] A feature map (450) having a predetermined size can be obtained by performing a convolution operation between the pixel values ​​in the second image (135) and the parameters of the filter kernel (430) while the filter kernel (430) moves according to the stride until it reaches the last pixel of the second image (135).

[0075] According to the present disclosure, through the combined training of the first neural network and the second neural network, the values ​​of the parameters of the second neural network, for example, the parameters of the filter kernel used in the convolution layers of the second neural network (for example, F1, F2, F3, F4, F5, F6, F7, F8, and F9 of the filter kernel (430)), can be optimized. The AI ​​setting unit (238) can determine an upscale target corresponding to a downscale target of the first neural network based on AI data, and determine the parameters corresponding to the determined upscale target as the parameters of the filter kernel used in the convolution layers of the second neural network.

[0076] The convolution layers included in the first neural network and the second neural network can perform processing according to the convolution operation process described in relation to FIG. 4, but the convolution operation process described in FIG. 4 is merely one example and is not limited thereto.

[0077] Referring again to FIG. 3, the feature maps output from the first convolution layer (310) are input to the first activation layer (320).

[0078] The first activation layer (320) can impart non-linear characteristics to each feature map. The first activation layer (320) may include a sigmoid function, a Tanh function, a ReLU (Rectified Linear Unit) function, etc., but is not limited thereto.

[0079] Imposing non-linear characteristics in the first activation layer (320) means changing and outputting some sample values ​​of the feature map, which is the output of the first convolution layer (310). At this time, the change is performed by applying non-linear characteristics.

[0080] The first activation layer (320) determines whether to pass sample values ​​of feature maps output from the first convolution layer (310) to the second convolution layer (330). For example, some sample values ​​of the feature maps are activated by the first activation layer (320) and passed to the second convolution layer (330), while some sample values ​​are deactivated by the first activation layer (320) and not passed to the second convolution layer (330). The unique characteristics of the second image (135) represented by the feature maps are emphasized by the first activation layer (320).

[0081] The feature maps (325) output from the first activation layer (320) are input to the second convolution layer (330). One of the feature maps (325) illustrated in FIG. 3 is the result of processing the feature map (450) described in relation to FIG. 4 in the first activation layer (320).

[0082] The 3x3x4 shown in the second convolution layer (330) exemplifies convolution processing on input feature maps (325) using four filter kernels of size 3 x 3. The output of the second convolution layer (330) is input to the second activation layer (340). The second activation layer (340) can impart non-linear characteristics to the input data.

[0083] Feature maps (345) output from the second activation layer (340) are input to the third convolution layer (350). The 3x3x1 shown in the third convolution layer (350) illustrated in FIG. 3 exemplifies performing convolution processing to create one output image using one filter kernel of size 3 x 3. The third convolution layer (350) is a layer for outputting a final image and generates one output using one filter kernel. According to an example of the present disclosure, the third convolution layer (350) can output a third image (145) through convolution operations.

[0084] Neural network configuration information indicating the number of filter kernels, parameters of filter kernels, etc. of the first convolution layer (310), the second convolution layer (330), and the third convolution layer (350) of the second neural network (300) may be multiple as described below, and the multiple neural network configuration information must correspond to the multiple neural network configuration information of the first neural network. The correspondence relationship between the multiple neural network configuration information of the second neural network and the multiple neural network configuration information of the first neural network can be implemented through linked learning of the first neural network and the second neural network.

[0085] FIG. 3 illustrates that the second neural network (300) includes three convolution layers (310, 330, 350) and two activation layers (320, 340), but this is merely an example, and depending on the embodiment, the number of convolution layers and activation layers may vary. Additionally, depending on the embodiment, the second neural network (300) may be implemented through a recurrent neural network (RNN). In this case, it means changing the CNN structure of the second neural network (300) according to the example of the present disclosure to an RNN structure.

[0086] In one embodiment, the AI ​​upscaler (236) may include at least one Arithmetic Logic Unit (ALU) for the aforementioned convolution operation and the operation of the activation layer. The ALU may be implemented as a processor. For the convolution operation, the ALU may include a multiplier that performs a multiplication operation between the sample values ​​of the feature map output from the second image (135) or the previous layer and the sample values ​​of the filter kernel, and an adder that adds the result values ​​of the multiplication. Additionally, for the operation of the activation layer, the ALU may include a multiplier that multiplies the input sample value by a weight used in a predetermined sigmoid function, Tanh function, or ReLU function, and a comparator that compares the result of the multiplication with a predetermined value to determine whether to pass the input sample value to the next layer.

[0087] In the following, the AI ​​setting unit (238) determines an upscale target, and the AI ​​upscale unit (236) explains how to AI upscale the second image (135) according to the upscale target.

[0088] In one embodiment, the AI ​​setting unit (238) can store a plurality of neural network setting information that can be set in the second neural network.

[0089] Here, the neural network configuration information may include information on at least one of the number of convolution layers included in the second neural network, the number of filter kernels per convolution layer, and the parameters of each filter kernel.

[0090] Multiple neural network configuration information can correspond to various upscale targets, and a second neural network can operate based on neural network configuration information corresponding to a specific upscale target. Depending on the neural network configuration information, the second neural network may have different structures. For example, depending on one neural network configuration information, the second neural network may include three convolution layers, and depending on another neural network configuration information, the second neural network may include four convolution layers.

[0091] In one embodiment, the neural network configuration information may include only the parameters of the filter kernel used in the second neural network. In this case, the structure of the second neural network is not changed, but only the parameters of the internal filter kernel may change according to the neural network configuration information.

[0092] The AI ​​setting unit (238) can obtain neural network setting information for AI upscaling of the second image (135) among a plurality of neural network setting information. Each of the plurality of neural network setting information used here is information for obtaining a third image (145) of a predetermined resolution and / or a predetermined quality, and is trained in conjunction with the first neural network.

[0093] For example, one of the neural network setting information may include information for obtaining a third image (145) with a resolution twice as large as the resolution of the second image (135), for example, a third image (145) of 4K (4096*2160) which is twice as large as the second image (135) of 2K (2048*1080), and another neural network setting information may include information for obtaining a third image (145) with a resolution four times larger than the resolution of the second image (135), for example, a third image (145) of 8K (8192*4320) which is four times larger than the second image (135) of 2K (2048*1080).

[0094] Each of the plurality of neural network setting information is created in conjunction with the neural network setting information of the first neural network of the AI ​​encoding device (700), and the AI ​​setting unit (238) obtains one of the neural network setting information among the plurality of neural network setting information according to an enlargement ratio corresponding to the reduction ratio of the neural network setting information of the first neural network. To do this, the AI ​​setting unit (238) must verify the information of the first neural network. In order for the AI ​​setting unit (238) to verify the information of the first neural network, the AI ​​decoding device (200) according to one embodiment receives AI data including the information of the first neural network from the AI ​​encoding device (700).

[0095] In other words, the AI ​​setting unit (238) can use the information received from the AI ​​encoding device (700) to identify the information targeted by the neural network setting information of the first neural network used to acquire the first image (115), and acquire the neural network setting information of the second neural network trained in conjunction with it.

[0096] When neural network setting information for AI upscaling of the second image (135) is obtained among multiple neural network setting information, the neural network setting information is transmitted to the AI ​​upscaling unit (236), and input data can be processed based on the second neural network that operates according to the neural network setting information.

[0097] For example, when the AI ​​upscaler (236) obtains any one neural network configuration information, for each of the first convolution layer (310), the second convolution layer (330), and the third convolution layer (350) of the second neural network (300) shown in FIG. 3, the number of filter kernels included in each layer and the parameters of the filter kernels are set to the values ​​included in the obtained neural network configuration information.

[0098] Specifically, the AI ​​upscaler (236) can replace the parameters of the 3 x 3 filter kernel (430) shown in FIG. 4 with {1, 1, 1, 1, 1, 1, 1, 1} and, if there is a change in the neural network configuration information, with the parameters of the filter kernel (430) being replaced with the parameters included in the changed neural network configuration information, namely {2, 2, 2, 2, 2, 2, 2, 2}.

[0099] The AI ​​setting unit (238) can obtain neural network setting information for upscaling the second image (135) among a plurality of neural network setting information based on information included in the AI ​​data, and the AI ​​data used to obtain the neural network setting information is described in detail.

[0100] In one embodiment, the AI ​​setting unit (238) can obtain neural network setting information for upscaling a second image (135) among a plurality of neural network setting information based on difference information included in the AI ​​data. For example, if it is confirmed based on the difference information that the resolution of the original image (105) (e.g., 4K (4096*2160)) is twice as large as the resolution of the first image (115) (e.g., 2K (2048*1080)), the AI ​​setting unit (238) can obtain neural network setting information that can increase the resolution of the second image (135) by two times.

[0101] In another embodiment, the AI ​​setting unit (238) may obtain neural network setting information for AI upscaling a second image (135) among a plurality of neural network setting information based on information related to a first image (115) included in the AI ​​data. The AI ​​setting unit (238) may predetermine a mapping relationship between the image-related information and the neural network setting information and obtain neural network setting information mapped to the information related to the first image (115).

[0102] Figure 5 is an exemplary diagram showing the mapping relationship between various image-related information and various neural network configuration information.

[0103] Through the embodiment according to FIG. 5, it can be seen that the AI ​​encoding / AI decoding process according to one embodiment of the present disclosure does not consider only changes in resolution. As shown in FIG. 5, the selection of neural network configuration information can be made by considering resolutions such as SD, HD, and Full HD, bitrates such as 10 Mbps, 15 Mbps, and 20 Mbps, and codec information such as AV1, H.264, and HEVC, either individually or all together. For such consideration, training that considers each of these elements must be performed in conjunction with the encoding and decoding processes during the AI ​​training process (see FIG. 11).

[0104] Accordingly, when a plurality of neural network setting information corresponding to image-related information, such as codec type and image resolution, is provided through training as illustrated in FIG. 5, the AI ​​setting unit (238) can select neural network setting information for AI upscaling of the second image (135) based on the first image (115) related information received during the AI ​​decoding process.

[0105] That is, the AI ​​setting unit (238) matches the image-related information shown on the left side of the table shown in FIG. 5 with the neural network setting information on the right side of the table, thereby enabling the use of neural network setting information according to the image-related information.

[0106] As illustrated in FIG. 5, if it is confirmed from the information related to the first image (115) that the resolution of the first image (115) is SD, the bit rate of the image data obtained as a result of the first encoding of the first image (115) is 10 Mbps, and the first image (115) is first encoded with an AV1 codec, the AI ​​setting unit (238) can obtain 'A' neural network setting information among a plurality of neural network setting information.

[0107] Additionally, if it is confirmed from the information related to the first image (115) that the resolution of the first image (115) is HD, the bit rate of the image data obtained as a result of the first encoding is 15 Mbps, and the first image (115) was first encoded with an H.264 codec, the AI ​​setting unit (238) can obtain the 'B' neural network setting information among the plurality of neural network setting information.

[0108] Additionally, if it is confirmed from the information related to the first image (115) that the resolution of the first image (115) is Full HD, the bitrate of the image data obtained as a result of the first encoding of the first image (115) is 20 Mbps, and the first image (115) is first encoded with an HEVC codec, the AI ​​setting unit (238) can obtain 'C' neural network setting information among the plurality of neural network setting information, and if it is confirmed that the resolution of the first image (115) is Full HD, the bitrate of the image data obtained as a result of the first encoding of the first image (115) is 15 Mbps, and the first image (115) is first encoded with an HEVC codec, the AI ​​setting unit (238) can obtain 'D' neural network setting information among the plurality of neural network setting information. Depending on whether the bit rate of the image data obtained as a result of the first encoding of the first image (115) is 20 Mbps or 15 Mbps, either the 'C' neural network setting information or the 'D' neural network setting information is selected. When the first image (115) of the same resolution is first encoded with the same codec, the fact that the bit rates of the image data are different means that the quality of the restored image is different. Accordingly, the first neural network and the second neural network can be trained in conjunction based on a predetermined quality, and accordingly, the AI ​​setting unit (238) can obtain neural network setting information according to the bit rate of the image data representing the quality of the second image (135).

[0109] In another embodiment, the AI ​​setting unit (238) may obtain neural network setting information for AI upscaling the second image (135) among a plurality of neural network setting information by considering all information provided from the first decoding unit (234) (prediction mode information, motion information, quantization parameter information, etc.) and information related to the first image (115) included in the AI ​​data. For example, the AI ​​setting unit (238) may receive quantization parameter information used in the first encoding process of the first image (115) from the first decoding unit (234), check the bit rate of the image data obtained as a result of encoding the first image (115) from the AI ​​data, and obtain neural network setting information corresponding to the quantization parameter and bit rate. Even with the same bit rate, there may be differences in the quality of the restored image depending on the complexity of the image, and the bit rate is a value that represents the entire first image (115) being encoded, so the quality of each frame within the first image (115) may differ. Therefore, by considering together the prediction mode information, motion information, and / or quantization parameters that can be obtained from each frame from the first decoding unit (234), neural network setting information that is more suitable for the second image (135) can be obtained compared to using only AI data.

[0110] Additionally, according to an embodiment, the AI ​​data may include an identifier of mutually agreed-upon neural network configuration information. The identifier of the neural network configuration information is information for distinguishing pairs of neural network configuration information linked between the first neural network and the second neural network so that the second image (135) can be AI upscaled as an upscale target corresponding to the downscale target of the first neural network. The AI ​​configuration unit (238) obtains the identifier of the neural network configuration information included in the AI ​​data, obtains neural network configuration information corresponding to the identifier of the neural network configuration information, and the AI ​​upscale unit (236) can AI upscale the second image (135) using the corresponding neural network configuration information. For example, an identifier pointing to each of a plurality of neural network configuration information that can be configured in the first neural network and an identifier pointing to each of a plurality of neural network configuration information that can be configured in the second neural network may be pre-specified. In this case, the same identifier may be assigned to pairs of neural network setting information that can be set for each of the first neural network and the second neural network. The AI ​​data may include an identifier of neural network setting information set in the first neural network for AI downscaling of the original image (105). The AI ​​setting unit (238) that receives the AI ​​data obtains the neural network setting information pointed to by the identifier included in the AI ​​data among a plurality of neural network setting information, and the AI ​​upscaling unit (236) can AI upscale the second image (135) using the corresponding neural network setting information.

[0111] Additionally, depending on the embodiment, the AI ​​data may include neural network setting information itself. The AI ​​setting unit (238) obtains the neural network setting information included in the AI ​​data, and the AI ​​upscale unit (236) may use the neural network setting information to AI upscale the second image (135).

[0112] According to an embodiment, if the information constituting the neural network setting information (e.g., the number of convolution layers, the number of filter kernels per convolution layer, the parameters of each filter kernel, etc.) is stored in the form of a lookup table, the AI ​​setting unit (238) obtains the neural network setting information by combining a selected portion of the lookup table values ​​based on the information included in the AI ​​data, and the AI ​​upscale unit (236) may AI upscale the second image (135) using the neural network setting information.

[0113] According to an embodiment, when the structure of the neural network corresponding to the upscale target is determined, the AI ​​setting unit (238) may obtain neural network setting information corresponding to the determined neural network structure, for example, parameters of a filter kernel.

[0114] As described above, the AI ​​setting unit (238) obtains neural network setting information of the second neural network through AI data containing information related to the first neural network, and the AI ​​upscale unit (236) AI upscales the second image (135) through the second neural network set with the neural network setting information, which can reduce memory usage and computational amount compared to upscale by directly analyzing the features of the second image (135).

[0115] In one embodiment, when the second image (135) is composed of a plurality of frames, the AI ​​setting unit (238) may independently obtain neural network setting information for each of a predetermined number of frames, or may obtain common neural network setting information for all frames.

[0116] FIG. 6 is a drawing illustrating a second image (135) composed of multiple frames.

[0117] As shown in FIG. 6, the second image (135) may consist of frames corresponding to t0 to tn.

[0118] In one example, the AI ​​setting unit (238) obtains neural network setting information of the second neural network through AI data, and the AI ​​upscaling unit (236) can AI upscale frames corresponding to t0 to tn based on the neural network setting information. That is, frames corresponding to t0 to tn can be AI upscaled based on common neural network setting information.

[0119] In another example, the AI ​​setting unit (238) may obtain 'A' neural network setting information from AI data for some of the frames corresponding to t0 to tn, for example, frames corresponding to t0 to ta, and obtain 'B' neural network setting information from AI data for frames corresponding to ta+1 to tb. Additionally, the AI ​​setting unit (238) may obtain 'C' neural network setting information from AI data for frames corresponding to tb+1 to tn. In other words, the AI ​​setting unit (238) independently obtains neural network setting information for each group containing a predetermined number of frames among a plurality of frames, and the AI ​​upscale unit (236) may AI upscale the frames included in each group using the independently obtained neural network setting information.

[0120] In another example, the AI ​​setting unit (238) may independently acquire neural network setting information for each frame constituting the second image (135). For example, if the second image (135) consists of three frames, the AI ​​setting unit (238) may acquire neural network setting information in relation to the first frame, acquire neural network setting information in relation to the second frame, and acquire neural network setting information in relation to the third frame. That is, neural network setting information may be independently acquired for each of the first frame, the second frame, and the third frame. Depending on the method in which neural network setting information is acquired based on the information provided from the aforementioned first decoding unit (234) (prediction mode information, motion information, quantization parameter information, etc.) and the information related to the first image (115) included in the AI ​​data, neural network setting information may be independently acquired for each frame constituting the second image (135). This is because the mode information, quantization parameter information, etc., can be independently determined for each frame constituting the second image (135).

[0121] In another example, the AI ​​data may include information indicating up to which frame the neural network configuration information obtained based on the AI ​​data is valid. For example, if the AI ​​data contains information that the neural network configuration information is valid up to frame ta, the AI ​​configuration unit (238) obtains the neural network configuration information based on the AI ​​data, and the AI ​​upscale unit (236) AI upscales frames t0 to ta with the corresponding neural network configuration information. And, if other AI data contains information that the neural network configuration information is valid up to frame tn, the AI ​​configuration unit (238) obtains the neural network configuration information based on the other AI data, and the AI ​​upscale unit (236) can AI upscale frames ta+1 to tn with the obtained neural network configuration information.

[0122] Hereinafter, with reference to FIG. 7, an AI encoding device (700) for AI encoding of an original image (105) will be described.

[0123] FIG. 7 is a block diagram showing the configuration of an AI encoding device (700) according to one embodiment.

[0124] Referring to FIG. 7, the AI ​​encoding device (700) may include an AI encoding unit (710) and a transmission unit (730). The AI ​​encoding unit (710) may include an AI downscale unit (712), a first encoding unit (714), a data processing unit (716), and an AI setting unit (718).

[0125] FIG. 7 illustrates the AI ​​encoding unit (710) and the transmission unit (730) as separate devices, but the AI ​​encoding unit (710) and the transmission unit (730) can be implemented through a single processor. In this case, it may be implemented with a dedicated processor, or it may be implemented through a combination of a general-purpose processor such as an AP, CPU, or GPU and S / W. In addition, in the case of a dedicated processor, it may include memory for implementing the embodiments of the present disclosure or a memory processing unit for using external memory.

[0126] Additionally, the AI ​​encoding unit (710) and the transmission unit (730) may be composed of multiple processors. In this case, they may be implemented by a combination of dedicated processors, or by a combination of multiple general-purpose processors such as AP, CPU, and GPU, and S / W. In one embodiment, the first encoding unit (714) is composed of a first processor, the AI ​​downscaling unit (712), the data processing unit (716), and the AI ​​setting unit (718) are implemented by a second processor different from the first processor, and the transmission unit (730) may be implemented by a third processor different from the first processor and the second processor.

[0127] The AI ​​encoding unit (710) performs AI downscaling of the original image (105) and first encoding of the first image (115), and transmits the AI ​​encoded data to the transmission unit (730). The transmission unit (730) transmits the AI ​​encoded data to the AI ​​decoding device (200).

[0128] The image data includes data obtained as a result of the first encoding of the first image (115). The image data may include data obtained based on pixel values ​​within the first image (115), for example, residual data which is the difference between the first image (115) and the prediction data of the first image (115). Additionally, the image data includes information used in the first encoding process of the first image (115). For example, the image data may include prediction mode information, motion information, and quantization parameter-related information used in the first encoding of the first image (115).

[0129] The AI ​​data includes information that enables the AI ​​upscale unit (236) to AI upscale the second image (135) to an upscale target corresponding to the downscale target of the first neural network.

[0130] In one example, the AI ​​data may include difference information between the original image (105) and the first image (115).

[0131] In one example, the AI ​​data may include information related to the first image (115). The information related to the first image (115) may include information about at least one of the resolution of the first image (115), the bitrate of the image data obtained as a result of the first encoding of the first image (115), and the type of codec used during the first encoding of the first image (115).

[0132] In one embodiment, the AI ​​data may include an identifier of mutually agreed neural network setting information so that the second image (135) can be AI upscaled to an upscale target corresponding to the downscale target of the first neural network.

[0133] In addition, in one embodiment, the AI ​​data may include neural network setting information that can be set in the second neural network.

[0134] The AI ​​downscaler (712) can obtain a first image (115) that has been AI downscaled from the original image (105) through a first neural network. The AI ​​downscaler (712) can AI downscale the original image (105) using neural network setting information provided by the AI ​​setting unit (718).

[0135] The AI ​​setting unit (718) can determine the downscale target of the original image (105) based on predetermined criteria.

[0136] In order to acquire a first image (115) corresponding to a downscale target, the AI ​​setting unit (718) can store multiple neural network setting information that can be set in the first neural network. The AI ​​setting unit (718) acquires neural network setting information corresponding to the downscale target among the multiple neural network setting information and provides the acquired neural network setting information to the AI ​​downscale unit (712).

[0137] Each of the above multiple neural network configuration information may be trained to obtain a first image (115) of a predetermined resolution and / or predetermined quality. For example, one of the neural network setting information may include information for obtaining a first image (115) with a resolution that is 1 / 2 times smaller than the resolution of the original image (105), for example, a first image (115) of 2K (2048*1080) that is 1 / 2 times smaller than the original image (105) of 4K (4096*2160), and another neural network setting information may include information for obtaining a first image (115) with a resolution that is 1 / 4 times smaller than the resolution of the original image (105), for example, a first image (115) of 2K (2048*1080) that is 1 / 4 times smaller than the original image (105) of 8K (8192*4320).

[0138] According to an implementation example, if the information constituting the neural network configuration information (e.g., the number of convolution layers, the number of filter kernels per convolution layer, the parameters of each filter kernel, etc.) is stored in the form of a lookup table, the AI ​​configuration unit (718) may obtain the neural network configuration information by combining a selected portion of the lookup table values ​​according to the downscale target, and provide the obtained neural network configuration information to the AI ​​downscale unit (712).

[0139] According to an embodiment, the AI ​​setting unit (718) may determine the structure of a neural network corresponding to a downscale target and obtain neural network setting information corresponding to the determined neural network structure, such as parameters of a filter kernel.

[0140] Multiple neural network configuration information for AI downscaling of the original image (105) can have optimized values ​​by training the first neural network and the second neural network in conjunction. Here, each neural network configuration information includes at least one of the number of convolution layers included in the first neural network, the number of filter kernels per convolution layer, and the parameters of each filter kernel.

[0141] The AI ​​downscaler (712) sets the first neural network with neural network setting information determined for AI downscale of the original image (105), and can obtain the first image (115) of a predetermined resolution and / or predetermined quality through the first neural network. When neural network setting information for AI downscale of the original image (105) is obtained among a plurality of neural network setting information, each layer within the first neural network can process input data based on the information included in the neural network setting information.

[0142] Below, the method by which the AI ​​setting unit (718) determines a downscale target is described. The downscale target may indicate, for example, how much the resolution of the first image (115) should be reduced from the original image (105).

[0143] The AI ​​setting unit (718) acquires one or more input information. In one embodiment, the input information may include at least one of the target resolution of the first image (115), the target bitrate of the image data, the bitrate type of the image data (e.g., variable bitrate type, constant bitrate type or average bitrate type, etc.), the color format to which AI downscale is applied (luminance component, chrominance component, red component, green component or blue component, etc.), the codec type for the first encoding of the first image (115), compression history information, the resolution of the original image (105), and the type of the original image (105).

[0144] One or more input information may include information that is pre-stored in the AI ​​encoding device (700) or received from the user.

[0145] The AI ​​setting unit (718) controls the operation of the AI ​​downscale unit (712) based on input information. In one embodiment, the AI ​​setting unit (718) determines a downscale target according to the input information and can provide neural network setting information corresponding to the determined downscale target to the AI ​​downscale unit (712).

[0146] In one embodiment, the AI ​​setting unit (718) may transmit at least a portion of the input information to the first encoding unit (714) so ​​that the first encoding unit (714) first encodes the first image (115) with a bit rate of a specific value, a bit rate of a specific type, and a specific codec.

[0147] In one embodiment, the AI ​​setting unit (718) can determine a downscale target based on at least one of a compression rate (e.g., a difference in resolution between the original image (105) and the first image (115), a target bitrate), a compression quality (e.g., a bitrate type), compression history information, and a type of the original image (105).

[0148] In one example, the AI ​​setting unit (718) can determine a downscale target based on a compression rate or compression quality that is pre-set or input by the user.

[0149] As another example, the AI ​​setting unit (718) may determine a downscale target using compression history information stored in the AI ​​encoding device (700). For example, according to the compression history information available to the AI ​​encoding device (700), the encoding quality or compression ratio preferred by the user may be determined, and the downscale target may be determined according to the encoding quality, etc. determined based on the compression history information. For example, the resolution, image quality, etc. of the first video (115) may be determined according to the encoding quality that has been used most frequently according to the compression history information.

[0150] As another example, the AI ​​setting unit (718) may determine a downscale target based on encoding quality that has been used more than a predetermined threshold value according to compression history information (e.g., the average quality of encoding quality that has been used more than a predetermined threshold value).

[0151] As another example, the AI ​​setting unit (718) may determine a downscale target based on the resolution and type (e.g., file format) of the original image (105).

[0152] In one embodiment, when the original image (105) is composed of a plurality of frames, the AI ​​setting unit (718) may independently acquire neural network setting information for a predetermined number of frames and provide the independently acquired neural network setting information to the AI ​​downscale unit (712).

[0153] In one example, the AI ​​setting unit (718) may divide the frames constituting the original image (105) into a predetermined number of groups and independently obtain neural network setting information for each group. For each group, identical or different neural network setting information may be obtained. The number of frames included in the groups may be the same or different for each group.

[0154] In another example, the AI ​​setting unit (718) can independently determine neural network setting information for each frame constituting the original image (105). For each frame, identical or different neural network setting information can be obtained.

[0155] Below, an exemplary structure of the first neural network (800) that forms the basis of AI downscaling is described.

[0156] FIG. 8 is an exemplary diagram showing a first neural network (800) for AI downscaling of an original image (105).

[0157] As illustrated in FIG. 8, the original image (105) is input to the first convolution layer (810). The first convolution layer (810) performs convolution processing on the original image (105) using 32 filter kernels of size 5 x 5. The 32 feature maps generated as a result of the convolution processing are input to the first activation layer (820).

[0158] The first activation layer (820) can impart non-linear characteristics to 32 feature maps.

[0159] The first activation layer (820) determines whether to transmit sample values ​​of feature maps output from the first convolution layer (810) to the second convolution layer (830). For example, some sample values ​​of the feature maps are activated by the first activation layer (820) and transmitted to the second convolution layer (830), while some sample values ​​are deactivated by the first activation layer (820) and not transmitted to the second convolution layer (830). The information represented by the feature maps output from the first convolution layer (810) is highlighted by the first activation layer (820).

[0160] The output (825) of the first activation layer (820) is input to the second convolution layer (830). The second convolution layer (830) performs convolution processing on the input data using 32 filter kernels of size 5 x 5. The 32 feature maps output as a result of the convolution processing are input to the second activation layer (840), and the second activation layer (840) can impart non-linear characteristics to the 32 feature maps.

[0161] The output (845) of the second activation layer (840) is input to the third convolution layer (850). The third convolution layer (850) performs convolution processing on the input data using one filter kernel of size 5 x 5. As a result of the convolution processing, one image can be output from the third convolution layer (850). The third convolution layer (850) obtains one output using one filter kernel as a layer for outputting the final image. According to an example of the present disclosure, the third convolution layer (850) can output the first image (115) through the result of the convolution operation.

[0162] There may be multiple neural network configuration information indicating the number of filter kernels, parameters of filter kernels, etc. of the first convolution layer (810), the second convolution layer (830), and the third convolution layer (850) of the first neural network (800), and the multiple neural network configuration information must correspond to multiple neural network configuration information of the second neural network. The correspondence relationship between the multiple neural network configuration information of the first neural network and the multiple neural network configuration information of the second neural network can be implemented through linked learning of the first neural network and the second neural network.

[0163] FIG. 8 illustrates that the first neural network (800) includes three convolution layers (810, 830, 850) and two activation layers (820, 840), but this is merely an example, and depending on the embodiment, the number of convolution layers and activation layers may vary. Additionally, depending on the embodiment, the first neural network (800) may be implemented through a recurrent neural network (RNN). In this case, it means changing the CNN structure of the first neural network (800) according to the example of the present disclosure to an RNN structure.

[0164] In one embodiment, the AI ​​downscaler (712) may include at least one ALU for convolution operations and operations of an activation layer. The ALU may be implemented as a processor. For convolution operations, the ALU may include a multiplier that performs a multiplication operation between sample values ​​of a feature map output from the original image (105) or a previous layer and sample values ​​of a filter kernel, and an adder that adds the results of the multiplication. Additionally, for operations of an activation layer, the ALU may include a multiplier that multiplies an input sample value by a weight used in a predetermined sigmoid function, Tanh function, or ReLU function, and a comparator that compares the result of the multiplication with a predetermined value to determine whether to pass the input sample value to the next layer.

[0165] Referring again to FIG. 7, the AI ​​setting unit (718) transmits AI data to the data processing unit (716). The AI ​​data includes information that enables the AI ​​upscale unit (236) to AI upscale the second image (135) to an upscale target corresponding to the downscale target of the first neural network.

[0166] The first encoding unit (714), which receives the first image (115) from the AI ​​downscale unit (712), can reduce the amount of information contained in the first image (115) by first encoding the first image (115) according to a frequency conversion-based image compression method. As a result of the first encoding through a predetermined codec (e.g., MPEG-2, H.264, MPEG-4, HEVC, VC-1, VP8, VP9, ​​or AV1), image data is obtained. The image data is obtained according to the rules of the predetermined codec, e.g., syntax. The image data may include residual data, which is the difference between the first image (115) and the prediction data of the first image (115), prediction mode information used to first encode the first image (115), motion information, and quantization parameter related information used to first encode the first image (115).

[0167] The image data obtained as a result of the first encoding of the first encoding unit (714) is provided to the data processing unit (716).

[0168] The data processing unit (716) generates AI encoded data including image data received from the first encoding unit (714) and AI data received from the AI ​​setting unit (718).

[0169] In one embodiment, the data processing unit (716) can generate AI encoded data that includes video data and AI data in separate states. For example, the AI ​​data may be included in a VSIF (Vendor Specific InfoFrame) within an HDMI stream.

[0170] In another embodiment, the data processing unit (716) may include AI data within the image data obtained as a result of the first encoding by the first encoding unit (714) and generate AI encoded data including the image data. For example, the data processing unit (716) may combine a bitstream corresponding to the image data and a bitstream corresponding to the AI ​​data to generate image data in the form of a single bitstream. To this end, the data processing unit (716) may represent the AI ​​data as bits having a value of 0 or 1, i.e., a bitstream. In one embodiment, the data processing unit (716) may include the bitstream corresponding to the AI ​​data in the SEI (Supplemental enhancement information), which is an additional information area of ​​the bitstream obtained as a result of the first encoding.

[0171] AI encoded data is transmitted to the transmission unit (730). The transmission unit (730) transmits the AI ​​encoded data obtained as an AI encoding result through the network.

[0172] In one embodiment, AI encoded data may be stored on a data storage medium including a magnetic medium such as a hard disk, a floppy disk and a magnetic tape, an optical recording medium such as a CD-ROM and a DVD, a magneto-optical medium such as a floptical disk, etc.

[0173] FIG. 9 is a diagram showing the structure of AI encoded data (900) according to one embodiment.

[0174] As described above, AI data (912) and video data (932) may be included separately within the AI ​​encoded data (900). Here, the AI ​​encoded data (900) may be a container format such as MP4, AVI, MKV, FLV, etc. The AI ​​encoded data (900) may consist of a metadata box (910) and a media data box (930).

[0175] The metadata box (910) contains information regarding video data (932) included in the media data box (930). For example, the metadata box (910) may contain information regarding the type of the first video (115), the type of codec used to encode the first video (115), and the playback time of the first video (115). Additionally, the metadata box (910) may contain AI data (912). The AI ​​data (912) may be encoded according to an encoding method provided by a predetermined container format and stored in the metadata box (910).

[0176] The media data box (930) may include image data (932) generated according to the syntax of a predetermined image compression method.

[0177] FIG. 10 is a diagram showing the structure of AI encoded data (1000) according to another embodiment.

[0178] Referring to FIG. 10, AI data (1034) may be included in video data (1032). AI encoded data (1000) may include a metadata box (1010) and a media data box (1030), but when AI data (1034) is included in video data (1032), the AI ​​data (1034) may not be included in the metadata box (1010).

[0179] The media data box (1030) contains video data (1032) including AI data (1034). For example, the AI ​​data (1034) may be included in an additional information area of ​​the video data (1032).

[0180] Hereinafter, with reference to FIG. 11, a method for training the first neural network (800) and the second neural network (300) in conjunction will be described.

[0181] FIG. 11 is a diagram illustrating a method for training a first neural network (800) and a second neural network (300).

[0182] In one embodiment, the original video (105) that has been AI encoded through an AI encoding process is restored to a third video (145) through an AI decoding process. In order to maintain similarity between the third video (145) obtained as a result of AI decoding and the original video (105), a correlation between the AI ​​encoding process and the AI ​​decoding process is required. That is, information lost during the AI ​​encoding process must be recoverable during the AI ​​decoding process. To achieve this, linked training of the first neural network (800) and the second neural network (300) is required.

[0183] For accurate AI decoding, it is ultimately necessary to reduce quality loss information (1130) corresponding to the comparison result between the third training image (1104) and the original training image (1101) shown in FIG. 11. Accordingly, quality loss information (1130) is used for training both the first neural network (800) and the second neural network (300).

[0184] First, the training process illustrated in Fig. 11 will be explained.

[0185] In FIG. 11, the original training image (1101) is the image to be AI downscaled, and the first training image (1102) is the image AI downscaled from the original training image (1101). Additionally, the third training image (1104) is the image AI upscaled from the first training image (1102).

[0186] The original training image (1101) includes a still image or a video composed of multiple frames. In one embodiment, the original training image (1101) may include a luminance image extracted from a still image or a video composed of multiple frames. In addition, in one embodiment, the original training image (1101) may include a patch image extracted from a still image or a video composed of multiple frames. When the original training image (1101) is composed of multiple frames, the first training image (1102), the second training image, and the third training image (1104) are also composed of multiple frames. When multiple frames of the original training image (1101) are sequentially input to the first neural network (800), multiple frames of the first training image (1102), the second training image, and the third training image (1104) can be sequentially acquired through the first neural network (800) and the second neural network (300).

[0187] For the combined training of the first neural network (800) and the second neural network (300), the original training image (1101) is input into the first neural network (800). The original training image (1101) input into the first neural network (800) is AI downscaled and output as the first training image (1102), and the first training image (1102) is input into the second neural network (300). The third training image (1104) is output as a result of AI upscaling the first training image (1102).

[0188] Referring to FIG. 11, a first training image (1102) is input to a second neural network (300). Depending on the embodiment, a second training image obtained through the first encoding and first decoding processes of the first training image (1102) may be input to the second neural network (300). Any one of MPEG-2, H.264, MPEG-4, HEVC, VC-1, VP8, VP9, ​​and AV1 may be used to input the second training image to the second neural network. Specifically, any one of MPEG-2, H.264, MPEG-4, HEVC, VC-1, VP8, VP9, ​​and AV1 may be used for the first encoding of the first training image (1102) and the first decoding of image data corresponding to the first training image (1102).

[0189] Referring to FIG. 11, a legacy downscaled reduced training image (1103) is obtained from the original training image (1101), separate from the output of the first training image (1102) through the first neural network (800). Here, the legacy downscale may include at least one of a bilinear scale, a bicubic scale, a lanczos scale, and a stair step scale.

[0190] In order to prevent the structural features of the first image (115) from deviating significantly from the structural features of the original image (105), a reduced training image (1103) that preserves the structural features of the original training image (1101) is obtained.

[0191] Before training proceeds, the first neural network (800) and the second neural network (300) can be set with predetermined neural network configuration information. As training proceeds, structural loss information (1110), complexity loss information (1120), and quality loss information (1130) can be determined.

[0192] Structural loss information (1110) can be determined based on the comparison result between the reduced training image (1103) and the first training image (1102). In one example, structural loss information (1110) may correspond to the difference between the structural information of the reduced training image (1103) and the structural information of the first training image (1102). Structural information may include various features extractable from the image, such as the image's brightness, contrast, and histogram. Structural loss information (1110) indicates the extent to which the structural information of the original training image (1101) is preserved in the first training image (1102). The smaller the structural loss information (1110), the more similar the structural information of the first training image (1102) is to the structural information of the original training image (1101).

[0193] Complexity loss information (1120) can be determined based on the spatial complexity of the first training image (1102). In one example, the total variance value of the first training image (1102) can be used as the spatial complexity. Complexity loss information (1120) is related to the bitrate of the image data obtained by first encoding the first training image (1102). It is defined that the smaller the complexity loss information (1120), the smaller the bitrate of the image data.

[0194] Quality loss information (1130) can be determined based on the comparison results between the original training video (1101) and the third training video (1104). Quality loss information (1130) may include at least one of an L1-norm value, an L2-norm value, a SSIM (Structural Similarity) value, a PSNR-HVS (Peak Signal-To-Noise Ratio-Human Vision System) value, a MS-SSIM (Multiscale SSIM) value, a VIF (Variance Inflation Factor) value, and a VMAF (Video Multimethod Assessment Fusion) value regarding the difference between the original training video (1101) and the third training video (1104). Quality loss information (1130) indicates the degree to which the third training video (1104) is similar to the original training video (1101). The smaller the quality loss information (1130), the more similar the third training image (1104) is to the original training image (1101).

[0195] Referring to FIG. 11, structural loss information (1110), complexity loss information (1120), and quality loss information (1130) are used for training the first neural network (800), and the quality loss information (1130) is used for training the second neural network (300). That is, the quality loss information (1130) is used for training both the first neural network (800) and the second neural network (300).

[0196] The first neural network (800) can update parameters so that the final loss information determined based on structural loss information (1110), complexity loss information (1120), and quality loss information (1130) is reduced or minimized. Additionally, the second neural network (300) can update parameters so that the quality loss information (1130) is reduced or minimized.

[0197] The final loss information for training the first neural network (800) and the second neural network (300) can be determined as shown in Equation 1 below.

[0199] [Mathematical Formula 1]

[0200]

[0202] In the above mathematical formula 1, LossDS represents the final loss information to be reduced or minimized for training the first neural network (800), and LossUS represents the final loss information to be reduced or minimized for training the second neural network (300). Additionally, a, b, c, and d may correspond to predetermined weights.

[0203] That is, the first neural network (800) updates the parameters in the direction in which the LossDS of Equation 1 decreases, and the second neural network (300) updates the parameters in the direction in which the LossUS decreases. When the parameters of the first neural network (800) are updated according to the LossDS derived during the training process, the first training image (1102) obtained based on the updated parameters becomes different from the first training image (1102) in the previous training process, and accordingly, the third training image (1104) also becomes different from the third training image (1104) in the previous training process. When the third training image (1104) becomes different from the third training image (1104) in the previous training process, the quality loss information (1130) is also newly determined, and accordingly, the second neural network (300) updates the parameters. When the quality loss information (1130) is newly determined, the LossDS is also newly determined, so the first neural network (800) updates its parameters according to the newly determined LossDS. That is, the parameter update of the first neural network (800) causes the parameter update of the second neural network (300), and the parameter update of the second neural network (300) causes the parameter update of the first neural network (800). In other words, since the first neural network (800) and the second neural network (300) are trained in conjunction through the sharing of the quality loss information (1130), the parameters of the first neural network (800) and the parameters of the second neural network (300) can be optimized with a correlation to each other.

[0204] Referring to Equation 1, it can be seen that LossUS is determined according to quality loss information (1130), but this is just one example, and LossUS may also be determined based on at least one of structural loss information (1110) and complexity loss information (1120) and quality loss information (1130).

[0205] Previously, it was explained that the AI ​​setting unit (238) of the AI ​​decoding device (200) and the AI ​​setting unit (718) of the AI ​​encoding device (700) store multiple neural network setting information. Now, a method for training each of the multiple neural network setting information stored in the AI ​​setting unit (238) and the AI ​​setting unit (718) will be described.

[0206] As explained in relation to Equation 1, in the case of the first neural network (800), parameters are updated by considering the degree of similarity between the structural information of the first training image (1102) and the structural information of the original training image (1101) (structural loss information (1110)), the bitrate of the image data obtained from the first encoding result of the first training image (1102) (complexity loss information (1120)), and the difference between the third training image (1104) and the original training image (1101) (quality loss information (1130)).

[0207] To explain in detail, the parameters of the first neural network (800) can be updated so that a first training image (1102) can be obtained which is similar to the structural information of the original training image (1101) and has a small bit rate of image data obtained when the first encoding is performed, and at the same time, a second neural network (300) that AI upscales the first training image (1102) can obtain a third training image (1104) similar to the original training image (1101).

[0208] By adjusting the weights of a, b, and c in mathematical formula 1, the direction in which the parameters of the first neural network (800) are optimized becomes different. For example, if the weight of b is set high, the parameters of the first neural network (800) may be updated with greater importance placed on lowering the bitrate than on the quality of the third training video (1104). Also, if the weight of c is set high, the parameters of the first neural network (800) may be updated with greater importance placed on increasing the quality of the third training video (1104) than on increasing the bitrate or maintaining the structural information of the original training video (1101).

[0209] In addition, the direction in which the parameters of the first neural network (800) are optimized may differ depending on the type of codec used to first encode the first training image (1102). This is because, depending on the type of codec, the second training image to be input to the second neural network (300) may differ.

[0210] That is, the parameters of the first neural network (800) and the parameters of the second neural network (300) can be updated in conjunction based on weights a, b, and c, and the type of codec for the first encoding of the first training image (1102). Accordingly, when weights a, b, and c are each determined to a predetermined value and the type of codec is determined to a predetermined type, and the first neural network (800) and the second neural network (300) are trained, the parameters of the first neural network (800) and the parameters of the second neural network (300) that are optimized in conjunction with each other can be determined.

[0211] And, after changing the weights a, b, c, and the type of codec, if the first neural network (800) and the second neural network (300) are trained, the parameters of the first neural network (800) and the parameters of the second neural network (300) that are optimized in conjunction with each other can be determined. In other words, if the first neural network (800) and the second neural network (300) are trained while changing the values ​​of each of the weights a, b, c, and the type of codec, multiple neural network configuration information that is trained in conjunction with each other can be determined in the first neural network (800) and the second neural network (300).

[0212] As previously explained in relation to FIG. 5, multiple neural network configuration information of the first neural network (800) and the second neural network (300) may be mapped to first image-related information. To establish this mapping relationship, a first training image (1102) output from the first neural network (800) may be first encoded with a specific codec according to a specific bit rate, and a second training image obtained by first decoding the bitstream obtained from the first encoding result may be input to the second neural network (300). That is, by setting up an environment so that a first training image (1102) of a specific resolution is first encoded by a specific codec at a specific bit rate, and then training the first neural network (800) and the second neural network (300), a pair of neural network setting information mapped to the resolution of the first training image (1102), the type of codec used for the first encoding of the first training image (1102), and the bit rate of the bitstream obtained as a result of the first encoding of the first training image (1102) can be determined. By varying the resolution of the first training image (1102), the type of codec used for the first encoding of the first training image (1102), and the bit rate of the bitstream obtained according to the first encoding of the first training image (1102), a mapping relationship between the multiple neural network setting information of the first neural network (800) and the second neural network (300) and the first image-related information can be determined.

[0213] FIG. 12 is a diagram illustrating the training process of the first neural network (800) and the second neural network (300) by the training device (1200).

[0214] Training of the first neural network (800) and the second neural network (300) described in relation to FIG. 11 can be performed by a training device (1200). The training device (1200) includes the first neural network (800) and the second neural network (300). The training device (1200) may be, for example, an AI encoding device (700) or a separate server. The neural network configuration information of the second neural network (300) obtained as a result of training is stored in an AI decoding device (200).

[0215] Referring to FIG. 12, the training device (1200) initializes neural network configuration information of the first neural network (800) and the second neural network (300) (S1240, S1245). By doing so, the first neural network (800) and the second neural network (300) can operate according to the predetermined neural network configuration information. The neural network configuration information may include information on at least one of the number of convolution layers included in the first neural network (800) and the second neural network (300), the number of filter kernels per convolution layer, the size of the filter kernels per convolution layer, and the parameters of each filter kernel.

[0216] The training device (1200) inputs the original training image (1101) into the first neural network (800) (S1250). The original training image (1101) may include at least one frame constituting a still image or a video.

[0217] The first neural network (800) processes the original training image (1101) according to the initial neural network setting information and outputs the first training image (1102) AI downscaled from the original training image (1101) (S1255). FIG. 12 is illustrated as the first training image (1102) output from the first neural network (800) being directly input to the second neural network (300), but the first training image (1102) output from the first neural network (800) may be input to the second neural network (300) by the training device (1200). Additionally, the training device (1200) may first encode and first decode the first training image (1102) using a predetermined codec, and then input the second training image to the second neural network (300).

[0218] The second neural network (300) processes the first training image (1102) or the second training image according to the initial neural network setting information and outputs a third training image (1104) that has been AI upscaled from the first training image (1102) or the second training image (S1260).

[0219] The training device (1200) produces complexity loss information (1120) based on the first training image (1102) (S1265).

[0220] The training device (1200) compares the reduced training image (1103) and the first training image (1102) to produce structural loss information (1110) (S1270).

[0221] The training device (1200) compares the original training image (1101) and the third training image (1104) to produce quality loss information (1130) (S1275).

[0222] The first neural network (800) updates the initially set neural network configuration information through a back propagation process based on the final loss information (S1280). The training device (1200) can calculate the final loss information for training the first neural network (800) based on the complexity loss information (1120), the structural loss information (1110), and the quality loss information (1130).

[0223] The second neural network (300) updates the initially set neural network configuration information through a backtranslation process based on quality loss information or final loss information (S1285). The training device (1200) can calculate final loss information for training the second neural network (300) based on quality loss information (1130).

[0224] Subsequently, the training device (1200), the first neural network (800), and the second neural network (300) update the neural network configuration information by repeating the process S1250 to S1285 until the final loss information is minimized. At this time, during each iteration process, the first neural network (800) and the second neural network (300) operate according to the neural network configuration information updated in the previous process.

[0225] Table 1 below shows the effects of AI encoding and AI decoding of the original image (105) according to one embodiment of the present disclosure, and of HEVC encoding and decoding of the original image (105).

[0226] [Table 1]

[0227]

[0228] As can be seen from Table 1, according to one embodiment of the present disclosure, when content consisting of 300 frames of 8K resolution is AI encoded and AI decoded, the average subjective quality is higher than when it is HEVC encoded and decoded, yet the bitrate is reduced by more than 50%.

[0229] The aforementioned AI encoding device (700) and AI decoding device (200) may be useful for a server-client structure. Specifically, the server obtains a first image (115) by AI downscaling the original image (105) in response to a client's image request, and transmits AI encoded data including image data obtained as a result of encoding the first image (115) to the client.

[0230] The client decodes the video data to obtain a second video (135) and displays a third video (145) obtained through AI upscaling of the second video (135).

[0231] Generally, since the capacity of the server's storage media (e.g., memory, hard disk, SSD (Solid State Drive), etc.) is very large, there is no major problem even if a large number of original videos (105) are continuously stored. However, if the aforementioned AI encoding device (700) is implemented as an electronic device with a small storage capacity, such as a mobile device like a smartphone, laptop, or tablet PC, it may be difficult to store the original videos (105) themselves. For example, if an 8k resolution video is captured with an electronic device, it is difficult to shoot for a long time due to the limitations of storage capacity, and storing the 8k resolution video may make it impossible to store other data.

[0232] Therefore, a method is required to prevent the storage capacity of the electronic device from becoming saturated as the original image (105) with high resolution is stored.

[0233] The embodiments described below may be useful for the structure of a mobile client. Specifically, the mobile device can prevent storage capacity from becoming saturated due to the large size of the original image (105) by AI downscaling and encoding the high-resolution original image (105) in advance.

[0234] The mobile device may provide the first image data of the first image (115) AI downscaled from the original image (105) to the client according to the client's request or the user's request, or provide the second image data of the third image (145) AI upscaled from the second image (135) to the client according to the client's function.

[0235] Hereinafter, with reference to FIGS. 13 to 24, an AI encoding process and an AI decoding process for preventing storage capacity from becoming saturated due to a large original image (105) will be described.

[0236] FIG. 13 is a block diagram illustrating the configuration of an image providing device (1300) according to one embodiment.

[0237] Referring to FIG. 13, the image providing device (1300) may include an AI encoding unit (1310) and a transmission unit (1330). The AI ​​encoding unit (1310) may include an AI scaling unit (1312), a first encoding unit (1314), a data processing unit (1316), an AI setting unit (1318), and a first decoding unit (1319).

[0238] FIG. 13 illustrates the AI ​​encoding unit (1310) and the transmission unit (1330) as separate devices, but the AI ​​encoding unit (1310) and the transmission unit (1330) can be implemented through a single processor. In this case, it may be implemented with a dedicated processor, or it may be implemented through a combination of a general-purpose processor such as an AP, CPU, or GPU and S / W. In addition, in the case of a dedicated processor, it may include memory for implementing the embodiments of the present disclosure or a memory processing unit for using external memory.

[0239] Additionally, the AI ​​encoding unit (1310) and the transmission unit (1330) may be composed of multiple processors. In this case, they may be implemented by a combination of dedicated processors, or by a combination of multiple general-purpose processors such as APs, CPUs, or GPUs and software.

[0240] In one embodiment, the first encoding unit (1314) and the first decoding unit (1319) are configured as a first processor, the AI ​​scaling unit (1312), the data processing unit (1316) and the AI ​​setting unit (1318) are implemented as a second processor different from the first processor, and the transmission unit (1330) may be implemented as a third processor different from the first processor and the second processor.

[0241] The AI ​​encoding unit (1310) acquires the original image (105). The AI ​​encoding unit (1310) may acquire the original image (105) captured by the camera of the image providing device (1300), or may acquire the original image (105) received from an external device via a network.

[0242] The AI ​​encoding unit (1310) obtains performance information of a display device to play a video. When the AI ​​encoding unit (1310) receives a request to provide a video from a display device or a user, it obtains performance information of the display device. Performance information of the display device may be received from the display device or input from the user.

[0243] Although not illustrated in FIG. 13, the image providing device (1300) may further include a receiver for receiving original image (105) and performance information of the display device.

[0244] Performance information of the display device may include information indicating whether the AI ​​upscaling function is supported.

[0245] The AI ​​encoding unit (1310) may perform only AI encoding on the original image (105) or perform both AI encoding and AI decoding on the original image (105), depending on whether the display device supports the AI ​​upscaling function.

[0246] The operation of the components included in the AI ​​encoding unit (1310) is described in detail below.

[0247] The AI ​​scaling unit (1312) obtains a first image (115) by AI downscaling the original image (105). The AI ​​scaling unit (1312) can AI downscale the original image (105) through a first neural network to which neural network setting information is applied. The first neural network used by the AI ​​scaling unit (1312) may have the structure of the first neural network (800) shown in FIG. 8, but the structure of the first neural network is not limited to the first neural network (800) shown in FIG. 8.

[0248] In the following, neural network configuration information applicable to a first neural network for AI downscaling is referred to as 'first neural network configuration information', and neural network configuration information applicable to a second neural network for AI upscaling is referred to as 'second neural network configuration information'. The first neural network configuration information (or the second neural network configuration information) may include information on at least one of the number of convolution layers included in the first neural network (or the second neural network), the number of filter kernels per convolution layer, and the parameters of each filter kernel.

[0249] The AI ​​setting unit (1318) provides first neural network setting information for AI downscaling to the AI ​​scale unit (1312), so that the AI ​​scale unit (1312) AI downscales the original image (105) according to the first neural network setting information.

[0250] The number of first neural network setting information stored in the AI ​​setting unit (1318) may be one or multiple. If one first neural network setting information is stored in the AI ​​setting unit (1318), the AI ​​setting unit (1318) transmits one first neural network setting information to the AI ​​scale unit (1312). If multiple first neural network setting information is stored in the AI ​​setting unit (1318), the AI ​​setting unit (1318) selects a first neural network setting information to be used for AI downscaling among the multiple first neural network setting information, and transmits the selected first neural network setting information to the AI ​​scale unit (1312).

[0251] One or more first neural network setting information and one or more second neural network setting information for AI upscaling stored in the AI ​​setting unit (1318) will be explained with reference to FIGS. 17 to 20.

[0252] The AI ​​setting unit (1318) provides AI data related to AI downscaling of the original image (105) to the data processing unit (1316). As described above, the AI ​​data includes information that enables the AI ​​upscaling unit of the display device to AI upscale the second image (135) to an upscaling target corresponding to the downscaling target of the first neural network.

[0253] In one example, the AI ​​data may include difference information between the original image (105) and the first image (115).

[0254] In another example, the AI ​​data may include information related to the first image (115). The information related to the first image (115) may include information about at least one of the resolution of the first image (115), the bitrate of the first image data obtained as a result of encoding the first image (115), and the codec type used when encoding the first image (115).

[0255] In another example, the AI ​​data may include an identifier of mutually agreed neural network configuration information so that the second image (135) can be AI upscaled to an upscale target corresponding to the downscale target of the first neural network.

[0256] In another example, the AI ​​data may include second neural network configuration information that can be set in the second neural network.

[0257] The first encoding unit (1314) encodes the first image (115) to obtain first image data. The first image data generated by the first encoding unit (1314) can be stored in a storage medium of the image providing device (1300).

[0258] The first image data includes data obtained as a result of encoding the first image (115). The first image data may include data obtained based on pixel values ​​within the first image (115), for example, residual data which is the difference between the first image (115) and the prediction data of the first image (115). Additionally, the first image data includes information used during the encoding process of the first image (115). For example, the first image data may include prediction mode information, motion information, and quantization parameter-related information used to encode the first image (115).

[0259] When a request for video provision is received from a display device or user, and the display device supports an AI upscale function, the first video data and AI data are provided from the first encoding unit (1314) and the AI ​​setting unit (1318) to the data processing unit (1316).

[0260] The data processing unit (1316) generates AI encoded data including the first image data received from the first encoding unit (1314) and the AI ​​data received from the AI ​​setting unit (1318).

[0261] In one embodiment, the data processing unit (1316) can generate AI encoded data that includes the first video data and the AI ​​data in separate states. For example, the AI ​​data may be included in a VSIF (Vendor Specific InfoFrame) within an HDMI stream.

[0262] In another embodiment, the data processing unit (1316) may include AI data within the first image data obtained as a result of encoding by the first encoding unit (1314) and generate AI encoded data including the first image data. For example, the data processing unit (1316) may combine a bitstream corresponding to the first image data and a bitstream corresponding to the AI ​​data to generate image data in the form of a single bitstream. To this end, the data processing unit (1316) may represent the AI ​​data as bits having a value of 0 or 1, i.e., a bitstream. In one embodiment, the data processing unit (1316) may include the bitstream corresponding to the AI ​​data in the SEI (Supplemental enhancement information), which is an additional information area of ​​the bitstream obtained as a result of encoding.

[0263] AI encoded data is transmitted to the transmission unit (1330). The transmission unit (1330) transmits the AI ​​encoded data obtained as an AI encoding result to a display device via a network.

[0264] When a request for video provision is received from a display device or user, and the display device does not support the AI ​​upscale function, the first video data obtained by the first encoding unit (1314) is provided to the first decoding unit (1319).

[0265] The first decoding unit (1319) decodes the first image data to obtain the second image (135) and provides the second image (135) to the AI ​​scale unit (1312).

[0266] The AI ​​scaling unit (1312) obtains a third image (145) by AI upscaling the second image (135). The AI ​​scaling unit (1312) can AI upscale the second image (135) through a second neural network to which the second neural network setting information is applied. The second neural network used by the AI ​​scaling unit (1312) may have the structure of the second neural network (300) shown in FIG. 3, but the structure of the second neural network is not limited to the second neural network (300) shown in FIG. 3.

[0267] The AI ​​setting unit (1318) provides second neural network setting information for AI upscaling to the AI ​​scaling unit (1312). However, if the display device does not support the AI ​​upscaling function, it may not provide AI data related to AI downscaling of the original image (105) and AI data related to AI upscaling of the second image (135) to the data processing unit (1316). This is because if the display device cannot perform AI upscaling, the AI ​​data cannot be used even if it is received from the image providing device (1300).

[0268] According to an embodiment, even if the display device does not support an AI upscale function, the AI ​​setting unit (1318) may provide AI data related to AI downscale of the original image (105) and / or AI data related to AI upscale of the second image (135) to the data processing unit (1316).

[0269] The AI ​​scaling unit (1312) provides the third image (145) obtained through AI upscaling of the second image (135) to the first encoding unit (1314), and the first encoding unit (1314) encodes the third image (145) to obtain the second image data.

[0270] The second image data includes data obtained as a result of encoding the second image (135). The second image data may include data obtained based on pixel values ​​within the second image (135), for example, residual data which is the difference between the second image (135) and the prediction data of the second image (135). Additionally, the second image data includes information used during the encoding process of the second image (135). For example, the second image data may include prediction mode information, motion information, and quantization parameter-related information used to encode the second image (135).

[0271] The first encoding unit (1314) provides the second image data to the data processing unit (1316).

[0272] The data processing unit (1316) generates AI encoded data including second image data received from the first encoding unit (1314). When AI data is received from the AI ​​setting unit (1318), the data processing unit (1316) generates AI encoded data including the second image data and AI data.

[0273] AI encoded data is transmitted to the transmission unit (1330). The transmission unit (1330) transmits the AI ​​encoded data obtained as an AI encoding result to a display device via a network.

[0274] According to one embodiment of the present disclosure, when an original image (105) is obtained, the image providing device (1300) performs AI downscaling on the original image (105) to obtain a first image (115), and the original image (105) can be deleted from the storage medium.

[0275] The image providing device (1300) can prevent the storage capacity of the storage medium from becoming saturated by storing a first image (115) (specifically, first image data) having a resolution lower than that of the original image (105) instead of the original image (105).

[0276] Additionally, the video providing device (1300) can transmit AI encoded data generated through AI encoding to the display device, or transmit second video data generated through AI encoding and AI decoding to the display device, by considering whether the display device to play the video supports an AI upscale function, so that a third video (145) of high resolution can be played through the display device.

[0277] Meanwhile, as previously mentioned, since the original image (105) and the third image (145) of high resolution occupy a large amount of storage space, it is important to quickly delete the original image (105) and the third image (145) in order to stably maintain the storage capacity of the image providing device (1300).

[0278] In order to perform AI downscaling of the original video (105) and encoding of the third video (145), the original video (105) and the third video (145) need to be temporarily stored in a storage medium, and in order to secure storage space, the storage time of the original video (105) and the third video (145) needs to be minimized.

[0279] In one embodiment of the present disclosure, by performing AI downscaling of the original image (105) and encoding of the third image (145) on a frame-by-frame basis, storage capacity can be prevented from becoming saturated due to the original image (105) and the third image (145) of high resolution. This will be explained with reference to FIGS. 14 to 16.

[0280] In the following, the frame constituting the original image (105) is referred to as the 'original frame', and the frame constituting the first image (115) is referred to as the 'first frame'. Additionally, the frame constituting the second image (135) is referred to as the 'second frame', and the frame constituting the third image (145) is referred to as the 'third frame'.

[0281] FIG. 14 is a diagram illustrating the processing of original frames when the display device does not support the AI ​​upscaling function.

[0282] Referring to FIG. 14, the camera (1410) responds to external light to sequentially acquire original frame a (105a), original frame b (105b), original frame c (105c), and original frame d (105d).

[0283] Original frame a (105a), original frame b (105b), original frame c (105c), and original frame d (105d) are sequentially AI downscaled by the first neural network (1420). FIG. 14 illustrates only four original frames constituting the original image (105), but this is for the sake of simplicity of explanation, and the number of frames constituting the original image (105) may vary.

[0284] It is just one example that original frames are obtained by the camera (1410). Original frames may also be obtained from an external device via a network.

[0285] The first neural network (1420) outputs a first frame a (115a), a first frame b (115b), a first frame c (115c), and a first frame d (115d) corresponding to the original frame a (105a), original frame b (105b), original frame c (105c), and original frame d (105d), respectively, as the original frame a (105a), original frame b (105b), original frame c (105c), and original frame d (105d) are input sequentially.

[0286] The image providing device (1300) can minimize the number of original frames stored in the image providing device (1300) by deleting original frames for which AI downscaling has been completed by the first neural network (1420) among original frame a (105a), original frame b (105b), original frame c (105c), and original frame d (105d). For example, while original frames are sequentially acquired by the camera (1410), AI downscaling by the first neural network (1420) is performed sequentially on the original frames, and original frames for which AI downscaling has been completed can be sequentially deleted.

[0287] Specifically, after the original frame a (105a) is acquired by the camera (1410), AI downscaling is performed on the original frame a (105a) to acquire the first frame a (115a), the image providing device (1300) deletes the original frame a (105a). Then, when the original frame b (105b) is acquired by the camera (1410), the image providing device (1300) inputs the original frame b (105b) into the first neural network (1420) to AI downscale the original frame b (105b), and deletes the original frame b (105b) after the AI ​​downscaling is completed.

[0288] If the display device supports an AI upscale function, first image data is obtained through encoding for the first frame a (115a), the first frame b (115b), the first frame c (115c), and the first frame d (115d), and as described above, AI encoded data including the first image data is transmitted to the display device.

[0289] If the display device does not support the AI ​​upscale function, the first frame a (115a), the first frame b (115b), the first frame c (115c), and the first frame d (115d) are encoded and decoded (1430) to obtain the second frame a (135a), the second frame b (135b), the second frame c (135c), and the third frame d (145d).

[0290] Compared to the original frames being deleted sequentially after AI downscaling is complete, the first frame a (115a), the first frame b (115b), the first frame c (115c), and the first frame d (115d) may be deleted together after encoding is complete. Additionally, the second frame a (135a), the second frame b (135b), the second frame c (135c), and the second frame d (135d) may also be deleted together after AI upscaling through the second neural network (1440) is complete. This is because the first frames and the second frames do not occupy much storage space due to their low resolution.

[0291] The second frame a (135a), the second frame b (135b), the second frame c (135c), and the second frame d (135d) are sequentially processed by the second neural network (1440), and the AI ​​upscaled third frame a (145a), the third frame b (145b), the third frame c (145c), and the third frame d (145d) are sequentially obtained from the second neural network (1440).

[0292] Compared to the original high-resolution frames (105a, 105b, 105c, 105d) being deleted immediately after the completion of AI downscaling, the third high-resolution frames (145a, 145b, 145c, 145d) are not deleted immediately after the completion of AI upscaling for the second frames (135a, 135b, 135c, 135d). This is because the third frames (145a, 145b, 145c, 145d) are additionally subject to encoding (1450).

[0293] Generally, frames constituting an image are encoded using intra-prediction or inter-prediction. Intra-prediction is an encoding tool that reduces spatial redundancy by predicting samples within a frame from surrounding samples, while inter-prediction is an encoding tool that reduces temporal redundancy by predicting samples within a frame from samples in other frames. Since intra-prediction and inter-prediction are widely used in general-purpose codecs such as HEVC, a detailed explanation is omitted.

[0294] A group of mutually referenced frames is referred to as a group of pictures (GOP) for inter-prediction. Since another third frame must exist when one third frame is encoded (1450) in order to predict a sample within one third frame from another third frame, the image providing device (1300) deletes the encoded third frames when the encoding (1450) for the third frames constituting the GOP is completed.

[0295] Referring to FIG. 14, third frame a (145a) and third frame b (145b) form the first GOP, and third frame c (145c) and third frame d (145d) form the second GOP.

[0296] Encoding (1450) for the third frame a (145a) and the third frame b (145b) is performed to obtain second image data corresponding to the third frame a (145a) and the third frame b (145b), and encoding (1450) for the third frame c (145c) and the third frame d (145d) is performed to obtain second image data corresponding to the third frame c (145c) and the third frame d (145d).

[0297] The video providing device (1300) deletes the third frame a (145a) and the third frame b (145b) when the encoding (1450) for the third frame a (145a) and the third frame b (145b) constituting the first GOP is completed, and deletes the third frame c (145c) and the third frame d (145d) when the encoding (1450) for the third frame c (145c) and the third frame d (145d) constituting the second GOP is completed.

[0298] The image providing device (1300) can minimize the number of third frames stored in the image providing device (1300) by performing encoding (1450) in GOP units and immediately deleting the GOP that has been encoded (1450).

[0299] FIG. 15 is a flowchart for explaining the AI ​​downscaling process of the original frames shown in FIG. 14.

[0300] In step S1510, the video providing device (1300) sequentially AI downscales the original frames. AI downscales for the original frames can be performed based on the first neural network (1420).

[0301] In step S1520, the video providing device (1300) deletes the original frame that has been AI downscaled.

[0302] In step S1530, the image providing device (1300) determines whether there is an original frame for which AI downscaling has not been completed.

[0303] If there is an original frame for which AI downscaling has not been completed, the video providing device (1300) repeats steps S1510, S1520, and S1530.

[0304] FIG. 16 is a flowchart for explaining the process of AI upscaling the second frames shown in FIG. 14.

[0305] In step S1610, the image providing device (1300) sequentially AI upscales the second frames to obtain the third frames. The AI ​​upscaling of the second frames can be performed based on the second neural network (1440).

[0306] In step S1620, the image providing device (1300) encodes the third frames constituting the GOP. The image providing device (1300) can encode the third frames constituting the GOP when the third frames constituting the GOP are acquired while AI upscaling for the second frames is performed sequentially.

[0307] In step S1630, the image providing device (1300) deletes the third frames constituting the encoded GOP.

[0308] In step S1640, the image providing device (1300) determines whether there is a second frame for which AI upscaling has not been completed.

[0309] If there is a second frame for which AI upscaling is not completed, the video providing device (1300) repeats steps S1610, S1620, S1630, and S1640.

[0310] Hereinafter, with reference to FIGS. 17 to 20, first neural network setting information that can be set in the first neural network (1420) and second neural network setting information that can be set in the second neural network (1440) will be described.

[0311] FIG. 17 is a diagram illustrating first neural network configuration information applicable to a first neural network (1420) for AI downscaling, and second neural network configuration information applicable to a second neural network (1440) for AI upscaling.

[0312] The AI ​​setting unit (1318) can store one 'A1' neural network setting information applicable to the first neural network (1420) and one 'A2' neural network setting information applicable to the second neural network (1440).

[0313] The AI ​​setting unit (1318) provides the 'A1' neural network setting information to the AI ​​scaling unit (1312) so that the original image (105) is AI downscaled through the first neural network (1420) to which the 'A1' neural network setting information is applied (or set).

[0314] Additionally, the AI ​​setting unit (1318) provides 'A2' neural network setting information to the AI ​​scaling unit (1312) so that the second image (135) is AI upscaled through the second neural network (1440) to which the 'A2' neural network setting information is applied (or set).

[0315] A pair of 'A1' neural network configuration information and 'A2' neural network configuration information can be obtained through the combined training of the first neural network (1420) and the second neural network (1440) described above in relation to FIGS. 11 and FIGS. 12.

[0316] Specifically, the first neural network (1420) and the second neural network (1440) are set with predetermined neural network configuration information. Then, the first neural network (1420) updates the neural network configuration information so that the final loss information determined from structural loss information (1110), complexity loss information (1120), and quality loss information (1130) is reduced or minimized. The second neural network (1440) updates the neural network configuration information so that the final loss information determined from quality loss information (1130) is reduced or minimized.

[0317] Through the combined training of the first neural network (1420) and the second neural network (1440) by sharing quality loss information (1130), a pair of corresponding 'A1' neural network configuration information and 'A2' neural network configuration information can be obtained.

[0318] In one embodiment, the scaling factor of the first neural network (1420) to which the 'A1' neural network setting information is applied may be 1 / d (where d is a rational number greater than 1), and the scaling factor of the second neural network (1440) to which the 'A2' neural network setting information is applied may be d (where d is a rational number greater than 1). The AI ​​scaling unit (1312) may AI upscale the second image (135) n times (where n is a natural number greater than or equal to 1) by considering the scaling factor of the second neural network (1440) to which the 'A2' neural network setting information is applied, the resolution of the second image (135), and the supported resolution of the display device. AI upscale n times means that the image output from the second neural network (1440) is input back into the second neural network (1440) n-1 times. For example, when the second image (135) is AI upscaled three times, the second image (135) is AI upscaled by being processed by the second neural network (1440), the output image of the second neural network (1440) is AI upscaled by being processed again by the second neural network (1440), and the output image of the second neural network (1440) is AI upscaled by being processed again by the second neural network (1440) so that the third image (145) can be obtained.

[0319] When the AI ​​scaling unit (1312) AI upscales the second image (135) two or more times, it may be useful when the supported resolution of the display device is greater than the resolution of the third image (145) obtained through one AI downscale of the second image (135).

[0320] For example, if the scaling factor of the second neural network (1440) to which the 'A2' neural network setting information is applied is 2, the supported resolution of the display device is 16k, and the resolution of the second image (135) input to the second neural network (1440) is 2k, then the AI ​​scaling unit (1312) can obtain a third image (145) with a resolution of 16k by AI upscaling the second image (135) three times through the second neural network (1440) to which the 'A2' neural network setting information is applied.

[0321] Information regarding the supported resolution of the display device may be included in the performance information of the display device received by the AI ​​encoding unit (1310). The supported resolution of the display device may indicate at least one of the resolutions that the display can play. For example, the supported resolution of the display device may indicate the largest resolution that the display can play.

[0322] FIG. 18 is a diagram illustrating a plurality of first neural network configuration information applicable to a first neural network (1420) for AI downscaling, and a plurality of second neural network configuration information applicable to a second neural network (1440) for AI upscaling.

[0323] The AI ​​setting unit (1318) can store a plurality of first neural network setting information applicable to the first neural network (1420) and a plurality of second neural network setting information applicable to the second neural network (1440). The plurality of first neural network setting information and the plurality of second neural network setting information can correspond to each other. The plurality of first neural network setting information and the plurality of second neural network setting information can correspond to each other in a 1:1 manner.

[0324] A plurality of first neural network configuration information and a plurality of second neural network configuration information can be obtained through the combined training of the first neural network (1420) and the second neural network (1440) described above in relation to FIGS. 11 and FIGS. 12.

[0325] Specifically, the first neural network (1420) and the second neural network (1440) are set with predetermined neural network configuration information. Then, the first neural network (1420) updates the neural network configuration information so that the final loss information determined from structural loss information (1110), complexity loss information (1120), and quality loss information (1130) is reduced or minimized. The second neural network (1440) updates the neural network configuration information so that the final loss information determined from quality loss information (1130) is reduced or minimized.

[0326] When the first neural network (1420) and the second neural network (1440) are trained in conjunction while changing the training conditions, a plurality of corresponding first neural network configuration information and a plurality of second neural network configuration information can be obtained. Since the method for obtaining a plurality of corresponding first neural network configuration information and a plurality of second neural network configuration information has been explained with reference to FIG. 11 and FIG. 12, a detailed explanation is omitted.

[0327] The AI ​​setting unit (1318) selects a first neural network setting information corresponding to a downscale target of the original image (105) from among a plurality of first neural network setting information and provides it to the AI ​​scaling unit (1312).

[0328] The AI ​​setting unit (1318) acquires one or more input information. In one embodiment, the input information may include at least one of the target resolution of the first image (115), the target bitrate of the first image data, the bitrate type of the first image data (e.g., variable bitrate type, constant bitrate type, or average bitrate type, etc.), the color format to which AI downscale is applied (luminance component, chrominance component, red component, green component, or blue component, etc.), the codec type for encoding the first image (115), compression history information, the resolution of the original image (105), the type of the original image (105), and the content included in the original image (105) (e.g., whether it is a person image, a landscape image, etc.).

[0329] One or more input information may include information that is pre-stored in the image providing device (1300) or received from the user.

[0330] The AI ​​setting unit (1318) determines a downscale target based on input information and can select a first neural network setting information corresponding to the downscale target among a plurality of first neural network setting information. If the information constituting the first neural network setting information (e.g., the number of convolution layers, the number of filter kernels per convolution layer, parameters of each filter kernel, etc.) is stored in the form of a lookup table, the AI ​​setting unit (1318) may select the first neural network setting information by combining some of the selected values ​​from the lookup table according to the downscale target.

[0331] In one embodiment, the AI ​​setting unit (1318) may transmit at least a portion of the input information to the first encoding unit (1314) so ​​that the first encoding unit (1314) can encode the first image (115) with a bit rate of a specific value, a bit rate of a specific type, and a specific codec.

[0332] In one embodiment, the AI ​​setting unit (1318) determines a downscale target based on at least one of a compression rate (e.g., a difference in resolution between the original image (105) and the first image (115), a target bitrate), a compression quality (e.g., a bitrate type), compression history information and a type of the original image (105), and can select a first neural network setting information corresponding to the downscale target.

[0333] As another example, the AI ​​setting unit (1318) may determine a downscale target using compression history information stored in the image providing device (1300) and select a first neural network setting information corresponding to the downscale target. For example, according to the compression history information available to the image providing device (1300), the encoding quality or compression rate preferred by the user may be determined, and the downscale target may be determined according to the encoding quality, etc. determined based on the compression history information. For example, if the encoding quality that has been used most frequently is identified from the compression history information, the resolution, image quality, etc. of the first image (115) may be determined according to the identified encoding quality.

[0334] As another example, if the AI ​​setting unit (1318) identifies a encoding quality that has been used more than a predetermined threshold value from the compression history information (e.g., the average quality of encoding qualities that have been used more than a predetermined threshold value), it may determine a downscale target based on the identified encoding quality.

[0335] As another example, the AI ​​setting unit (1318) may determine a downscale target based on the resolution and type (e.g., file format) of the original image (105).

[0336] In one embodiment, the AI ​​setting unit (1318) may independently select first neural network setting information for each of a predetermined number of original frames and provide the independently selected first neural network setting information to the AI ​​scale unit (1312).

[0337] In another embodiment, the AI ​​setting unit (1318) divides the original frames into a predetermined number of groups and can independently select first neural network setting information for each group. For each group, first neural network setting information that is identical or different from each other may be selected. The number of original frames included in the groups may be the same or different for each group.

[0338] In another embodiment, the AI ​​setting unit (1318) may independently select first neural network setting information for each original frame. For each original frame, first neural network setting information that is the same or different from each other may be selected.

[0339] When AI upscaling is required, the AI ​​setting unit (1318) selects a second neural network setting information corresponding to a first neural network setting information selected for AI downscaling among a plurality of second neural network setting information, and provides the selected second neural network setting information to the AI ​​scaling unit (1312).

[0340] The AI ​​scaling unit (1312) AI upscales the second image (135) through the second DNN using the second neural network setting information received from the AI ​​setting unit (1318).

[0341] In one embodiment, the scaling factor of the first neural network (1420) to which the first neural network setting information selected for AI downscaling among the plurality of first neural network setting information is applied is 1 / d (where d is a rational number greater than 1), and the scaling factor of the second neural network (1440) to which the second neural network setting information corresponding to the first neural network setting information selected for AI downscaling is applied may be d (where d is a rational number greater than 1).

[0342] The AI ​​scaling unit (1312) can AI upscale the second image (135) n times (n is a natural number greater than or equal to 1) by considering the scaling ratio of the second neural network (1440) to which the second neural network setting information is applied, the resolution of the second image (135), and the supported resolution of the display device. When the AI ​​scaling unit (1312) AI upscales the second image (135) two or more times, it may be useful when the supported resolution of the display device is greater than the resolution of the third image (145) obtained through one AI downscale of the second image (135).

[0343] For example, if the scaling factor of the second neural network (1440) to which the 'B2' neural network setting information is applied is 2, the supported resolution of the display device is 16k, and the resolution of the second image (135) input to the second neural network (1440) is 2k, then the AI ​​scaling unit (1312) can obtain a third image (145) with a resolution of 16k by AI upscaling the second image (135) three times through the second neural network (1440) to which the 'B2' neural network setting information is applied.

[0344] Information regarding the supported resolution of the display device may be included in the performance information of the display device received by the AI ​​encoding unit (1310). The supported resolution of the display device may indicate at least one of the resolutions that the display can play. For example, the supported resolution of the display device may indicate the largest resolution that the display can play.

[0345] FIG. 19 is a diagram illustrating first neural network configuration information applicable to a first neural network (1420) for AI downscaling, and a plurality of second neural network configuration information applicable to a second neural network (1440) for AI upscaling.

[0346] As illustrated in FIG. 19, the AI ​​setting unit (1318) can store one 'A1' neural network setting information that can be set in the first neural network (1420) and a plurality of second neural network setting information that can be set in the second neural network (1440), namely, 'A2' neural network setting information, 'B2' neural network setting information, 'C2' neural network setting information, and 'D2' neural network setting information. FIG. 19 illustrates four second neural network setting information applicable to the second neural network (1440), but the number of second neural network setting information applicable to the second neural network (1440) may vary.

[0347] Among the multiple second neural network configuration information, the 'A2' neural network configuration information may correspond to the 'A1' neural network configuration information. The 'A1' neural network configuration information and the 'A2' neural network configuration information can be obtained through the combined training of the first neural network (1420) and the second neural network (1440), as described in relation to FIG. 17.

[0348] The scaling factor of the 'A1' neural network configuration information may be 1 / 4, and the scaling factors of the 'A2' neural network configuration information, 'B2' neural network configuration information, 'C2' neural network configuration information, and 'D2' neural network configuration information, respectively, may be 4, 2, 3, and 8. Since the 'A1' neural network configuration information and the 'A2' neural network configuration information are obtained through the combined training of the first neural network (1420) and the second neural network (1440), the scaling factor of the 'A1' neural network configuration information and the scaling factor of the 'A2' neural network configuration information have an inverse relationship.

[0349] The 'B2' neural network configuration information, the 'C2' neural network configuration information, and the 'D2' neural network configuration information can be obtained through individual training of the second neural network (1440) while the 'A1' neural network configuration information is set in the first neural network (1420) and the 'A2' neural network configuration information is set in the second neural network (1440). The individual training process of the second neural network (1440) will be described later with reference to FIGS. 21 and FIGS. 22.

[0350] When AI upscaling for the second image (135) is required, the AI ​​setting unit (1318) selects a second neural network setting information to be applied to the second neural network (1440) from among a plurality of second neural network setting information. In one embodiment, the AI ​​setting unit (1318) may select a second neural network setting information to be applied to the second neural network (1440) based on the supported resolution of the display device.

[0351] For example, if the resolution of the original image (105) is 8k, the resolution of the first image (115) obtained by the first neural network (1420) to which the 'A1' neural network setting information is applied becomes 2k. Since the resolution of the second image (135) is the same as that of the first image (115), the resolution of the second image (135) also becomes 2k. If the supported resolution of the display device is 8k, the AI ​​setting unit (1318) selects the 'A2' neural network setting information having a scaling factor of 4 and provides it to the AI ​​scaling unit (1312). Additionally, if the supported resolution of the display device is 4k, the AI ​​setting unit (1318) selects the 'B2' neural network setting information having a scaling factor of 2 and provides it to the AI ​​scaling unit (1312).

[0352] The AI ​​scaling unit (1312) can obtain a third image (145) that matches the supported resolution of the display device by AI upscaling the second image (135) using the second neural network setting information provided by the AI ​​setting unit (1318).

[0353] FIG. 20 is a diagram illustrating a plurality of first neural network configuration information applicable to a first neural network (1420) for AI downscaling, and a plurality of second neural network configuration information applicable to a second neural network (1440) for AI upscaling.

[0354] As illustrated in FIG. 20, the AI ​​setting unit (1318) can store a plurality of first neural network setting information that can be set in the first neural network (1420) and a plurality of second neural network setting information that can be set in the second neural network (1440).

[0355] Referring to FIG. 20, it can be seen that there are multiple second neural network configuration information corresponding to one first neural network configuration information. The multiple first neural network configuration information that can be configured in the first neural network (1420) may correspond to the multiple first neural network configuration information described above in relation to FIG. 18. In addition, there are multiple second neural network configuration information having different scaling ratios corresponding to each first neural network configuration information.

[0356] For example, as illustrated in FIG. 20, regarding the 'A1' neural network configuration information applicable to the first neural network (1420), there may be 'A2-1' neural network configuration information, 'A2-2' neural network configuration information, 'A2-3' neural network configuration information, and 'A2-4' neural network configuration information applicable to the second neural network (1440). The 'A1' neural network configuration information and the 'A2-1' neural network configuration information can be obtained through the combined training of the first neural network (1420) and the second neural network (1440), as described in relation to FIG. 17. And, the 'A2-2' neural network setting information, the 'A2-3' neural network setting information, and the 'A2-4' neural network setting information can be obtained through individual training of the second neural network (1440), which will be described with reference to FIGS. 21 and FIGS. 22, while the 'A1' neural network setting information is set in the first neural network (1420) and the 'A2-1' neural network setting information is set in the second neural network (1440).

[0357] Additionally, regarding the 'B1' neural network configuration information applicable to the first neural network (1420), there may be 'B2-1' neural network configuration information, 'B2-2' neural network configuration information, 'B2-3' neural network configuration information, and 'B2-4' neural network configuration information applicable to the second neural network (1440). The 'B1' neural network configuration information and the 'B2-1' neural network configuration information can be obtained through the combined training of the first neural network (1420) and the second neural network (1440), as described in relation to FIG. 17. And, the 'B2-2' neural network setting information, the 'B2-3' neural network setting information, and the 'B2-4' neural network setting information can be obtained through individual training of the second neural network (1440), which will be described with reference to FIGS. 21 and FIGS. 22, while the 'B1' neural network setting information is set in the first neural network (1420) and the 'B2-1' neural network setting information is set in the second neural network (1440).

[0358] The AI ​​setting unit (1318) selects a first neural network setting information to be applied to the first neural network (1420) from among a plurality of first neural network setting information and provides it to the AI ​​scale unit (1312). The method by which the AI ​​setting unit (1318) selects the first neural network setting information has been described in relation to FIG. 18, so a detailed description is omitted.

[0359] If AI upscaling of the second image (135) is required, the AI ​​setting unit (1318) can select a second neural network setting information to be applied to the second neural network (1440) according to the supported resolution of the display device from among a plurality of second neural network setting information having different scaling ratios corresponding to the first neural network setting information selected for AI downscaling, and provide it to the AI ​​scaling unit (1312).

[0360] For example, if the resolution of the original image (105) is 8k and the ‘A1’ neural network setting information having a scaling factor of 1 / 4 is selected for AI downscaling, the resolution of the first image (115) becomes 2k. If the AI ​​setting unit (1318) has a supported resolution of 4k for the display device, the ‘A2-2’ neural network setting information capable of outputting a third image (145) of 4k can be selected for AI upscaling.

[0361] As another example, if the resolution of the original image (105) is 8k and the ‘A1’ neural network setting information with a scaling factor of 1 / 4 is selected for AI downscaling, the resolution of the first image (115) becomes 2k. The AI ​​setting unit (1318) can select the ‘A2-4’ neural network setting information for AI upscaling, which can output a third image (145) of 16k if the supported resolution of the display device is 16k.

[0362] Below, individual training of the second neural network (1440) will be described with reference to FIGS. 21 and FIGS. 22.

[0363] FIG. 21 is an exemplary drawing showing a second neural network (1440) according to one embodiment.

[0364] The second neural network (1440) illustrated in FIG. 21 may include a first convolution layer (2110), a first activation layer (2120), a scaler (2130), a second convolution layer (2140), a second activation layer (2150), and a third convolution layer (2160).

[0365] Comparing the second neural network (1440) shown in FIG. 21 with the second neural network (300) shown in FIG. 3, the second neural network (1440) shown in FIG. 21 includes more scalers (2130) than the second neural network (300) shown in FIG. 3. The scalers (2130) are used to obtain second neural network configuration information having various scaling ratios. In other words, the scaling ratio of the scalers (2130) must be adjusted to obtain second neural network configuration information having various scaling ratios as described in relation to FIG. 19 and FIG. 20.

[0366] The scaler (2130) may be a bilinear scaler, a bicubic scaler, a lanczos scaler, or a stair-step scaler. Depending on the embodiment, the scaler (2130) may be implemented as a convolution layer.

[0367] The first convolution layer (2110), the first activation layer (2120), the second convolution layer (2140), the second activation layer (2150), and the third convolution layer (2160) are the same as those described in relation to FIG. 3, so a detailed description is omitted.

[0368] FIG. 22 is a diagram illustrating a method for individually training a second neural network (1440) after the training in conjunction with the first neural network (1420) has been completed.

[0369] In FIG. 22, the original training image (2201) is the image to be AI downscaled, and the first training image (2202) is the image AI downscaled from the original training image (2201). Additionally, the second training image (2203) is the image AI upscaled from the first training image (2202).

[0370] The original training image (2201) includes a still image or a video composed of multiple frames. In one embodiment, the original training image (2201) may include a luminance image extracted from a still image or a video composed of multiple frames. Additionally, in one embodiment, the original training image (2201) may include a patch image extracted from a still image or a video composed of multiple frames. If the original training image (2201) consists of multiple frames, the first training image (2202) and the second training image (2203) are also composed of multiple frames.

[0371] When multiple frames of the original training video (2201) are sequentially input into the first neural network (1420), multiple frames of the first training video (2202) and the second training video (2203) can be sequentially acquired through the first neural network (1420) and the second neural network (1440).

[0372] In FIG. 22, the first neural network (1420) and the second neural network (1440) are set with the first neural network setting information and the second neural network setting information obtained through linked training. Then, to adjust the scaling factor of the second neural network (1440), the setting value of the scaler (2130) shown in FIG. 21 is adjusted. For example, if one wants to obtain the second neural network setting information with a scaling factor of 2, the scaling factor of the scaler (2130) can be set to 2, and if one wants to obtain the second neural network setting information with a scaling factor of 3, the scaling factor of the scaler (2130) can be set to 3.

[0373] For individual training of the second neural network (1440), the original training video (2201) is input into the first neural network (1420). The first neural network (1420) operates according to the first neural network configuration information obtained through the combined training of the first neural network (1420) and the second neural network (1440).

[0374] The original training image (2201) input to the first neural network (1420) is AI downscaled and output as the first training image (2202), and the first training image (2202) is input to the second neural network (1440). The second training image (2203) is output as a result of AI upscaling the first training image (2202).

[0375] Referring to FIG. 22, a first training image (2202) is input to a second neural network (1440), and according to an embodiment, a training image obtained through the encoding and decoding process of the first training image (2202) may be input to the second neural network (1440).

[0376] Apart from the original training image (2201) being input into the first neural network (1420), the original training image (2201) is processed by a scaler (2210). Here, the scaler (2210) may include at least one of a bilinear scaler, a bicubic scaler, a lanczos scaler, and a stair step scaler.

[0377] The scaler (2210) scales the original training image (2201) according to the ratio between the resolution of the second training image (2203) to be output from the second neural network (1440) and the resolution of the original training image (2201) according to the scaling ratio of the second neural network (1440). For example, if the resolution of the second training image (2203) is output as 4k and the resolution of the original training image (2201) is 8k through the adjustment of the scaler (2130) included in the second neural network (1440), the scaler (2210) can reduce the resolution of the original training image (2201) by half. The reason for matching the resolutions of the original training image (2201) and the second training image (2203) is to calculate loss information through the comparison of the two images (2201, 2203).

[0378] Quality loss information (2230) corresponding to the difference between the second training image (2203) output from the second neural network (1440) and the original training image (2201) scaled by the scaler (2210) is calculated. The quality loss information (2230) may include at least one of an L1-norm value, an L2-norm value, a SSIM (Structural Similarity) value, a PSNR-HVS (Peak Signal-To-Noise Ratio-Human Vision System) value, an MS-SSIM (Multiscale SSIM) value, a VIF (Variance Inflation Factor) value, and a VMAF (Video Multimethod Assessment Fusion) value for the difference between the scaled original training image (2201) and the second training image (2203).

[0379] Quality loss information (2230) is used only for training the second neural network (1440). Neural network configuration information set in the second neural network (1440) is updated so that quality loss information (2230) is reduced or minimized. As shown in FIG. 22, quality loss information (2230) is not used for training the first neural network (1420).

[0380] By individually training the second neural network (1440) while adjusting the scaling factor of the scaler (2130) included in the second neural network (1440), second neural network configuration information corresponding to various scaling factors can be obtained. Specifically, if the scaling factor of the scaler (2130) included in the second neural network (1440) is set to 2 and the second neural network (1440) is individually trained, second neural network configuration information with a scaling factor of 2 is obtained. And, if the scaling factor of the scaler (2130) included in the second neural network (1440) is set to 3 and the second neural network (1440) is individually trained, second neural network configuration information with a scaling factor of 3 is obtained.

[0381] FIG. 23 is a diagram illustrating the individual training process of a second neural network (1440) by a training device (2300).

[0382] Individual training of the second neural network (1440) can be performed by a training device (2300). The training device (2300) includes the first neural network (1420) and the second neural network (1440). The training device (2300) may be, for example, an image providing device (1300) or a separate server. The neural network configuration information of the second neural network (1440) obtained as a result of training is stored in the image providing device (1300).

[0383] Before individual training of the second neural network (1440), the first neural network setting information and the second neural network setting information obtained through the combined training of the first neural network (1420) and the second neural network (1440) are set in the first neural network (1420) and the second neural network (1440), respectively. Then, the scaling ratio of the scaler (2130) included in the second neural network (1440) is adjusted according to the scaling ratio of the second neural network setting information to be obtained.

[0384] Referring to FIG. 23, the training device (2300) inputs the original training image (2201) into the first neural network (1420) (S2310). The original training image (2201) may include at least one frame constituting a still image or a video.

[0385] The first neural network (1420) processes the original training image (2201) according to the set first neural network setting information and outputs the first training image (2202) that has been AI downscaled from the original training image (2201) (S2320). FIG. 23 is illustrated as the first training image (2202) output from the first neural network (1420) being directly input into the second neural network (1440), but the first training image (2202) output from the first neural network (1420) may be input into the second neural network (1440) by the training device (2300). Additionally, the training device (2300) may input the first training image (2202) into the second neural network (1440) after encoding and decoding it with a predetermined codec.

[0386] The second neural network (1440) processes the first training image (2202) or the training image obtained through encoding and decoding of the first training image (2202) and outputs the second training image (2203) (S2330).

[0387] The training device (2300) scales the original training image (2201) (S2340). The training device (2300) scales the original training image (2201) to have the same resolution as the second training image (2203) output from the second neural network (1440).

[0388] The training device (2300) compares the scaled original training image (2201) and the second training image (2203) to produce quality loss information (2230) (S2350).

[0389] The second neural network (1440) updates the neural network configuration information through a back propagation process based on the final loss information determined from the quality loss information (2230) (S2360).

[0390] Afterward, the training device (2300), the first neural network (1420), and the second neural network (1440) repeat the process S2310 to S2360 until the final loss information is minimized. At this time, during each iteration, the second neural network (1440) operates according to the neural network setting information updated in the previous process.

[0391] FIG. 24 is a flowchart for explaining a method of providing an image by an image providing device (1300) according to one embodiment.

[0392] In step S2410, the image providing device (1300) obtains a first image (115) by performing AI downscaling on the original image (105). The image providing device (1300) can prevent the storage capacity of the image providing device (1300) from becoming saturated by deleting from the image providing device (1300) the original frames for which AI downscaling has been completed among the original frames constituting the original image (105).

[0393] In step S2420, the image providing device (1300) encodes the first image (115) to obtain the first image data.

[0394] In step S2430, when the image providing device (1300) receives a request for an image from a display device or a user, it determines whether the display device supports an AI upscaling function. The image providing device (1300) can determine whether the display device supports an AI upscaling function from performance information received from the display device or the user.

[0395] If the display device does not support the AI ​​upscale function, in step S2440, the image providing device (1300) decodes the first image data to obtain the second image (135).

[0396] In step S2450, the image providing device (1300) AI upscales the second image (135) to obtain the third image (145).

[0397] In step S2460, the image providing device (1300) encodes the third image (145) to obtain second image data. When the encoding of the third frame constituting the GOP among the third frames constituting the third image (145) is completed, the image providing device (1300) can delete the corresponding third frames to prevent the storage capacity of the image providing device (1300) from becoming saturated.

[0398] In step S2470, the image providing device (1300) provides the second image data to the display. The display device can decode the second image data to restore the third image (145) and play the restored third image (145).

[0399] In step S2430, if the display device does not support the AI ​​upscaling function, the video providing device (1300) provides AI encoded data containing the first video data to the display device. The display device obtains a third video (145) through AI decoding of the AI ​​encoded data and can play the obtained third video (145).

[0400] Meanwhile, the structure of the display device supporting the AI ​​upscaling function may be the same as the AI ​​decoding device (200) shown in FIG. 2.

[0401] That is, the receiving unit (210) transmits the AI ​​encoded data received from the video providing device (1300) to the parsing unit (232). The parsing unit (232) transmits the first video data among the AI ​​encoded data to the first decoding unit (234) and transmits the AI ​​data to the AI ​​setting unit (238).

[0402] The first decoding unit (234) decodes the first image data to obtain the second image (135) and transmits the second image (135) to the AI ​​upscaler (236). The first decoding unit (234) may provide information related to decoding to the AI ​​setting unit (238).

[0403] The AI ​​setting unit (238) selects second neural network setting information to be used for AI upscaling based on AI data.

[0404] The AI ​​setting unit (238) can store one or more second neural network setting information for AI upscaling. When one second neural network setting information is stored in the AI ​​setting unit (238), the AI ​​setting unit (238) transmits one second neural network setting information to the AI ​​upscaling unit (236).

[0405] When multiple second neural network setting information is stored in the AI ​​setting unit (238), the AI ​​setting unit (238) selects a second neural network setting information to be set in the second neural network (300 or 1440) from among the multiple second neural network setting information based on AI data, and transmits the selected second neural network setting information to the AI ​​upscale unit (236). The AI ​​setting unit (238) can verify the first neural network setting information set in the first neural network (1420) from the AI ​​data and select a second neural network setting information corresponding to the verified first neural network setting information. As for the method of selecting the second neural network setting information from the AI ​​data, it has been explained above in relation to the AI ​​decoding device (200), so a detailed explanation is omitted here.

[0406] According to an embodiment, the AI ​​setting unit (238) may select a second neural network setting information corresponding to the supported resolution of the display device among a plurality of second neural network setting information having different scaling ratios and transmit it to the AI ​​upscale unit (236).

[0407] The AI ​​upscaler (236) sets the second neural network (300 or 1440) with the second neural network setting information received from the AI ​​setting unit (238), and obtains the third image (145) by AI upscaler the second image (135) through the second neural network (300 or 1440). According to an embodiment, the AI ​​upscaler (236) can obtain the third image (145) by AI upscaler the second image (135) n times (n is a natural number greater than or equal to 1) according to the supported resolution of the display device.

[0408] The display device can play the third image (145), or if necessary, post-process the third image (145) and play the post-processed third image (145).

[0409] A display device that does not support an AI upscaling function may not include the AI ​​upscaling unit (236) and the AI ​​setting unit (238) among the components included in the AI ​​decoding device (200) shown in FIG. 2. Accordingly, the display device decodes the second image data received from the image providing device (1300) to obtain a third image (145). The display device may play the restored third image (145), or post-process the restored third image (145) and play the post-processed third image (145).

[0410] Meanwhile, the embodiments of the present disclosure described above can be written as a program or instruction that can be executed on a computer, and the written program or instruction can be stored on a storage medium.

[0411] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory storage medium' simply means that it is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, a 'non-transitory storage medium' may include a buffer in which data is stored temporarily.

[0412] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0413] Meanwhile, the model related to the neural network described above may be implemented as a software module. When implemented as a software module (e.g., a program module containing instructions), the neural network model may be stored on a computer-readable recording medium.

[0414] Additionally, the neural network model may be integrated in the form of a hardware chip and become part of the aforementioned AI decoder (200), AI encoder (700), image provider (1300), and display device. For example, the neural network model may be manufactured in the form of a dedicated hardware chip for artificial intelligence, or it may be manufactured as part of an existing general-purpose processor (e.g., CPU or application processor) or a graphics-dedicated processor (e.g., GPU).

[0415] Additionally, the neural network model may be provided in the form of downloadable software. The computer program product may include a product in the form of a software program (e.g., a downloadable application) that is distributed electronically through a manufacturer or an electronic market. For electronic distribution, at least a portion of the software program may be stored on a storage medium or temporarily created. In this case, the storage medium may be a server of the manufacturer or electronic market, or a storage medium of a relay server.

[0416] Although the technical concept of the present disclosure has been described in detail with reference to preferred embodiments, the technical concept of the present disclosure is not limited to the above embodiments, and various modifications and changes can be made by those skilled in the art within the scope of the technical concept of the present disclosure.

Claims

Claim 1 An electronic device that provides an image using artificial intelligence (AI), comprising a processor that executes one or more instructions stored in the electronic device, wherein, as the one or more instructions are executed by the processor, the electronic device obtains a first image by performing AI downscaling on an original image through a neural network for downscaling, obtains first image data by encoding the first image, and if the display device does not support an AI upscaling function, obtains a second image by decoding the first image data, obtains a third image by performing AI upscaling on the second image through a neural network for upscaling, and provides the second image data obtained through encoding the third image to the display device, wherein if the display device supports the AI ​​upscaling function, the first image data is provided to the display device. Claim 2 In claim 1, the electronic device further comprises a camera for acquiring the original image. Claim 3 delete Claim 4 The electronic device according to claim 1, wherein the electronic device obtains the first image by AI downscaling the original image through a downscaling neural network to which predetermined first neural network setting information is applied, selects a second neural network setting information corresponding to the supported resolution of the display device among a plurality of second neural network setting information, and obtains the third image by AI upscaling the second image through an upscaling neural network to which the selected second neural network setting information is applied. Claim 5 An electronic device according to claim 4, wherein the first neural network setting information that is predetermined and the second neural network setting information corresponding to the first neural network setting information among the plurality of second neural network setting information are obtained through joint training of the downscale neural network and the upscale neural network, and the second neural network setting information other than the second neural network setting information corresponding to the first neural network setting information among the plurality of second neural network setting information is obtained through individual training of the upscale neural network after the joint training. Claim 6 The electronic device according to claim 1, wherein the electronic device obtains the first image by AI downscaling the original image through a downscaling neural network to which predetermined first neural network setting information is applied, and obtains the third image by AI upscaling the second image n times (n is a natural number) through the upscaling neural network to which predetermined second neural network setting information is applied, taking into account the scaling ratio of the upscaling neural network to which predetermined second neural network setting information is applied and the supported resolution of the display device. Claim 7 The electronic device according to claim 1, wherein the electronic device selects a first neural network setting information for AI downscaling among a plurality of first neural network setting information, performs AI downscaling on the original image through a downscaling neural network to which the selected first neural network setting information is applied to obtain the first image, selects a second neural network setting information corresponding to the selected first neural network setting information among a plurality of second neural network setting information, and obtains the third image by performing AI upscaling on the second image n times (n is a natural number) through the upscaling neural network, taking into account the scaling ratio of the upscaling neural network to which the selected second neural network setting information is applied and the supported resolution of the display device. Claim 8 In claim 1, the original frames constituting the original image are sequentially AI downscaled through the downscale neural network, thereby obtaining the first frames constituting the first image, and the electronic device sequentially deletes the original frames for which AI downscaled is completed. Claim 9 An electronic device according to claim 1, wherein third frames constituting the third image are obtained as the second frames constituting the second image are sequentially AI upscaled through the upscale neural network, and wherein, when third frames constituting a group of pictures (GOP) are obtained through the AI ​​upscale, the electronic device encodes the third frames constituting the GOP and deletes the encoded third frames constituting the GOP. Claim 10 In claim 1, the electronic device receives performance information from the display device and checks whether the AI ​​upscale function is supported and the supported resolution of the display device from the performance information. Claim 11 delete Claim 12 A method for providing an image using an electronic device utilizing artificial intelligence (AI), comprising: a step of obtaining a first image by performing AI downscaling on an original image through a neural network for downscaling; a step of obtaining first image data by encoding the first image; a step of obtaining a second image by decoding the first image data if the display device does not support an AI upscaling function; a step of obtaining a third image by AI upscaling the second image through a neural network for upscaling; and a step of providing the second image data obtained through encoding the third image to the display device, wherein if the display device supports the AI ​​upscaling function, the first image data is provided to the display device. Claim 13 A computer-readable recording medium having a program recorded thereon for performing the method of paragraph 12 on a computer. Claim 14 delete