Image processing apparatus and method of operating the same
By using a single deep neural network in an image processing device and determining the processing purpose in combination with a classifier, the problem of difficulty in processing images according to different purposes in the prior art is solved, and flexible and efficient image processing is achieved.
Patent Information
- Application Number
- CN201980051536.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-08-02
- Filing Date
- 2019-07-29
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2039-07-29
AI Technical Summary
The prior art is difficult to process input images according to different purposes through a single deep neural network, which lacks flexibility and efficiency.
An image processing device is designed to process input images based on different purposes using a single deep neural network (DNN), and to determine the processing purpose by adjusting the image processing level and using classifiers.
It realizes the rapid and flexible processing of images according to different purposes, and improves the efficiency and accuracy of image processing.
Smart Images

Figure CN112534443B_ABST
Abstract
Description
Technical Field
[0001] Various embodiments relate to an image processing apparatus and an operating method thereof for processing an image by using a deep neural network, and more particularly, to an image processing apparatus and an operating method thereof for processing an input image according to different purposes by using a single network. Background Art
[0002] As data traffic has exponentially increased with the development of computer technology, artificial intelligence (AI) has become a major trend driving future innovation. Since AI mimics how humans think, it can be infinitely applied to essentially all industries. Representative technologies of AI may include pattern recognition, machine learning, expert systems, neural networks, natural language processing, etc.
[0003] A neural network mimics the characteristics of human biological neurons by using mathematical expressions and uses an algorithm that mimics the human ability called learning. By using such an algorithm, a neural network can generate a mapping between input data and output data, and the ability to generate the mapping can be represented as the learning ability of the neural network. In addition, a neural network has a generalization ability, through which corrected output data for input data that has not been used for training can be generated based on the trained result. Summary of the Invention
[0004] Technical Problem
[0005] Various embodiments may provide an image processing apparatus and an operating method thereof that can process an input image according to various purposes by using a single deep neural network.
[0006] Advantageous Effects of the Disclosure
[0007] An image processing apparatus according to an embodiment can process an image according to various purposes by using a single deep neural network.
[0008] An image processing apparatus according to an embodiment can adjust an image processing level according to various purposes. Brief Description of the Drawings
[0009] Figure 1 Shows a first deep neural network, a second deep neural network, and a third deep neural network trained using different training data sets.
[0010] Figure 2 Shows a process of processing an input image according to different purposes performed by an image processing apparatus according to an embodiment.
[0011] Figure 3It is a flowchart of an operation method of an image processing apparatus according to an embodiment.
[0012] Figure 4 It is a reference diagram for describing a method of processing an image performed by an image processing apparatus according to an embodiment.
[0013] Figure 5 It is a reference diagram for describing a method of generating an input image and a classifier performed by an image processing apparatus according to an embodiment.
[0014] Figure 6 It is for describing the process of performing a convolution operation by Figure 4 the first convolutional layer.
[0015] Figure 7 It shows Figure 4 the input data and output data of the first convolutional layer according to an embodiment.
[0016] Figure 8 It shows Figure 1 the input data and output data of the first convolutional layer in the third deep neural network.
[0017] Figure 9 It shows the input data and output data of a deep neural network according to an embodiment.
[0018] Figure 10 It shows the input data and output data of a deep neural network according to an embodiment.
[0019] Figure 11 It is a reference diagram for describing a method of generating a classifier according to an embodiment.
[0020] Figure 12 It shows an identification image generated based on a first image according to an embodiment.
[0021] Figure 13 It is a reference diagram for describing a method of training a deep neural network according to an embodiment.
[0022] Figure 14 It is a reference diagram for describing a deep neural network configured to perform image processing according to multiple purposes according to an embodiment.
[0023] Figure 15 It is a reference diagram for describing a method of generating a training data set for training a deep neural network according to an embodiment.
[0024] Figure 16 It is a block diagram of an image processing apparatus according to an embodiment.
[0025] Figure 17 It is a block diagram of a processor according to an embodiment.
[0026] Figure 18 It is a block diagram of an image processor according to an embodiment. Detailed implementation
[0027] According to an embodiment of the present disclosure, there is provided an image processing apparatus, which includes: a memory that stores one or more instructions; and a processor configured to execute the one or more instructions stored in the memory, wherein the processor is further configured to: obtain a first image and a classifier indicating the purpose of image processing, and process the first image by using a deep neural network (DNN), and the DNN processes an input image according to the purpose indicated by the classifier, wherein the DNN processes the input image according to different purposes.
[0028] The DNN may include N convolutional layers, and the processor may further be configured to generate an input image based on the first image, extract feature information by performing a convolution operation, in which one or more kernels are applied to the input image and the classifier in the N convolutional layers, and generate a second image based on the extracted feature information.
[0029] The processor may further be configured to: convert the R, G, and B channels included in the first image into the Y channel, U channel, and V channel in the YUV mode, and determine the image of the Y channel among the Y channel, U channel, and V channel as the input image.
[0030] The processor may further be configured to: generate a second image based on the image output by processing the image of the Y channel in the DNN and based on the images of the U and V channels among the Y channel, U channel, and V channel.
[0031] Pixels included in the classifier may have at least one of a first value and a second value greater than the first value, the first value may indicate a first purpose, and the second value may indicate a second purpose.
[0032] The processor may further be configured to: when all pixels included in the classifier have the first value, process the first image according to the first purpose, and when all pixels included in the classifier have the second value, process the first image according to the second purpose.
[0033] When pixels in a first region included in the classifier have the first value and pixels in a second region included in the classifier have the second value, the third region corresponding to the first region in the first image may be processed according to the first purpose, and the fourth region corresponding to the second region in the first image may be processed according to the second purpose.
[0034] The processor may also be configured to process a first image according to an image processing level for a first purpose and an image processing level for a second purpose, where the first purpose and the second purpose are determined based on values of pixels included in a classifier, a first value, and a second value.
[0035] The processor may also be configured to generate a classifier based on characteristics of the first image.
[0036] The processor may also be configured to generate a mapping image indicating text and edges included in the first image, and determine values of pixels included in the classifier based on the mapping image.
[0037] The DNN may be trained by a first training dataset including first image data, a first classifier having a first value as a pixel value, and first label data obtained by processing the first image data according to the first purpose, and may be trained by a second training dataset including second image data, a second classifier having a second value as a pixel value, and second label data obtained by processing the second image data according to the second purpose.
[0038] The processor may also be configured to adjust weights of one or more kernels included in the DNN to reduce a difference between the first label data and image data output when the first image data and the first classifier are input to the DNN; and adjust weights of one or more kernels included in the DNN to reduce a difference between the second label data and image data output when the second image data and the second classifier are input to the DNN.
[0039] According to another embodiment of the present disclosure, there is provided a method of operating an image processing device, the method including: obtaining a first image and a classifier indicating an image processing purpose; and processing the first image by using a deep neural network (DNN) according to the purpose indicated by the classifier, where the DNN processes multiple images according to different purposes.
[0040] According to another embodiment of the present disclosure, there is provided a computer program product including one or more computer-readable recording media storing a program configured to perform the following operations: obtaining a first image and a classifier indicating an image processing purpose; and processing the first image according to the purpose indicated by the classifier by using a deep neural network (DNN) configured to process multiple images according to different purposes.
[0041] Disclosed embodiments
[0042] Terms used in this specification will be schematically described and then the present invention will be described in detail.
[0043] The terms used in the present invention are those general terms currently widely used in the art, but these terms may vary according to the intention of those of ordinary skill in the art, precedent, or new technologies in the art. In addition, terms designated by the applicant may be selected, and in such cases, their detailed meanings will be described in the detailed description. Therefore, the terms used in the present invention should not be construed as simple names, but should be interpreted based on the meaning of the terms and the overall description.
[0044] Throughout the specification, it should also be understood that when a component "comprises" an element, unless otherwise described to the contrary, it should be understood that the component does not exclude another element, but may further include another element. In addition, terms such as "unit", "module", etc. refer to a unit that performs at least one function or operation, and these units can be implemented as hardware or software, or implemented as a combination of hardware and software.
[0045] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings so that those of ordinary skill in the art to which the present invention pertains can easily implement the embodiments. However, the present invention can be implemented in many different forms and should not be construed as limited to the embodiments herein. In the drawings, parts irrelevant to the description are omitted to clearly describe the present invention, and the same reference numerals throughout the specification denote the same elements.
[0046] Figure 1 The first deep neural network (DNN1), the second deep neural network (DNN2), and the third deep neural network (DNN3) trained using different training data sets are shown.
[0047] Referring to Figure 1 , DNN1 can be trained by the first training data set 11. Here, the first training data set 11 can be a training data set for enhancing the details of images (e.g., for a first purpose). For example, the first training data set 11 can include multiple pieces of first image data D1 and multiple pieces of first label data L1. As the high-resolution image data corresponding to each piece of first image data D1, each piece of first label data L1 can be the image data obtained by enhancing the texture representation when each piece of first image data D1 is converted into a high-resolution image.
[0048] DNN2 can be trained by the second training dataset 21. Here, the second training dataset 21 can be a training dataset for enhancing text or edges included in an image (second purpose). For example, the second training dataset 21 can include second image data D2 and second label data L2. As high-resolution image data corresponding to the second image data D2, each piece of the second label data L2 can be image data obtained by reducing jaggedness and the like that appear around text or edges (enhanced text or edge representation) when the second image data D2 is converted into a high-resolution image.
[0049] In addition, DNN3 can be trained by the first training dataset 11 and the second training dataset 21.
[0050] As Figure 1 shown, when the first image 10 is input into DNN1, DNN1 can output a first output image 15, and the first output image 15 can be a high-resolution image obtained by enhancing the texture representation of the first image 10. In addition, when the second image 20 is input into DNN2, DNN2 can output a second output image 25, and the second output image 25 can be a high-resolution image obtained by enhancing the text or edge representation of the second image 20.
[0051] On the other hand, when the first image 10 is input into DNN3, DNN3 can output a third output image 35, and in this case, the degree of detail enhancement (texture representation enhancement) of the third output image 35 is less than that of the first output image 15. In addition, when the second image 20 is input into DNN3, DNN3 can output a fourth output image 45, and in this case, the degree of enhancement of the text or edge representation of the fourth output image 45 is less than that of the second output image 25.
[0052] DNN3 trained by the first training dataset 11 for the first training purpose (detail enhancement) and the second training dataset 21 for the second training purpose (text or edge enhancement) does not have the performance of DNN1 trained by the first training dataset 11 and the performance of DNN2 trained by the second training dataset 21.
[0053] Figure 2 The process of processing an input image according to different purposes by an image processing device according to an embodiment is shown.
[0054] The image processing apparatus 100 according to an embodiment may process an input image by using a deep neural network (DNN) 150. The DNN 150 according to an embodiment may be trained by a first training data set 210 for a first purpose (e.g., detail enhancement) and a second training data set 220 for a second purpose (e.g., text or edge enhancement). Here, the first training data set 210 may include first image data D1, a first classifier C1 indicating the first purpose, and first label data L1. Each piece of the first image data D1 and each piece of the first label data L1 are the same as the corresponding pieces of the first image data D1 and the first label data L1 described with reference to Figure 1 When each piece of the first image data D1 and the first classifier C1 are input, the DNN 150 may be trained to output each piece of the first label data L1. In addition, the second training data set 220 may include second image data D2, a second classifier C2 indicating the second purpose, and second label data L2. The second image data D2 and the second label data L2 are the same as the corresponding second image data D2 and the second label data L2 described with reference to Figure 1 When the second image data D2 and the second classifier C2 are input, the DNN 150 may be trained to output the second label data L2.
[0055] The image processing apparatus 100 according to an embodiment may input a first image 10 and a first classifier C1 indicating a first purpose (detail enhancement) to the DNN 150, and the DNN 150 may output a first output image 230 by processing the first image 10 according to the first purpose (e.g., detail enhancement) indicated by the first classifier C1. Here, the degree of detail enhancement in the first output image 230 may be the same as or similar to the degree of detail enhancement of the first output image 15 of the DNN1 described with reference to Figure 1
[0056] In addition, the image processing apparatus 100 may input a second image 20 and a second classifier C2 indicating a second purpose (e.g., text or edge enhancement) to the DNN 150, and the DNN 150 may output a second output image 240 by processing the second image 20 according to the second purpose (e.g., text or edge enhancement) indicated by the second classifier C2. Here, the degree of text or edge enhancement in the second output image 240 may be the same as or similar to the degree of the second output image 25 of the DNN2 described with reference to Figure 1
[0057] The image processing apparatus 100 according to an embodiment may process a plurality of images for different purposes by using a single DNN trained with image data, a classifier indicating a processing purpose of the image data, and label data corresponding to the image data.
[0058] Figure 3 It is a flowchart of an operation method of an image processing apparatus according to an embodiment.
[0059] Referring to Figure 3 , in operation S310, the image processing apparatus 100 according to an embodiment may obtain a first image and a classifier indicating an image processing purpose.
[0060] The image processing apparatus 100 may generate an input image to be input to the DNN and a classifier based on the first image. For example, the image processing apparatus 100 may convert the R channel, G channel, and B channel included in the first image into the Y channel, U channel, and V channel by color space conversion, and determine the image of the Y channel among the Y channel, U channel, and V channel as the input image. When attempting to perform image processing on the first image according to the first purpose, the image processing apparatus 100 may determine the pixel value of the classifier as a first value corresponding to the first purpose, and when attempting to perform image processing on the first image according to the second purpose, determine the pixel value of the classifier as a second value corresponding to the second purpose. Alternatively, the image processing apparatus 100 may extract edges and text from the first image, and generate a classifier based on the extracted edges and text.
[0061] In operation S320, the image processing apparatus 100 according to an embodiment may process the first image by using the DNN according to the purpose indicated by the classifier.
[0062] The DNN may include N convolutional layers, and the image processing apparatus 100 may extract feature information (feature maps) by performing a convolution operation, in which one or more kernels are applied to the input image and the identity image in each of the N convolutional layers, and based on the extracted feature information, process the first image according to the purpose indicated by the classifier.
[0063] The image processing apparatus 100 may determine an image processing level according to the first purpose and an image processing level according to the second purpose based on the value of the pixel included in the classifier, the first value corresponding to the first purpose, and the second value corresponding to the second purpose, and process the first image according to the determined levels.
[0064] On the other hand, the image processing apparatus 100 may train the DNN by using a plurality of training data sets having different purposes. For example, the image processing apparatus 100 may train the DNN by using a first training data set and a second training data set. The first training data set includes each piece of image data, a first classifier, and each piece of first label data obtained by processing each piece of image data according to the first purpose. The second training data set includes each piece of second image data, a second classifier, and each piece of second label data obtained by processing each piece of second image data according to the second purpose.
[0065] The image processing device 100 can adjust the weights of one or more kernels included in the DNN to reduce the difference between each piece of first label data and the image data output when the first pieces of image data and the first classifier are input into the DNN. In addition, the image processing device 100 can adjust the weights of one or more kernels included in the DNN to reduce the difference between each piece of second label data and the image data output when the second image data and each piece of second classifier are input into the DNN.
[0066] Figure 4 FIG. is a reference diagram for describing a method of processing an image performed by an image processing device according to an embodiment.
[0067] Referring to Figure 4 , the image processing device 100 according to an embodiment can receive a first image 410. The image processing device 100 can input the first image 410 into the DNN 150, or generate an input image 420 to be input into the DNN 150 based on the first image 410. In addition, the image processing device 100 according to an embodiment can generate a classifier 430 that indicates the purpose of performing image processing on the first image 410. The method of generating the input image 420 and the classifier 430 will be described in detail with reference to Figure 5 FIG.
[0068] Figure 5 FIG. is a reference diagram for describing a method of generating an input image and a classifier performed by an image processing device according to an embodiment.
[0069] Referring to Figure 5 , the first image 410 may include an R channel, a G channel, and a B channel (RGB 3ch). The image processing device 100 according to an embodiment can determine the R channel, the G channel, and the B channel (RGB 3ch) included in the first image 410 as the input image to be input into the DNN 150.
[0070] Alternatively, the image processing device 100 can convert the R channel, the G channel, and the B channel (RGB 3ch) included in the first image 410 into a Y channel, a U channel, and a V channel (YUV 3ch) through color space conversion. Here, the Y channel 420 indicates a luminance signal, the U channel indicates the difference between the luminance signal and the blue component, and the V channel indicates the difference between the luminance signal and the red component. The image processing device 100 can determine the image of the Y channel 420 among the converted Y channel, U channel, and V channel (YUV 3ch) as the input image to be input into the DNN 150. However, this embodiment is not limited thereto.
[0071] In addition, the image processing apparatus 100 according to an embodiment may generate a classifier 430 indicating the purpose of image processing. Here, the classifier 430 may have one pixel value or different pixel values for each region. In addition, the pixel values of the classifier 430 may be determined based on a user input or an image to be processed (e.g., the first image 410).
[0072] For example, when attempting to perform image processing on the entire first image 410 according to a first purpose, the image processing apparatus 100 may generate a classifier 430 such that all pixels of the classifier 430 have a first value corresponding to the first purpose. Or, when attempting to perform image processing on the entire first image 410 according to a second purpose, the image processing apparatus 100 may generate a classifier 430 such that all pixels of the classifier 430 have a second value corresponding to the second purpose.
[0073] Or, the image processing apparatus 100 may generate a classifier 430 based on the first image 410. For example, the image processing apparatus 100 may determine whether to process the first image 410 according to a first purpose or a second purpose by analyzing the image of the first image 410. When the first image 410 mainly includes textures, the image processing apparatus 100 may determine that the first image 410 should be processed according to the first purpose, and when the first image 410 mainly includes text or edges, the image processing apparatus 100 may determine that the first image 410 should be processed according to the second purpose. In addition, the image processing apparatus 100 may determine for each region included in the first image 410 whether the corresponding region should be processed according to the first purpose or the second purpose. Thus, when assuming that the first image 410 or a partial region (the first region) of the first image 410 is to be processed according to the first purpose, the image processing apparatus 100 may generate a classifier 430 such that the pixels included in the entire region or a partial region (the region corresponding to the first region) of the classifier 430 have the first value. Or, when assuming that the first image 410 or a partial region (the second region) of the first image 410 is to be processed according to the second purpose, the image processing apparatus 100 may generate a classifier 430 such that the pixels included in the entire region or a partial region (the region corresponding to the second region) of the classifier 430 have the second value. This will be described in detail with reference to Figure 11 and Figure 12 in detail.
[0074] The image processing apparatus 100 may add the generated classifier 430 as an input channel to the DNN 150. Thus, the input image 420 and the classifier 430 may be input to the DNN 150.
[0075] Returning to Figure 4, the image processing apparatus 100 can process the input image 420 according to the purpose indicated by the classifier 430 by using the DNN (150), and output the processed image.
[0076] For example, the DNN 150 may include N convolutional layers (two or more convolutional layers). The DNN 150 has a structure in which input data (e.g., the input image 420 and the classifier 430) is input and passed through N convolutional layers to output output data. In addition, in the DNN 150, other processing operations may be performed in addition to the convolution operation performed by applying one or more kernels to the feature map (feature information), and these processing operations may be performed between the convolutional layers. For example, operations such as an activation function and pooling may be performed.
[0077] The image processing apparatus 100 according to an embodiment can extract "features" such as contours, lines, and colors from the input image 420 by using the DNN 150. Each of the N convolutional layers included in the DNN 150 can receive data, process the received data, and generate output data. For example, the image processing apparatus 100 performs a convolution operation by applying one or more kernels or filters to the image (e.g., the input image 420 and the classifier 430) input to the first convolutional layer 440, and extracts a first feature map (first feature information) as a result of the convolution operation. This will be described in detail with reference to Figure 6 and Figure 7 in detail.
[0078] In addition, the image processing apparatus 100 can apply an activation function to change the value of the extracted first feature map to a non-linear value indicating the "presence or absence" of a feature. In this case, the Relu function may be used, but the present embodiment is not limited thereto. In addition, the image processing apparatus 100 can perform downsampling (pooling) to reduce the size of the extracted feature map, and in this case, max pooling, average pooling, L2 norm pooling, etc. may be used, but the present embodiment is not limited thereto.
[0079] Referring to Figure 4 , the DNN 150 may include M residual blocks. The residual block 450 may include one or more convolutional layers, and the image processing apparatus 100 can perform an operation (element-wisely summing) 470 of element-wisely summing the data 460 (e.g., the first feature map to which the Relu function has been applied), thereby skipping the convolutional layers included in the residual block 450 and the data that has passed through the convolutional layers included in the residual block 450 (e.g., the third feature map extracted from the third convolutional layer and to which the Relu function has been applied).
[0080] In addition, the DNN 150 may include a deconvolution layer 480. Here, the size of the feature information extracted from the deconvolution layer 480 may be greater than the size of the feature information input to the deconvolution layer 480.
[0081] In addition, the image processing device 100 may magnify the data input to the DNN 150. Here, the image processing device 100 may magnify the input data by applying at least one of a bilinear interpolation method, a bicubic interpolation method, and a convolutional interpolation method to the data input to the DNN 150. The image processing device 100 may generate output data by performing an element-wise summation operation (element-wise summation) on the magnified image data and the feature information extracted from the deconvolution layer 480. However, this embodiment is not limited thereto.
[0082] In addition, the image processing device 100 may magnify the images of the U channel and the V channel in the Y channel, U channel, and V channel (YUV3ch). Here, the image processing device 100 may magnify the images of the U channel and the V channel by applying at least one of a bilinear interpolation method, a bicubic interpolation method, and a convolutional interpolation method to the images of the U channel and the V channel. The second image 490 may be generated by concatenating the magnified part of the image data and the output data. In this case, the second image 490 may be obtained by processing the first image 410 according to the purpose indicated by the classifier 430.
[0083] Figure 6 is a reference diagram for describing the process of performing a convolutional operation of the first convolutional layer 440 Figure 4 executed.
[0084] Referring to Figure 6 , assume that the input data 600 of the first convolutional layer 440 has a size of 8*8, and the number of channels is 2 (for example, the input image 610 and the classifier 620). In addition, assume that the size of the kernel applied to the input data 600 is 3*3*2 (horizontal * vertical * depth), and the number of kernels is n. Here, the depth of the kernel has the same value as the number of channels of the input data 600. It can be considered that one kernel includes two sub-kernels, each sub-kernel having a size of 3*3, and the two sub-kernels may respectively correspond to the two channels of the input data 600 (the input image 610 and the classifier 620).
[0085] Referring to Figure 6, it shows the process of extracting the features of the input data 600 by applying the sub-kernels 631 and 632 included in the first kernel Kernel1 to the upper left end to the lower right end of the input data 600. For example, the convolution operation can be performed by applying the first kernel Kernel1 to the pixels included in the 3*3*2 regions 611 and 621 at the upper left end of the input data 600. That is, the pixel values 641 mapped to the 3*3*2 regions 611 and 621 at the upper left end can be generated by multiplying the pixel values included in the 3*3*2 regions 611 and 621 at the upper left end by the weights included in the first kernel Kernel1 respectively and summing the multiplication results.
[0086] In addition, the pixel values 642 mapped to the 3*3*2 regions 612 and 622 that are offset by one pixel from the 3*3*2 regions 611 and 621 at the upper left end of the input data 600 respectively can be generated by multiplying the pixel values included in the 3*3*2 regions 612 and 622 by the weights included in the first kernel Kernel1 respectively and summing the multiplication results.
[0087] In the same way, pixel values can be generated by performing the multiplication and addition of the weights included in the first kernel Kernel1 while scanning the convolution operation region pixel by pixel from left to right and from top to bottom within the input data 600. Therefore, a 6*6 feature map (output data) 640 can be output. In this case, the data of the convolution operation can be scanned while shifting pixel by pixel, but the data of the convolution operation can also be scanned while shifting by two or more pixels at a time. The number of pixels by which the input data is shifted during the scanning process is called the stride, and the size of the feature map to be output can be determined according to the size of this stride.
[0088] Referring to Figure 6 , the input data 600 has a size of 8*8, but the output data 640 has a size of 6*6, which is smaller than the size of the input data 600. The DNN includes multiple convolutional layers, and when passing through multiple convolutional layers, the size of the data keeps decreasing. In this case, when the size of the data decreases before the features are fully extracted, the features of the input data may be lost, and to prevent this problem, padding can be performed. Padding means increasing the size of the input data by assigning a certain value (for example, "0") to the edges of the input data to prevent the output data from decreasing. However, this embodiment is not limited thereto.
[0089] Although Figure 6 only shows the convolution operation result of the first kernel Kernel1, when performing the convolution operation on n kernels, n feature maps can be output. That is, the number of channels of the output data is determined according to the number of kernels (n), and therefore, the number of channels of the input data in the next layer can also be determined.
[0090] Figure 7 Shows the input data and output data of the first convolutional layer 440 according to an embodiment. Figure 4 of the first convolutional layer 440.
[0091] Referring to Figure 7 , the image processing device 100 may input the first input data Input1 into the first convolutional layer 440. Here, the first input data Input1 may include two channels, including the first input image I1 and the first classifier C1. Here, as an image mainly including textures, the first input image I1 may be an image to be processed so as to enhance texture representation when the image is converted into a high-resolution image. In addition, the pixel values of the first classifier C1 may be the first values indicating the first purpose (e.g., detail enhancement of the image). When the first input image I1 and the first classifier C1 are input, the first convolutional layer 440 may perform a convolution operation by applying n kernels to the first input image I1 and the first classifier C1, as described in reference to Figure 6 . As a result of the convolution operation, as Figure 7 shown, n first feature maps 710 may be extracted.
[0092] In addition, the image processing device 100 may input the second input data Input2 into the first convolutional layer 440. Here, the second input data Input2 may include two channels, including the second input image I2 and the second classifier C2. Here, as an image mainly including text or edges, the second input image I2 may be an image to be processed so as to enhance text or edge representation when the image is converted into a high-resolution image. In addition, the pixel values of the second classifier C2 may be the second values indicating the second purpose (e.g., text or edge enhancement).
[0093] When the second input image I2 and the second classifier C2 are input, the first convolutional layer 440 may perform a convolution operation by applying n kernels to the second input image I2 and the second classifier C2, as described in reference to Figure 6 . As a result of the convolution operation, as Figure 7 shown, n second feature maps 720 may be extracted.
[0094] When Figure 7 the first feature maps 710 and the second feature maps 720 are compared with each other, texture features are better represented in the first feature maps 710 than in the second feature maps 720, and text or edge features are better represented in the second feature maps 720 than in the first feature maps 710.
[0095] Figure 8 Shows Figure 1The input data and output data of the first convolutional layer in DNN3.
[0096] Referring to Figure 8 , when the first input image I1 is input into the first convolutional layer included in the DNN3 described in Figure 1 , the third feature map 810 can be extracted. In addition, when the second input image I2 is input into the first convolutional layer included in DNN3, the fourth feature map 820 can be extracted.
[0097] When Figure 7 's first feature map 710 is compared with Figure 8 's third feature map 810, the texture features included in the first input image I1 are better represented in the first feature map 710 than in the third feature map 810. In addition, when Figure 7 's second feature map 720 is compared with Figure 8 's fourth feature map 820, the text or edge features are better represented in the second feature map 720 than in the fourth feature map 820.
[0098] Therefore, the DNN 150 according to the embodiment can perform image processing on the input image according to the purpose indicated by the classifier.
[0099] Figure 9 Shows the input data and output data of the DNN according to the embodiment.
[0100] Referring to Figure 9 , when the first image 910 and the first classifier 931 are input into the DNN 150, the third image 930 is output, and when the second image 920 and the first classifier 931 are input into the DNN150, the fourth image 940 is output. Here, the pixel values of the first classifier 931 can have a first value indicating detail enhancement (enhanced texture representation). When the image to be processed and the first classifier 931 are both input into the DNN 150, the purpose of the image processing can be determined as the detail enhancement of the first classifier 931.
[0101] The first image 910 is an image mainly including texture. Therefore, the output third image 930 can be a high-resolution image obtained by enhancing the details of the first image 910. However, the second image 920 is an image mainly including texture or edges. Therefore, the output fourth image 940 may not show the effect of text or edge enhancement.
[0102] When the first image 910 and the second classifier 932 are input into the DNN 150, the fifth image 950 is output, and when the second image 920 and the second classifier 932 are input into the DNN 150, the sixth image 960 is output. Here, the pixel values of the second classifier 932 may have a second value indicating text or edge enhancement. When both the image to be processed and the second classifier 932 are input into the DNN 150, the purpose of image processing may be determined as text or edge enhancement by the second classifier 932.
[0103] The second image 920 is an image mainly including text or edges. Therefore, the output sixth image 960 may be a high-resolution image obtained by enhancing the text or edges included in the second image 920. However, the first image 910 is an image mainly including textures. Therefore, the output fifth image 950 may not show the effect of texture enhancement.
[0104] Figure 10 The input data and output data of the DNN according to an embodiment are shown.
[0105] Referring to Figure 10 , when the first image 910 and the third classifier 933 are input into the DNN 150, the seventh image 970 may be output. The pixel values of the third classifier 933 may have a third value. Here, when the first value indicating detail enhancement is less than the second value indicating text or edge enhancement, the third value may be less than the first value. Or, when the first value is greater than the second value, the third value may be greater than the first value. The output seventh image 970 may exhibit a better detail enhancement effect than the detail enhancement effect exhibited in the third image 930 described in Figure 9 .
[0106] In addition, when the second image 920 and the fourth classifier 934 are input into the DNN 150, the eighth image 980 may be output. The pixel values of the fourth classifier 934 may have a fourth value. Here, when the second value is greater than the first value, the fourth value may be greater than the second value. Or, when the second value is less than the first value, the fourth value may be less than the second value. In this case, the output eighth image 980 may exhibit a better text or edge enhancement effect than the text or edge enhancement effect exhibited in the sixth image 960 described in Figure 9 .
[0107] Figure 11 is a reference diagram for describing a method of generating a classifier according to an embodiment.
[0108] Referring to Figure 11 , the image processing device 100 according to an embodiment may generate a classifier 1030 based on the first image 1010 to be processed.
[0109] For example, the image processing device 100 may generate a mapping image 1020 indicating the edges and text of the first image 1010 by extracting the edges and text of the first image 1010. The image processing device 100 may extract the edges and text of the first image 1010 by using various known edge extraction filters or text extraction filters.
[0110] In this case, in the mapping image 1020 indicating the edges and text, the pixel values in the edge and text regions may be set to have a second value, while the pixel values in other regions may be set to have a first value. However, this embodiment is not limited thereto.
[0111] The image processing device 100 may generate a classifier 1030 by smoothing the mapping image 1020 indicating the edges and text. Here, smoothing may be an image process for adjusting the pixel values around the edges and text so that the pixel values change smoothly.
[0112] Referring to Figure 11 , in the classifier 1030, the region corresponding to the first region A1 (which is represented as texture) of the first image 1010 may include pixels mainly having the first value, and the region corresponding to the second region A2 (which is represented as text) of the first image 1010 may include pixels mainly having the second value.
[0113] In addition, the image processing device 100 may input the first image 1010 and the generated classifier 1030 into the DNN 150, thereby generating an output image 1040 obtained by converting the first image 1010 into a high-resolution image. Here, in the output image 1040, the region 1041 represented as texture corresponding to the first region A1 of the first image 1010 may exhibit an effect of enhanced detail (enhanced texture representation), and the region 1042 represented as text corresponding to the second region A2 of the first image 1010 may exhibit an effect of enhanced text.
[0114] Figure 12 Shows an identification image generated based on a first image according to an embodiment.
[0115] Referring to Figure 12 , the image processing device 100 may receive an image to be processed. The received image may be classified into an image mainly including a texture representation and an image mainly including a text or edge representation. For example, the first to third images 1210, 1220, and 1230 may be images mainly including a texture representation, and the fourth to sixth images 1240, 1250, and 1260 may be images mainly including a text or edge representation.
[0116] When the received image is an image mainly including a texture representation, the image processing device 100 may determine the pixel values of the classifier as a first value and generate a first classifier C1 in which each pixel value has the first value. In this case, the DNN 150 included in the image processing device 100 may be a network trained by inputting the first classifier C1 and training data for texture representation enhancement when training the network with the training data.
[0117] Alternatively, when the received image is an image mainly including a text or edge representation, the image processing device 100 may determine the pixel values of the classifier as a second value and generate a second classifier C2 having the second value as each pixel value. In this case, the DNN 150 included in the image processing device 100 may be a network trained by inputting the second classifier C2 and training data for text or edge representation enhancement when training the network with the training data.
[0118] Figure 13 is a reference diagram for describing a method of training a DNN according to an embodiment.
[0119] Referring to Figure 13 , the DNN 150 according to an embodiment may be trained by a plurality of training data sets having different purposes. For example, the training data set may include a first training data set D1 and L1 for training image processing according to a first purpose and a second training data set D2 and L2 for training image processing according to a second purpose. Here, the image processing according to the first purpose may include processing an input image to enhance details (texture representation) when the input image is converted into a high-resolution image. In addition, the image processing according to the second purpose may include processing an input image to enhance text or edge representation when the input image is converted into a high-resolution image. Although Figure 13 shows an example of training according to two purposes, the present embodiment is not limited thereto.
[0120] The first training data set D1 and L1 may include first image data D1 and first label data L1. As each piece of image data obtained by converting each piece of the first image data D1 into a high-resolution image, each piece of the first label data L1 may be each piece of image data in which the texture representation is enhanced when each piece of the first image data D1 is converted into a high-resolution image.
[0121] In addition, the second training data set D2 and L2 may include second image data D2 and second label data L2. As image data obtained by converting the second image data D2 into a high-resolution image, each piece of the second label data L2 may be each piece of image data obtained by reducing jaggedness or the like that appears around text or edges (enhanced text or edge representation) when converting the second image data D2 into a high-resolution image.
[0122] The image processing apparatus 100 according to an embodiment may input the first classifier C1 together with the first image data D1 to the DNN 150, and input the second classifier C2 together with the second image data D2 to the DNN 150. Here, the pixel values of the first classifier C1 and the second classifier C2 may be set by a user, and the first classifier C1 and the second classifier C2 may each have a single pixel value. In addition, the pixel value of the first classifier C1 is different from the pixel value of the second classifier C2.
[0123] The image processing apparatus 100 may train the DNN 150 such that when the first image data D1 and the first classifier C1 are input, the first label data L1 corresponding to the first image data D1 is output. For example, the image processing apparatus 100 may adjust the weights of one or more kernels included in the DNN 150 in order to reduce the difference between each piece of the first label data L1 and each piece of image data output when each piece of the first image data D1 and the first classifier C1 are input to the DNN 150.
[0124] In addition, the image processing apparatus 100 may train the DNN 150 such that when the second image data D2 and the second classifier C2 are input, the second label data L2 corresponding to the second image data D2 is output. For example, the image processing apparatus 100 may adjust the weights of one or more kernels included in the DNN 150 in order to reduce the difference between the second label data L2 and the image data output when the second image data D2 and the second classifier C2 are input to the DNN 150.
[0125] Figure 14 is a reference diagram according to an embodiment for describing a DNN configured to perform image processing according to multiple purposes.
[0126] Refer to Figure 14, the DNN 150 according to the embodiment can be trained by the training dataset 1420. The training dataset 1420 can include first to nth image data and first to nth label data corresponding to n pieces of image data respectively. For example, as the image data obtained by converting n pieces of image data into high-resolution images respectively, the n pieces of label data can be the image data obtained by processing each of the n pieces of image data according to at least one of a first purpose (e.g., detail enhancement (texture representation enhancement)), a second purpose (e.g., noise reduction), and a third purpose (e.g., coding artifact reduction). However, this embodiment is not limited thereto.
[0127] The DNN 150 can receive a first classifier C1, a second classifier C2, and a third classifier C3 (four channels) as well as image data. Here, the first classifier C1 can be an image indicating the degree of detail enhancement, the second classifier C2 can be an image indicating the degree of noise reduction, and the third classifier C3 can be an image indicating the degree of coding artifact reduction.
[0128] For example, based on the first image data 1410 and the first label data, the pixel values of the first classifier C1, the second classifier C2, and the third classifier C3 input to the DNN 150 and the first image data 1410 can be determined. By comparing the first image data 1410 with the first label data obtained by converting the first image data 1410 into a high-resolution image, the pixel value of the first classifier C1 can be determined according to the degree of detail enhancement presented in the first label data, the pixel value of the second classifier C2 can be determined according to the degree of noise reduction, and the pixel value of the third classifier C3 can be determined according to the degree of coding artifact reduction. As the degree of detail enhancement increases, the pixel value of the first classifier C1 can be smaller.
[0129] For example, when the degree of detail enhancement presented in the first label data when comparing the first image data 1410 with the first label data is greater than the degree of detail enhancement presented in the second label data when comparing the second image data with the second label data, the pixel value of the first classifier C1 input together with the first image data 1410 can be smaller than the pixel value of the first classifier C1 input together with the second image data.
[0130] In addition, as the degree of noise reduction increases, the pixel value of the second classifier C2 can be smaller, and as the degree of coding artifact reduction increases, the pixel value of the third classifier C3 can be smaller. However, this embodiment is not limited thereto, and the pixel values of the identification images can be determined by various methods.
[0131] In addition, the first to third classifiers C1, C2, and C3, which are input together with each of the n image data, can be determined differently according to the degree of detail enhancement, the degree of noise reduction, and the degree of reduction of coding artifacts presented in the n label data.
[0132] In addition, the image processing device 100 can train the DNN 150 such that the first label data is output when the first image data 1410, the first classifier C1, the second classifier C2, and the third classifier C3 are input. For example, the image processing device 100 can adjust the weights of one or more kernels included in the DNN 150 to reduce the difference between the first label data and the image data 1430 output when the first image data 1410, the first classifier C1, the second classifier C2, and the third classifier C3 are input to the DNN 150.
[0133] When processing an input image by using the DNN 150 trained in the same manner as described above, the image processing purpose of the input image and the image processing level according to the purpose can be determined by adjusting the pixel values of the first to third classifiers C1, C2, and C3. For example, when a greater degree of detail enhancement, a smaller degree of noise reduction, and a smaller degree of reduction of coding artifacts are required in the input image, the pixel value of the first classifier C1 can be set to a smaller value, and the pixel values of the second and third classifiers C2 and C3 can be set to larger values. However, this embodiment is not limited thereto.
[0134] Figure 15 is a reference diagram for describing a method of generating a training data set for training a DNN according to an embodiment.
[0135] Referring to Figure 15 , the training data set can include multiple pieces of image data (e.g., first to third image data) 1510, 1520, and 1530 and one piece of label data 1540. Here, multiple pieces of image data 1510, 1520, and 1530 can be generated by using the label data 1540. For example, the image processing device 100 can generate the first image data 1510 by blurring the label data 1540 with a first intensity, generate the second image data 1520 by blurring the label data 1540 with a second intensity, and generate the third image data 1530 by blurring the label data 1540 with a third intensity.
[0136] The image processing device 100 can input each of the first to third image data 1510, 1520, and 1530 and a classifier indicating the degree of detail enhancement to the DNN 150. Here, the pixel value of the first classifier C1 input together with the first image data 1510 can be set to a first value. In addition, the pixel value of the second classifier C2 input together with the second image data 1520 can be set to a second value, and the second value can be smaller than the first value. In addition, the pixel value of the third classifier C3 input together with the third image data 1530 can be set to a third value, and the third value can be smaller than the second value. However, this embodiment is not limited thereto.
[0137] The image processing device 100 can train the DNN 150 so as to reduce the difference between the label data 1540 and the output data output corresponding to each of the input first image data 1510, second image data 1520, and third image data 1530.
[0138] Although only Figure 15 the method for generating a training data set for detail enhancement is described, a training data set for noise reduction or coding artifact reduction can also be generated in the same manner.
[0139] Figure 16 is a block diagram of an image processing device according to an embodiment.
[0140] Referring to Figure 16 , the image processing device 100 according to an embodiment can include a processor 130 and a memory 120.
[0141] The processor 130 according to an embodiment can generally control the image processing device 100. The processor 130 according to an embodiment can execute one or more programs stored in the memory 120.
[0142] The memory 120 according to an embodiment can store various data, programs, or applications for operating and controlling the image processing device 100. The programs stored in the memory 120 can include one or more instructions. The programs (one or more instructions) or applications stored in the memory 120 can be executed by the processor 130.
[0143] The processor 130 according to an embodiment can obtain a first image and a classifier indicating the purpose of image processing, and process the first image by using the DNN according to the purpose indicated by the classifier. Here, the DNN can be the DNN shown and Figures 2 to 15 referred to in Figures 2 to 15 the description.
[0144] Alternatively, the processor 130 may generate an input image and a classifier to be input to the DNN based on the first image. For example, the processor 130 may convert the R channel, G channel, and B channel included in the first image into the Y channel, U channel, and V channel through color space conversion, and determine the image of the Y channel among the Y channel, U channel, and V channel as the input image. When attempting to perform image processing on the first image according to the first purpose, the processor 130 may determine the pixel value of the classifier as the first value corresponding to the first purpose, and when attempting to perform image processing on the first image according to the second purpose, the processor 130 may determine the pixel value of the classifier as the second value corresponding to the second purpose. Alternatively, the processor 130 may extract edges and text from the first image, and generate a classifier based on the extracted edges and text.
[0145] The DNN may include N convolutional layers, and the processor 130 may extract feature information (feature maps) by performing a convolution operation in which one or more kernels are applied to the input image and the identity image in each of the N convolutional layers, and process the first image based on the extracted feature information according to the purpose indicated by the classifier.
[0146] The processor 130 may determine the image processing level according to the first purpose and the image processing level according to the second purpose based on the value of the pixels included in the classifier, the first value corresponding to the first purpose, and the second value corresponding to the second purpose, and process the first image according to the determined levels.
[0147] The processor 130 may train the DNN by using multiple training data sets having different purposes. For example, the processor 130 may train the DNN by using a first training data set and a second training data set. The first training data set includes multiple first image data, a first classifier, and multiple first label data obtained by processing the multiple first image data according to the first purpose. The second training data set includes multiple second image data, a second classifier, and multiple second label data obtained by processing the multiple second image data according to the second purpose.
[0148] For example, the processor 130 may adjust the weights of one or more kernels included in the DNN to reduce the difference between the first label data and the multiple image data output when the multiple first image data and the first classifier are input to the DNN. In addition, the processor 130 may adjust the weights of one or more kernels included in the DNN to reduce the difference between the second label data and the multiple image data output when the second image data and the second classifier are input to the DNN.
[0149] Figure 17 is a block diagram of the processor 130 according to an embodiment.
[0150] Reference Figure 17 Referring to Figure 17 , the processor 130 according to the embodiment may include a network trainer 1400 and an image processor 1500.
[0151] The network trainer 1400 may obtain training data for training the DNN according to the embodiment. The network trainer 1400 may obtain a plurality of training data sets with different purposes. For example, the training data set may include a first training data set for training of image processing according to a first purpose and a second training data set for training of image processing according to a second purpose. Here, the image processing according to the first purpose may include processing an input image to enhance details (texture representation) when the input image is converted into a high-resolution image. In addition, the image processing according to the second purpose may include processing an input image to enhance text or edge representation when the input image is converted into a high-resolution image. However, the present embodiment is not limited thereto.
[0152] Alternatively, the network trainer 1400 may generate training data for training the DNN. For example, the training data set may be generated by referring to the method described in Figure 15 . Figure 15 described.
[0153] The network trainer 1400 may learn a reference on how to process an input image based on a plurality of training data sets with different purposes. Alternatively, the network trainer 1400 may learn which reference of the training data set is considered for processing the input image. For example, the network trainer 1400 may train the DNN by using a plurality of training data sets by referring to the methods described in Figure 13 and Figure 14 . Figure 13 and Figure 14 described.
[0154] The network trainer 1400 may store the trained network (e.g., DNN) in the memory of the image processing device. Alternatively, the network trainer 1400 may store the trained network in the memory of a server connected to the image processing device via a wired or wireless network.
[0155] The memory storing the trained network may also store, for example, commands or data associated with at least one other component of the image processing device 100. In addition, the memory may store software and / or programs. The programs may include, for example, kernels, middleware, application programming interfaces (APIs), and / or applications (or “apps”).
[0156] Alternatively, the network trainer 1400 may train the DNN based on the high-resolution image generated by the image processor 1500. For example, the DNN may be trained by using a training algorithm including error backpropagation or gradient descent.
[0157] The image processor 1500 may process the first image based on the first image and a classifier indicating the purpose of image processing. For example, the image processor 1500 may process the first image by using a trained DNN according to the purpose indicated by the classifier.
[0158] At least one of the network trainer 1400 and the image processor 1500 may be manufactured in the form of a hardware chip and installed in an image processing device. For example, at least one of the network trainer 1400 and the image processor 1500 may be manufactured in the form of a dedicated hardware chip for artificial intelligence (AI), or may be manufactured as part of a conventional general-purpose processor (e.g., a central processing unit (CPU) or an application processor) or a graphics dedicated processor (e.g., a graphics processing unit (GPU)) and installed in various image processing devices.
[0159] In this case, the network trainer 1400 and the image processor 1500 may be installed in one image processing device or in separate image processing devices respectively. For example, a part of the network trainer 1400 and the image processor 1500 may be included in the image processing device 100, while the remaining part may be included in a server.
[0160] In addition, at least one of the network trainer 1400 and the image processor 1500 may be implemented by a software module. When at least one of the network trainer 1400 and the image processor 1500 is implemented by a software module (or a program module including instructions), the software module may be stored in a non-transitory computer-readable recording medium. In addition, in this case, at least one software module may be provided by an operating system (OS) or a certain application. Alternatively, a part of at least one software module may be provided by the OS, and the remaining part may be provided by a specific application.
[0161] Figure 18 is a block diagram of the image processor 1500 according to an embodiment.
[0162] Referring to Figure 18 , the image processor 1500 may include an input image generator 1510, a classifier generator 1520, a DNN unit 1530, an output image generator 1540, and a network updater 1550.
[0163] The input image generator 1510 may receive a first image to be processed and generate an input image to be input to the DNN based on the first image. For example, the input image generator 1510 may convert the R, G, and B channels included in the first image into the Y channel, U channel, and V channel through color space conversion, and determine the image of the Y channel among the Y channel, U channel, and V channels as the input image. Alternatively, the input image generator 1510 may determine the R, G, and B channels included in the first image as the input image.
[0164] The classifier generator 1520 may generate a first classifier having a first value corresponding to the first purpose as a pixel value when attempting to perform image processing on the first image according to the first purpose, and generate a second classifier having a second value corresponding to the second purpose as a pixel value when attempting to perform image processing on the first image according to the second purpose. The classifier generator 1520 may generate a classifier having a single pixel value or a classifier having different pixel values for each region.
[0165] The classifier generator 1520 may extract edges and text from the first image and generate a classifier based on the extracted edges and text. For example, the classifier generator 1520 may generate a mapping image indicating the edges and text of the first image by extracting the edges and text of the first image. In this case, the classifier generator 1520 may extract the edges and text of the first image by using various known edge extraction filters or text extraction filters. In addition, the classifier generator 1520 may generate a classifier by smoothing the mapping image indicating the edges and text.
[0166] The DNN unit 1530 according to an embodiment may perform image processing on the first image by using the DNN trained by the network trainer 1400. For example, the input image generated by the input image generator 1510 and the classifier generated by the classifier generator 1520 may be input as input data of the DNN. The DNN according to an embodiment may extract feature information by performing a convolution operation in which one or more kernels are applied to the input image and the classifier. The DNN unit 1530 may process the first image based on the extracted feature information according to the purpose indicated by the classifier.
[0167] The output image generator 1540 may generate a final image (second image) based on the data output from the DNN. For example, the output image generator 1540 may enlarge the images of the U and V channels of the first image and generate a final image by concatenating the enlarged images with the image (image of the Y channel of the first image) processed by the DNN and output from the DNN. Here, the final image may be an image obtained by processing the first image according to the purpose indicated by the classifier.
[0168] The network updater 1550 may update the DNN based on an evaluation of the output image provided by the DNN unit 1530 or the final image provided by the output image generator 1540. For example, the network updater 1550 may allow the network trainer 1400 to update the DNN by providing the network trainer 1400 with the image data provided by the DNN unit 1530 or the output image generator 1540.
[0169] At least one of the input image generator 1510, the classifier generator 1520, the DNN unit 1530, the output image generator 1540, and the network updater 1550 may be fabricated in the form of a hardware chip and installed in an image processing device. For example, at least one of the input image generator 1510, the classifier generator 1520, the DNN unit 1530, the output image generator 1540, and the network updater 1550 may be fabricated in the form of a dedicated hardware chip for AI, or may be fabricated as part of a conventional general-purpose processor (e.g., a CPU or an application processor) or a graphics dedicated processor (e.g., a GPU) and installed in various image processing devices.
[0170] In this case, the input image generator 1510, the classifier generator 1520, the DNN unit 1530, the output image generator 1540, and the network updater 1550 may be installed in one image processing device or in separate image processing devices. For example, some of the input image generator 1510, the classifier generator 1520, the DNN unit 1530, the output image generator 1540, and the network updater 1550 may be included in the image processing device, while some others may be included in the server.
[0171] In addition, at least one of the input image generator 1510, the classifier generator 1520, the DNN unit 1530, the output image generator 1540, and the network updater 1550 may be implemented by a software module. When at least one of the input image generator 1510, the classifier generator 1520, the DNN unit 1530, the output image generator 1540, and the network updater 1550 is implemented by a software module (or a program module including instructions), the software module may be stored in a non-transitory computer-readable recording medium. In addition, in this case, at least one software module may be provided by the OS or a certain application. Alternatively, a part of at least one software module may be provided by the OS, and the remaining part may be provided by a specific application.
[0172] As Figures 16 to 18The block diagrams of the image processing apparatus 100, the processor 130, and the image processor 1500 shown are only for one embodiment. Each component in the block diagrams may be integrated, added, or omitted according to the specifications of the image processing apparatus 100. That is, depending on the environment, two or more components may be integrated into one component, or one component may be divided into two or more components. In addition, the functions performed by each block describe the embodiments, and the specific operations or devices do not limit the proper scope of the present invention.
[0173] The method of operating an image processing apparatus according to an embodiment may be implemented in the form of program commands executable by various computer devices and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., alone or in combination. The program commands recorded on the medium may be those specially designed and constructed for the present invention, or those known and available to those of ordinary skill in the computer software field. Examples of the computer-readable recording medium include: magnetic media, such as hard disks, floppy disks, or magnetic tapes; optical media, such as compact disc read-only memory (CD-ROM); or digital versatile discs (DVD); magneto-optical media, such as optical discs; and hardware devices specifically configured to store and execute program commands, such as ROM, RAM, or flash memory. Examples of program commands include high-level language codes executable by a computer using an interpreter and machine language codes generated by a compiler.
[0174] In addition, an image processing apparatus and a method of operating the image processing apparatus for generating a high-resolution video according to an embodiment may be provided by being included in a computer program product. The computer program product may be traded as a product between a seller and a buyer.
[0175] The computer program product may include a software (S / W) program and a computer-readable storage medium in which the S / W program is stored. For example, the computer program product may include an S / W program in the form of a downloadable application, such as a product electronically distributed by a manufacturing company of an electronic device or an electronic market (e.g., Google PlayStore or App Store). For electronic distribution, at least a part of the S / W program may be stored in a storage medium or generated temporarily. In this case, the storage medium may be included in a server of the manufacturing company, a server of the electronic market, or a relay server configured to temporarily store the S / W program.
[0176] A computer program product may include a storage medium of a server or a storage medium of a client device in a system including a server and a client device. Alternatively, when there is a third device (e.g., a smart phone) communicatively connected to the server or the client device, the computer program product may include a storage medium of the third device. Alternatively, the computer program product may include an S / W program that is to be transmitted from the server to the client device or the third device or from the third device to the client device.
[0177] In this case, one of the server, the client device, and the third device may execute the computer program product and perform the method according to the embodiment. Alternatively, two or more of the server, the client device, and the third device may execute the computer program product and perform the method according to the embodiment in a distributed manner.
[0178] For example, a server (e.g., a cloud server or an AI server) may execute the computer program product stored in the server to control a client device communicatively connected to the server, where the client device performs the method according to the disclosed embodiment.
[0179] Although the embodiments have been described in detail, the proper scope of the present invention is not limited thereto, and various modifications and improvements obtained by those of ordinary skill in the art using the basic concepts of the present invention defined in the claims also fall within the scope of the present disclosure.
Claims
1. An image processing device, comprising: a memory storing one or more instructions; and a processor configured to execute the one or more instructions stored in the memory, wherein the processor is further configured to: obtain a first image and a classifier corresponding to the first image, wherein the classifier is an image including at least one region, and pixel values within the same region are the same, and different pixel values are used to indicate image processing methods with different image processing purposes; and input the first image and the classifier into a pre-trained deep neural network DNN so that the DNN processes the first image according to the purpose indicated by the classifier, wherein the DNN is obtained by training with a training image and a classifier as inputs and an image obtained by processing the training image in the image processing method indicated by the classifier as an output.
2. The image processing device according to claim 1, wherein the DNN includes N convolutional layers, and the processor is further configured to: generate the input image based on the first image, extract feature information by performing a convolution operation, in which one or more kernels are applied to the input image and the classifier in the N convolutional layers and generate a second image based on the extracted feature information.
3. The image processing device according to claim 2, wherein the processor is further configured to: convert the R channel, G channel, and B channel included in the first image into the Y channel, U channel, and V channel in YUV mode, and determine the image of the Y channel among the Y channel, the U channel, and the V channel as the input image.
4. The image processing device according to claim 3, wherein the processor is further configured to: generate the second image based on the image output by processing the image of the Y channel in the DNN and based on the images of the U channel and the V channel among the Y channel, the U channel, and the V channel.
5. The image processing device according to claim 1, wherein the pixels included in the classifier have at least one of a first value and a second value greater than the first value, the first purpose indicated by the first value is texture enhancement, and the second purpose indicated by the second value is text or edge enhancement.
6. The image processing device according to claim 5, wherein the processor is further configured to: when all the pixels included in the classifier have the first value, process the first image according to the first purpose; and when all the pixels included in the classifier have the second value, process the first image according to the second purpose.
7. The image processing device according to claim 5, wherein When pixels in a first region included in the classifier have the first value and pixels in a second region included in the classifier have the second value, a third region corresponding to the first region in the first image is processed according to the first purpose, and a fourth region corresponding to the second region in the first image is processed according to the second purpose.
8. The image processing apparatus according to claim 5, wherein, the processor is further configured to: process the first image according to an image processing level based on the first purpose and an image processing level based on the second purpose, and the image processing level based on the first purpose and the image processing level based on the second purpose are determined based on the values of the pixels included in the classifier, the first value, and the second value.
9. The image processing apparatus according to claim 1, wherein, the processor is further configured to: generate the classifier based on characteristics of the first image.
10. The image processing apparatus according to claim 9, wherein, the processor is further configured to: generate a mapping image indicating text and edges included in the first image, and determine the values of the pixels included in the classifier based on the mapping image.
11. The image processing apparatus according to claim 1, wherein, the DNN is trained by a first training data set including first image data, a first classifier having the first value as a pixel value, and first label data obtained by processing the first image data according to the first purpose, and is trained by a second training data set including second image data, a second classifier having the second value as a pixel value, and second label data obtained by processing the second image data according to the second purpose.
12. The image processing apparatus according to claim 11, wherein, the processor is further configured to: adjust the weights of one or more kernels included in the DNN to reduce the difference between the first label data and the image data output when the first image data and the first classifier are input to the DNN; and adjust the weights of the one or more kernels included in the DNN to reduce the difference between the second label data and the image data output when the second image data and the second classifier are input to the DNN.
13. An operation method of an image processing apparatus, the method comprising: acquiring a first image and a classifier corresponding to the first image, wherein the classifier is an image including at least one region, and pixel values within the same region are the same, and different pixel values are used to indicate image processing methods with different image processing purposes; and inputting the first image and the classifier into a pre-trained deep neural network DNN so that the DNN processes the first image according to the purpose indicated by the classifier; wherein the DNN is obtained by training with a training image and a classifier as inputs and an image obtained by processing the training image in the image processing method indicated by the classifier as an output.
14. The method according to claim 13, wherein, obtaining the first image and the classifier includes: generating an input image based on the first image, the DNN includes N convolutional layers, and processing the first image by the DNN according to the purpose indicated by the classifier includes: extracting feature information by performing a convolution operation, in which one or more kernels are applied to the input image and the classifier in the N convolutional layers; and generating a second image based on the extracted feature information.
15. The method according to claim 14, wherein, generating the input image based on the first image includes: converting the R channel, G channel, and B channel included in the first image into the Y channel, U channel, and V channel in YUV mode; and determining the image of the Y channel among the Y channel, the U channel, and the V channel as the input image.
Citation Information
Patent Citations
System and method for CNN layer sharing
US20180157916A1