Image restoration method and imaging system for executing same

Through the artificial intelligence model, the adversarial learning is performed in Fourier space, and the clear and correct data is directly learned in the frequency domain, solving the problem of image quality deterioration caused by optical aberration and image damage in the high spatial frequency domain in ultra-small imaging systems, achieving efficient image repair effects.

CN120167066APending Publication Date: 2025-06-17INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480002278.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-04
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of image quality deterioration caused by optical aberrations and image damage in high spatial frequency domains in ultra-small imaging systems.

Method used

By using an artificial intelligence model, image data directly captured by lenses with strong aberrations are learned and an adversarial learning method is performed in Fourier space to directly learn clear and correct data in the frequency domain to generate repair images.

Benefits of technology

Improvement of image quality degradation in ultra-small imaging systems, especially image damage repair in high spatial frequency domains, improving image clarity and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120167066A_ABST
    Figure CN120167066A_ABST
Patent Text Reader

Abstract

The imaging restoration device according to one embodiment of the present invention comprises: a memory for storing an original image captured from a camera and an artificial intelligence model; and a processor for training the artificial intelligence model, the artificial intelligence model comprising: an image restoration model for cutting the original image with a preset block and generating a restored image based on input data of coordinate information of the block embedded in the cut image block; and a discriminating model for performing Fourier transform on the inpainting image generated by the image inpainting model and discriminating the inpainting image after the Fourier transform, the processor executing adversarial learning of the image inpainting model and the discriminating model, and generating a restored image of a new original image shot by the camera from the learned image restoration model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image restoration method for improving the performance of an ultra-small imaging system with optical aberration using information embedding technology and an adversarial learning method performed in the Fourier space, and an imaging system for performing the method. Background Art

[0002] A metalens is an ultrathin lens composed of ultrashort-wavelength structures and has attracted attention as a technology that can overcome the limitations of existing lenses. However, according to recent research, there is a fundamental trade-off between broadband focusing efficiency and diameter in a large-area broadband metalens. Therefore, currently disclosed broadband metalenses have chromatic aberration or low focusing efficiency within a wide bandwidth, which will hinder the commercialization of small imaging systems based on metalenses.

[0003] On the one hand, previously disclosed image restoration techniques are as follows.

[0004] Prior Paper 1 relates to an image restoration method by deconvolution. In Prior Paper 1, assuming that the performance does not change due to position, deconvolution is performed on a damaged image using a pre-measured point spread function (PSF). That is, in the technique of Prior Paper 1, since it is assumed that the performance does not change due to position, there is a problem that the performance degradation according to position cannot be considered.

[0005] Prior Paper 2 is an image restoration technique by deep learning. The technique of Prior Paper 2 uses randomly cropped image patches from a given image instead of using the original full-resolution image for the physical limitations of the GPU (Graphic Processing Unit) and learning efficiency. In this case, when the model is trained on the given data, the position information of the image is completely lost, resulting in the inability to learn the performance degradation caused by the aberration of the lens.

[0006] Prior Paper 3 discloses a patch-wise deconvolution method based on a patch-based PSF. In Prior Paper 3, in order to repair image damage caused by general optical aberration with a PSF that changes according to the viewing angle, the PSF for each patch unit region on the image is directly measured by a person and trained, so there are problems in terms of cost and time.

[0007] Prior Paper 1

[0008] Krishnan, Dilip, and Rob Fergus. “Fast image deconvolution using hyper-Laplacian priors.” Advances in neural information processing systems 22 (2009).

[0009] Prior Paper 2

[0010] Zamir, Syed Waqas, et al. “Multi-stage progressive image restoration.” Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021.

[0011] Prior Paper 3

[0012] Li, Xiu, et al. “Universal and flexible optical aberration correction using deep-prior based deconvolution.” Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021. Summary of the Invention

[0013] Technical Problem

[0014] An embodiment of the present invention relates to an image restoration method and an imaging system for performing the method. By using an artificial intelligence model learned from image data directly captured using a lens with strong aberration and directly learning clear and correct data in the frequency domain, it is possible to solve the deterioration of image quality in an ultra-small imaging system and repair image damage in the high spatial frequency domain.

[0015] Technical Solution

[0016] An imaging restoration device according to an embodiment of the present invention, including: a memory for storing an original image captured by a camera and an artificial intelligence model; and a processor for training the artificial intelligence model, the artificial intelligence model including: an image restoration model for cropping the original image with a preset block, and generating a restored image based on input data embedding coordinate information of the block in the cropped image block; and a discriminant model for performing Fourier transform on the restored image generated by the image restoration model and distinguishing the Fourier-transformed restored image, the processor performing adversarial learning of the image restoration model and the discriminant model, and generating a restored image of a new original image captured by the camera from the learned image restoration model.

[0017] The processor may generate the coordinate information based on the middle pixel of the image block, convert the coordinate information into two-dimensional coordinate data by a Meshgrid method, and combine (concatenate) the coordinate information converted into the two-dimensional data to the image block through a 1×1 convolutional layer to generate the input data.

[0018] The processor may input the input data with the same size of the input data and the output data output from the 1×1 convolutional layer to a first neural network, and output the restored image from the first neural network.

[0019] The first neural network may include a Convolution Neural Network (CNN) model including at least one of MIRNet, MPRNet, or NAFNet, or a Transformer model including at least one of Restormer or Uformer.

[0020] The discriminant model may include a CNN model for discriminating the authenticity of the Fourier-transformed restored image, the CNN model including at least one of GoogleNet, AlexNet, and VGG network. The processor compares the Fourier-transformed restored image with a correct image through the discriminant model, and adversarially trains the image restoration model based on the comparison result, so that the image restoration model generates a high-frequency restored image from a low-frequency original image.

[0021] The processor may train the discriminant model to output a value above a preset first reference value for the correct image, and train the discriminant model to output a value below a preset second reference value for the restored image.

[0022] An imaging system according to other embodiments of the present invention includes: a camera including a meta-lens or a diffractive optical lens; a memory that receives a raw image captured by the camera and stores an artificial intelligence model; a processor for training the artificial intelligence model; and a display for outputting a restored image in which the processor restores the raw image through the artificial intelligence model that has completed learning. The artificial intelligence model includes: an image restoration model that crops the raw image with a preset block and generates a restored image based on input data in which coordinate information of the block is embedded in the cropped image block; and a discriminative model that performs a Fourier transform on the restored image generated by the image restoration model and discriminates the Fourier-transformed restored image.

[0023] The processor may generate the coordinate information based on the middle pixel of the image block, convert the coordinate information into two-dimensional coordinate data by a grid method, and combine the coordinate information converted into the two-dimensional data with the image block through a 1×1 convolutional layer to generate the input data.

[0024] The processor may input the input data output from the 1×1 convolutional layer, the size of which is the same as the size of the output data, into a first neural network and output the restored image from the first neural network.

[0025] The processor may compare the Fourier-transformed restored image with a correct image through the discriminative model and adversarially train the image restoration model based on the comparison result so that the image restoration model generates a high-frequency restored image from a low-frequency raw image.

[0026] The processor may train the discriminative model to output a value above a preset first reference value for the correct image and train the discriminative model to output a value below a preset second reference value for the restored image.

[0027] An image restoration method according to still another embodiment of the present invention includes: training an artificial intelligence model; a camera including a meta-lens acquires a raw image; generating a restored image of the raw image through the artificial intelligence model that has completed learning. The artificial intelligence model includes: an image restoration model that crops an image of learning data with a preset block and generates a restored image based on input data in which coordinate information of the block is embedded in the cropped image block; and a discriminative model that performs a Fourier transform on the restored image generated by the image restoration model and discriminates the Fourier-transformed restored image. Training the artificial intelligence model includes comparing the Fourier-transformed restored image with a correct image through the discriminative model and adversarially training the image restoration model based on the comparison result so that the image restoration model generates a high-frequency restored image from a low-frequency raw image.

[0028] The image inpainting model generates the inpainted image, which may include: generating the coordinate information based on the middle pixel of the image patch; converting the coordinate information into two-dimensional coordinate data through a grid method; and combining the coordinate information converted into the two-dimensional data with the image patch through a 1×1 convolutional layer to generate the input data.

[0029] The image inpainting model generates the inpainted image, which may include inputting the input data output from the 1×1 convolutional layer into a first neural network with the same size of the input data and the output data, and outputting the inpainted image from the first neural network.

[0030] Training the artificial intelligence model may include training the discriminative model to output a value above a preset first reference value for the correct image, and training the discriminative model to output a value below a preset second reference value for the inpainted image.

[0031] Advantages of the Invention

[0032] According to an embodiment of the present invention, an image inpainting method and an imaging system for performing the method can solve the problem of image quality degradation in a ultra-small imaging system by using an artificial intelligence model learned with image data directly captured by a lens with strong aberration.

[0033] In addition, the image inpainting method and the imaging system for performing the method of the present invention can repair image damage in a high spatial frequency domain by directly learning clear correct data in the frequency domain. Therefore, when the image inpainting method and the imaging system for performing the method of the present invention are installed on a wearable device such as a smart phone, the phenomenon of the camera of the smart phone protruding can be solved, and high-quality images / videos can be captured.

[0034] In addition, the image inpainting method and the imaging system for performing the method of the present invention also exhibit excellent performance in downstream applications that can be performed on images captured by cameras provided in unmanned aerial vehicles and AR (Augmented Reality) / VR (Virtual Reality) devices based on the improved image / video quality.

[0035] The image inpainting method and the imaging system for performing the method of the present invention can achieve the light weight of the lens, thereby reducing the weight of the mounted imaging system, and it is expected to reduce the weight of unmanned aerial vehicles, drones, etc. Thus, an improvement in power efficiency can also be expected. In addition, the image inpainting method and the imaging system for implementing the method of the present invention can reduce the weight of the device itself by installing a ultra-small imaging system in an imaging system that acts as an eye on an AR / VR device, and an improvement in power efficiency is expected. Moreover, by reducing the fatigue of the user, the user experience can be improved. Brief Description of the Drawings

[0036] Figure 1 is a schematic diagram for explaining the structures of an imaging system according to an embodiment of the present invention;

[0037] Figure 2 is a control block diagram of an imaging restoration device according to an embodiment of the present invention;

[0038] Figure 3 is an overall flowchart of an image restoration method according to an embodiment of the present invention;

[0039] Figure 4 is a diagram for specifically explaining an artificial intelligence model according to an embodiment of the present invention;

[0040] Figure 5 and Figure 6 is a diagram for explaining a method of generating and combining coordinate information;

[0041] Figure 7 is a flowchart for specifically explaining a method of performing learning in an image restoration method according to an embodiment of the present invention;

[0042] Figure 8 is a diagram for explaining a discrimination process according to an embodiment;

[0043] Figure 9 is an example of the result of the imaging system of the present invention;

[0044] Figure 10 is a table showing the performance comparison results of the image restoration method of the present invention. Detailed Description of the Embodiments

[0045] Throughout the specification, the same reference numerals refer to the same elements. This specification does not describe all elements of the embodiments, and common content in the technical field to which the present invention pertains or repeated content between embodiments is omitted.

[0046] Throughout the specification, when a part is "connected" to another part, this includes not only the case of direct connection but also the case of indirect connection.

[0047] In addition, unless otherwise stated to the contrary, when a part "includes" a certain component, this means that other components are not excluded, and other components can be further included.

[0048] When not specifically mentioned in the text, the singular form includes the plural form.

[0049] In addition, terms such as "~ section", "~ device", "~ block", "~ component", "~ module", etc. can mean a unit that processes at least one function or action. For example, these terms can mean at least one hardware such as FPGA (field-programmable gate array) / ASIC (application specific integrated circuit), at least one software stored in a memory, or at least one program processed by a processor.

[0050] The symbols attached to each step are used to identify each step, and these symbols do not indicate the mutual order of each step. Unless the specific order is clearly stated in the context, each step can be implemented in an order different from the described order.

[0051] Hereinafter, an embodiment of a speaker verification device according to one aspect will be described in detail with reference to the accompanying drawings.

[0052] Figure 1 It is a schematic diagram for explaining each structure of an imaging system according to an embodiment of the present invention.

[0053] An imaging system 1 according to an embodiment of the present invention includes: a camera 5 for photographing an object (Object, Ob); and an image restoration device 10-1, 10-2, 10-3, 10 that receives a raw image photographed (or acquired) by the camera 5 and generates a restored image by the image restoration method of the present invention.

[0054] Specifically, the camera 5 according to an embodiment of the present invention may be a CMOS (Complementary Metal Oxide Semiconductor) image sensor including a super-small lens such as a meta-lens and a diffractive optics lens. Due to the performance limitations of a single lens, commercial lenses are usually prepared by stacking multiple lenses. As a result, various electronic devices such as smartphones 10-2 and drones 10-3 that use commercial lenses have problems of gradually increasing weight and size, and decreasing power efficiency. To solve this problem, meta-lenses (or singlet lenses) have recently been developed, and the camera 5 including such super-small lenses is expected to reduce the weight of the user terminal 10 and improve power efficiency.

[0055] However, for the camera 5 equipped with such a ultra-small lens, due to process limitations and theoretical limitations, there is a problem of degraded performance in restoring images. In particular, for the camera 5 including such an ultra-small lens, strong aberration is generated due to the limitations of the small and thin lens. In addition, for the camera 5 including the ultra-small lens, there are significant differences in the performance of restoring images depending on the shooting space, and it is difficult to use existing general restoration models (including artificial intelligence models).

[0056] The user terminal 10 that receives the original image captured by the camera 5, that is, the original image with strong aberration, includes an artificial intelligence model ( Figure 2 20 in) trained by an adversarial learning method, and can generate an image (hereinafter, restored image) for restoring the original image received by the artificial intelligence model 20 that has completed learning.

[0057] As Figure 1As shown, the user terminal 10 can be implemented by the server 10-1, the smart phone 10-2, or the drone 10-3, and can also be implemented by a computer or a portable terminal that receives the original image through a network. Here, the computer includes, for example, a laptop equipped with a web browser, a desktop computer, a laptop, a tablet PC, a slate PC, etc., and the portable terminal is a wireless communication device that ensures portability and mobility, and can include a personal communication system (PCS), a global system for mobile communications (GSM), a personal digital cellular system (PDC), a personal handyphone system (PHS), a personal digital assistant (PDA), an international mobile telecommunications (IMT)-2000, a code division multiple access (CDMA)-2000, a wideband code division multiple access (W-CDMA), a wireless broadband (WiBro) terminal, a smart phone, etc., which are wireless communication devices based on various handheld devices, wearable devices, such as watches, glasses, contact lenses, or head-mounted devices (HMDs). That is, the user terminal 10 can include all kinds of structures that can store the learned artificial intelligence model 20 and generate a restored image from the original image captured by the camera 5. Below, the user terminal 10 including the camera 5 will be described as an embodiment (image restoration device).

[0058] Figure 2 is a control block diagram of an imaging restoration device according to an embodiment of the present invention.

[0059] Referring to Figure 2, the imaging restoration device 10 of the present invention may include: a camera 5; a communication unit 11 capable of communicating with the outside; a processor 12 that inputs the original image transmitted through the camera 5 or the communication unit 11 into the completed learning artificial intelligence model 20 to generate a restored image; a memory 13 that stores, in addition to the artificial intelligence model 20, the received original image, the image (correct image) required for learning, or the generated restored image; and an output unit 14 for displaying the original image or the generated restored image.

[0060] Specifically, the camera 5 according to an embodiment of the present invention includes an ultra-small lens and an image sensor, such as a meta-lens or a diffractive optical lens, etc., and can operate under the control of the processor 12 to obtain an original image with strong aberration. The camera 5 converts the original image into an electrical signal and transmits it to the processor 12 or the memory 13.

[0061] The communication unit 11 is configured to receive the original image taken from the outside through the ultra-small lens via a communication network. The communication unit 11 may include one or more components capable of communicating with the outside, and may include at least one of, for example, a short-range communication module, a wired communication module, and a wireless communication module.

[0062] The short-range communication module may include various short-range communication modules that transmit and receive signals within a short range using a wireless communication network, such as a Bluetooth module, an infrared communication module, an RFID (Radio Frequency Identification) communication module, a wireless local area network (WLAN, Wireless Local Access Network) communication module, an NFC communication module, a Zigbee communication module, etc.

[0063] The wired communication module may include various wired communication modules, such as a Controller Area Network (CAN) communication module, a Local Area Network (LAN) module, a Wide Area Network (WAN) module, or a Value Added Network (VAN) module, etc.; and various cable communication modules, such as Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), Digital Visual Interface (DVI), RS-232 (recommended standard 232), power line communication, or plain old telephone service (POTS), etc.

[0064] In addition to the Wi-Fi module and the WiBro (Wireless broadband) module, the wireless communication module may further include wireless communication modules that support various wireless communication methods, such as global System for Mobile Communication (GSM), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), universal mobile telecommunications system (UMTS), Time Division Multiple Access (TDMA), Long Term Evolution (LTE), etc.

[0065] The memory 13 can be implemented by at least one of non-volatile storage elements such as caches, read-only memories (ROMs), programmable ROMs (PROMs), erasable programmable ROMs (EPROMs), electrically erasable programmable ROMs (EEPROMs), flash memories; or volatile storage elements such as random access memories (RAMs); or storage media such as hard disk drives (HDDs), CD-ROMs, but is not limited thereto.

[0066] The output unit 14 is a display for displaying the original image or the restored image, and may include a digital light processing (DLP) panel, a plasma display panel, a liquid crystal display (LCD) panel, an electro-luminescence (EL) panel, an electrophoretic display (EPD) panel, an electrochromic display (ECD) panel, a light emitting diode (LED) panel, an organic light emitting diode (OLED) panel, etc., but is not limited thereto.

[0067] The processor 12 is configured to control the entire imaging restoration device 10. In particular, it trains the artificial intelligence model 20, generates a restored image through the learned artificial intelligence model 20, and displays it through the output unit 14. Additionally, the processor 12 can provide the generated restored image to a wearable device carried by the user through the communication unit 11.

[0068] Specifically, the artificial intelligence model 20 can be divided into an inpainting module 21 and a discriminant module 22. Among them, the inpainting module 21 crops the original image with a preset block (Crop), generates a repaired image based on the input data embedding the coordinate information of the block in the cropped image block, and the discriminant module 22 performs a Fourier transform on the repaired image generated by the inpainting model 21 and differentiates the repaired image after the Fourier transform. The processor 12 improves the performance of the inpainting model 21 by performing adversarial learning of the inpainting model 21 and the discriminant model 22, and then repairs the original image received from the camera 5 or the communication unit 11. The following specific description of the other attached drawings relates to the method for the processor 12 to train the artificial intelligence model 20 and the inpainting method executed by the trained artificial intelligence model 20.

[0069] The processor 12 may refer to a data processing device built into the hardware, which has a physical-structured circuit to execute the functions expressed by the codes or instructions included in the program. An example of the data processing device built into the hardware may include a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, and a processing device such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), and a Graphics Processing Unit (GPU), but is not limited thereto. The processor 12 may be provided in a plurality of structures including one or more processors.

[0070] In addition, the imaging inpainting device 10, in addition to the above-described structure in Figure 2 described, may further include various structures, and some specific structures may be omitted as needed.

[0071] Figure 3 is the overall flowchart of the inpainting method according to an embodiment of the present invention.

[0072] Referring to Figure 3 , for the inpainting method, first, the artificial intelligence model 20 (100) is trained.

[0073] The artificial intelligence model 20 learns through adversarial learning. Specifically, in the artificial intelligence model, the image inpainting model 21 functions as an adversarial learning generator that generates an inpainted image based on the original data (learning data) input by the user. The discriminative model 22 functions as a discriminator that performs adversarial learning by comparing the Fourier-transformed inpainted image generated by the image inpainting model 21 with the correct data after performing a Fourier transform on the inpainted image. Here, the correct data is a clear-quality image that the user wants to inpaint.

[0074] The image inpainting method of the present invention can enable the entire artificial intelligence model 20 to learn by continuously updating (learning) the image inpainting model 21 and the discriminative model 22, such that the image inpainting model 21 can generate a high-frequency inpainted image from a low-frequency original image.

[0075] When the artificial intelligence model 20 has finished learning, the camera 5 captures a new object (200).

[0076] Here, the image captured and transmitted by the camera 5 is not the original image for the learning of the artificial intelligence model 20, but a new original image of the captured object Ob.

[0077] The image inpainting method inputs the original image captured by the camera 5 into the artificial intelligence model 20 that has completed learning (300).

[0078] Specifically, the original image captured by the camera 5 is only input into the image inpainting model 21 that has completed learning, and the image inpainting model 21 that has completed learning generates a high-frequency inpainted image.

[0079] The image inpainting method outputs the inpainted image generated (output) from the artificial intelligence model 20 (400).

[0080] There are many ways to output the inpainted image generated by the image inpainting method. When the image inpainting method is implemented in the imaging inpainting device 10 including a display, the inpainted image can be displayed to the user through the output unit 14. When the imaging inpainting device 10 is provided on a structure that does not include the output unit 14 (for example, a drone without a display, etc.), the imaging inpainting device 10 transmits the generated inpainted image to a wearable device such as a smartphone through the communication unit 11.

[0081] Figure 4 It is a diagram for specifically illustrating an artificial intelligence model according to an embodiment of the present invention.

[0082] Refer to Figure 4, the following process is executed in the artificial intelligence model 20, that is, the image inpainting model 21 generates an inpainted image, and the discriminative model 22 distinguishes authenticity by comparing the inpainted image with the correct image. Based on the discrimination result of the discriminative model 22, the image inpainting model 21 that has adjusted the loss function generates an inpainted image again in an iterative process, that is, adversarial learning.

[0083] The artificial intelligence model 20 is characterized in that the image inpainting model 21 operates in the spatial domain, and the discriminative model 22 operates in the Fourier domain.

[0084] Specifically, the image inpainting model 21 learns the original image as an inpainted image based on the image patches 211. The image inpainting model 21 according to an embodiment of the present invention uses image patches 211 randomly cropped from the original image (learning data) to overcome the physical capacity limitation of the GPU provided in the image inpainting device 10. That is, the image patches are multiple images for cutting the original image into patches of a user-set size.

[0085] In addition, existing image inpainting methods using deconvolution also use randomly cropped image patches. However, the image inpainting model 21 of the present invention generates coordinate information 212 of the image patches 211 in order to improve the inpainting performance of the original image with optical aberration, and embeds the generated coordinate information 212 into the image patches 211.

[0086] Specifically, the image inpainting model 21 generates coordinate information 212 based on the patch size of the image patches 211. When generating coordinate information for each patch, the image inpainting model 21 converts each coordinate information in the image patches 211 into two-dimensional coordinate data (2D coordinate data) by the meshgrid method.

[0087] In order to convert to two-dimensional coordinate data, the image inpainting model 21 includes a 1×1 convolutional layer 213. The image inpainting model 21 passes the generated coordinate information 212 and the image patches 211 through the 1×1 convolutional layer 213, thereby creating image patches 211 combined with the coordinate information 212, that is, input data 214. Below, the positional embedding method performed by the image inpainting model 21 will be described in detail with other drawings.

[0088] The image inpainting model 21 inputs the input data 214 into the first neural network 215. The first neural network 215 is a neural network in which the size of the input data 214 is the same as the size of the output data, that is, the inpainted image. Figure 4The first neural network 215 in [description] is illustrated as representing a skip connection, which is a basic concept of NAFNet (Nonlinear Activation Free Network). However, the first neural network 215 is not necessarily limited to NAFNet and may also include a Convolution Neural Network (CNN) model including at least one of MIRNet or MPRNet, or a Transformer model including at least one of Restormer or Uformer.

[0089] When the image inpainting model 21 generates an inpainted image, the discriminative model 22 performs the Fourier transform 221 of the inpainting model.

[0090] The original image damaged by optical aberration usually loses many high-frequency components. Therefore, the discriminative model 22 of the present invention performs adversarial learning on the inpainted image through domain conversion, so that the learned image inpainting model 21 can generate high-frequency information well based on low-frequency information.

[0091] The Fourier-transformed inpainted image is input into the second neural network 222. Here, the discriminative model 22 compares the inpainted image with the original image and outputs whether the inpainted image is Real or Fake. The second neural network 222 may include a CNN model including at least one of GoogleNet, AlexNet, and VGG network.

[0092] Figure 5 and Figure 6 are diagrams for explaining the method of generating and combining coordinate information. Below, to avoid repeated description, they will be described together.

[0093] First, referring to Figure 5 , the image inpainting method of the present invention crops the original image (111) in block units.

[0094] The original images included in the learning data are cropped into blocks of a preset size (the same size) within the same batch, and the positions are randomly cropped.

[0095] The image inpainting method of the present invention generates coordinate information (112) based on the middle pixel of the image block 211.

[0096] In Figure 6 's embodiment, the image block 211 can be cropped at a specific position of the original image 201. The image inpainting method can generate the value of the midpoint (x, y,) as the coordinate information 212 based on the pixel (0, 0) in the upper left corner of the image block 211.

[0097] The image inpainting method converts the coordinate information 212 into two-dimensional coordinate data (113) through the Meshgrid method.

[0098] The Meshgrid method is a method that returns two-dimensional (or three-dimensional) grid coordinates based on the coordinates included in vector x and vector y. The image inpainting method of the present invention converts the coordinate information 212 of the image patch 211 into two-dimensional coordinate data, and then the image patch 211 and the coordinate information 212 can be combined through the 1×1 convolutional layer 213. Figure 6 The shown Meshgrid method involves an example of converting specific coordinates (x, y) into grid coordinate data.

[0099] Refer again to Figure 5 , and through the 1×1 convolutional layer 213, the coordinate information 212 (114) converted into two-dimensional data is combined with the image patch 211.

[0100] Here, the combination of the image patch 211 and the coordinate information 212 can be generated by inputting the image patch 211 and the coordinate information 212 into the 1×1 convolutional layer 213 in a manner of combining a specific sequence.

[0101] The image inpainting method generates input data 214 by combining the image patch 211 and the coordinate information 212, and based on this, generates an inpainted image through the first neural network 215. Here, the sizes of the input data and the inpainted image output through the first neural network 215 are the same.

[0102] Figure 7 is a flowchart for specifically illustrating the method of performing learning in the image inpainting method according to an embodiment of the present invention. Figure 8 is a diagram for illustrating the discrimination process according to an embodiment. To avoid repeated description, it will be described together below.

[0103] First, refer to Figure 7 , the image inpainting method embeds the coordinate information 212 into the image patch 211 (110), and the image inpainting model 21 generates an inpainted image (120).

[0104] It has been described in Figure 5 and Figure 6 the method of generating the coordinate information 212 and combining it with the image patch 211 to generate input data (or learning data 214), so it is omitted here.

[0105] When the image inpainting model 21 generates an inpainted image, the discrimination model 22 performs a Fourier transform on the inpainted image (130) and discriminates the Fourier-transformed inpainted image (140).

[0106] Refer to Figure 8, according to an embodiment, the discrimination model 22 can control the result output from the second neural network 222 to be output as a preset reference value.

[0107] Specifically, when the restored image by Fourier transform is restored well enough to be judged as a correct image, the second neural network 222 can output a value of 1 or more. If it is judged that the restored image by Fourier transform is not restored like a correct image, a value of -1 or less can be output. That is, the reference values (1: first reference value, 2: second reference value) set as the values output by the second neural network 222 can be changed by the user. In addition, Figure 8 The ranges of the shown decision boundary and margin can be set in various ways.

[0108] Refer to again Figure 7 , the image restoration method determines whether learning has been completed by learning data. If learning has not been completed (No in 150), the image restoration method adjusts the loss function of the artificial intelligence model 20, and then executes the image restoration step (120) again to continue learning.

[0109] The artificial intelligence model 20 of the present invention is updated in the direction of reducing the loss function, and the overall loss function (L Total ) of the artificial intelligence model 20 of the present invention is expressed as Equation 1 below.

[0110] [Equation 1]

[0111] L Total = L PSNR + λL a

[0112] Here, λ is a hyperparameter preset by the user and can be changed in various ways.

[0113] Specifically, L PSNR is the loss function between the restored image (restoration result) after Fourier transform and the correct image. L PSNR is defined by Equation 2 below.

[0114] [Equation 2]

[0115]

[0116] Here, X^ is the result restored from the image restoration model 21, and X represents the correct image. R is the maximum pixel value of the correct image X^, and MSE is an index representing the distance between X^ and X, which can be defined by Equation 3 below.

[0117] [Equation 3]

[0118]

[0119] In Equation 1, L a is the loss function (L G ) of the image inpainting model 21 and the loss function (L D ) of the discriminative model 22, and is defined by the following Equation 4.

[0120] [Equation 4]

[0121]

[0122] The image inpainting model learns in the direction that minimizes the sum of the loss function (L G ) of the image inpainting model 21 and the loss function (L D ) of the discriminative model 22.

[0123] The loss function (L D ) of the discriminative model 22 is defined by the following Equation 5, and the loss function (L G ) of the image inpainting model 21 is defined by the following Equation 6.

[0124] [Equation 5]

[0125]

[0126] [Equation 6]

[0127]

[0128] Here, F() represents the Fourier transform, and D() represents the output of the discriminative model 22.

[0129] In addition, when the learning is completed (yes in 150), the image inpainting method generates a new inpainted image (401) from the image inpainting model 21 that has completed learning.

[0130] Here, the new inpainted image is the inpainted image output by the image inpainting model 21 that has completed learning from the original image newly captured by the camera 5.

[0131] Figure 9 is an example diagram of the result of the imaging system of the present invention.

[0132] The imaging system 1 can acquire an image with strong aberration, such as Figure 9 the original image 302. The stronger the aberration, the more high-frequency components disappear in the original image 302.

[0133] The imaging system 1 of the present invention is based on Figure 9The correct image 301 is used to train the discrimination model 22, and the image restoration model 21 is trained through adversarial learning. The trained imaging system 1 can output a restored image 304 of the original image 302 through the trained image restoration model 21.

[0134] When comparing the image 303 restored using traditional techniques with the restored image 304, the image 303 restored using a traditional general deep learning model has no damage to the restoration edge. However, it can be seen that the restored image 304 restored using the imaging system 1 of the present invention repairs the high-frequency region more clearly than the image 303 restored using traditional techniques.

[0135] Figure 10 It is a table showing the performance comparison results of the image restoration method of the present invention.

[0136] Figure 10 The table in [ ] compares the currently disclosed image restoration models (such as MIRNet V2, SFNet, HINet, NAFNet) with the image restoration method of the present invention. The performance of each artificial intelligence model is compared using Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Map (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS).

[0137] Specifically, the image restoration method of the present invention shows a PSNR of 22.095 / 2.423, which can be confirmed to be higher than that of traditional models. The SSIM and LPIPS are 0.692 / 0.103 and 0.432 / 0.096 respectively, which can be confirmed to be higher than that of traditional models.

[0138] Thereby, the image restoration method of the present invention and the imaging system for executing the method can solve the image quality degradation in a ultra-small imaging system by using an artificial intelligence model learned using image data directly captured by a lens with strong aberration.

[0139] In addition, the image restoration method of the present invention and the imaging system for executing the method can repair image damage in the high spatial frequency domain by directly learning clear correct data in the frequency domain. Therefore, the image restoration method of the present invention and the imaging system for executing the method can solve the phenomenon of the protruding camera of a smartphone when installed on a wearable device such as a smartphone, and can capture high-quality images / videos.

[0140] In addition, the image restoration method and the imaging system for implementing the method according to the present invention also exhibit excellent performance in downstream applications that can be performed on images captured by cameras provided in unmanned aerial vehicles and AR (Augmented Reality) / VR (Virtual Reality) devices, based on improved image / video quality.

[0141] The image restoration method and the imaging system for implementing the method according to the present invention can achieve the lightweight of the lens, thereby reducing the weight of the mounted imaging system. It is expected to reduce the weight of unmanned aircraft, drones, etc. Thus, an improvement in power efficiency can also be expected. In addition, the image restoration method and the imaging system for implementing the method according to the present invention can reduce the weight of the device itself by installing a super-small imaging system in the imaging system that acts as eyes on the AR / VR device, and an improvement in power efficiency is expected. Moreover, by reducing the fatigue of the user, the user experience can be improved.

[0142] In addition, embodiments of the present invention can be implemented in the form of a recording medium storing instructions executable by a computer. The instructions can be stored in the form of program code and, when executed by a processor, can create program modules to perform the operations of the embodiments of the present invention. The recording medium can be implemented by a computer-readable recording medium.

[0143] Computer-readable recording media include all types of recording media storing instructions interpretable by a computer. For example, there can be read-only memory (ROM), random access memory (RAM), magnetic tapes, magnetic disks, flash memories, optical data memories, etc.

[0144] As described above, embodiments of the present invention have been described with reference to the accompanying drawings. Those of ordinary skill in the art to which the present invention pertains should understand that the present invention can be implemented in forms different from those of the embodiments of the present invention without changing the technical idea or essential features of the present invention. The embodiments of the present invention are illustrative and should not be construed as restrictive.

Claims

1. An imaging repair device, wherein: include: Memory for storing raw images captured from the camera and the AI ​​model; as well as A processor for training the artificial intelligence model, The artificial intelligence model comprises: An image restoration model, which crops the original image with a preset block and generates a restored image based on input data of embedding coordinate information of the block in the cropped image block; as well as A discriminant model is used to perform Fourier transform on the inpainted image generated by the image inpainting model and distinguish the inpainted image after the Fourier transform. The processor performs adversarial learning of the image restoration model and the discriminant model, and generates a restoration image of a new original image taken by the camera from the image restoration model that has completed learning.

2. The imaging repair device according to claim 1, wherein: The processor generates the coordinate information based on the middle pixel of the image block, The coordinate information is converted into two-dimensional coordinate data by a grid method, The input data is generated by combining the coordinate information converted into the two-dimensional data with the image block through a 1×1 convolution layer.

3. The imaging repair device according to claim 2, wherein: The processor inputs to a first neural network having the same size of input data and output data output from the 1×1 convolutional layer, and outputs the restored image from the first neural network.

4. The imaging repair device according to claim 3, wherein: The first neural network includes a CNN model including at least one of MIRNet, MPRNet or NAFNet, or a Transformer model including at least one of Restormer or Uformer.

5. The imaging repair device according to claim 3, wherein: The discriminant model includes a CNN model for discriminating the authenticity of a Fourier transformed restored image, wherein the CNN model includes at least one of a GoogleNet, an AlexNet, and a VGG network. The processor compares the Fourier transformed repaired image with the correct image through the discriminant model, and adversarially trains the image repair model based on the comparison result, so that the image repair model generates a high-frequency repaired image from a low-frequency original image.

6. The imaging repair device according to claim 5, wherein: The processor trains the discriminant model to output a value above a preset first reference value for the correct image, and trains the discriminant model to output a value below a preset second reference value for the restored image.

7. An imaging system, wherein: include: A camera including a superlens; A memory, receiving the original image captured by the camera and storing the artificial intelligence model; A processor, configured to train the artificial intelligence model; as well as a display, configured to output a repaired image of the original image repaired by the processor using the learned artificial intelligence model, The artificial intelligence model comprises: An image restoration model, which crops the original image with a preset block and generates a restored image based on input data of embedding coordinate information of the block in the cropped image block; as well as The discriminant model performs Fourier transform on the inpainted image generated by the image inpainting model and distinguishes the inpainted image after the Fourier transform.

8. The imaging system of claim 7, wherein: The processor generates the coordinate information based on the middle pixel of the image block, The coordinate information is converted into two-dimensional coordinate data by a grid method, The input data is generated by combining the coordinate information converted into the two-dimensional data with the image block through a 1×1 convolution layer.

9. The imaging system of claim 8, wherein: The processor inputs to a first neural network having the same size of input data and output data output from the 1×1 convolutional layer, and outputs the restored image from the first neural network.

10. The imaging system of claim 9, wherein: The processor compares the Fourier transformed repaired image with the correct image through the discriminant model, and adversarially trains the image repair model based on the comparison result, so that the image repair model generates a high-frequency repaired image from the low-frequency original image.

11. The imaging system of claim 10, wherein: The processor trains the discriminant model to output a value above a preset first reference value for the correct image, and trains the discriminant model to output a value below a preset second reference value for the restored image.

12. An image restoration method, wherein: include: Training AI models; A camera including a superlens or a diffractive optical lens acquires a raw image; Generate a restored image of the original image through the learned artificial intelligence model, The artificial intelligence model comprises: An image restoration model that crops an image of the learning data with a preset block, and generates a restoration image based on input data of embedding coordinate information of the block in the cropped image block; as well as A discriminant model is used to perform Fourier transform on the inpainted image generated by the image inpainting model and distinguish the inpainted image after the Fourier transform. The training artificial intelligence model includes comparing the Fourier transformed repaired image with the correct image through the discriminant model, and adversarially training the image repair model based on the comparison result, so that the image repair model generates a low-frequency original image into a high-frequency repaired image.

13. The image restoration method according to claim 12, wherein: The image restoration model generates the restoration image, including: Generating the coordinate information based on the middle pixel of the image block; Converting the coordinate information into two-dimensional coordinate data by a grid method; The input data is generated by combining the coordinate information converted into the two-dimensional data with the image block through a 1×1 convolution layer.

14. The image restoration method according to claim 13, wherein: The image restoration model generates the restoration image, including inputting to a first neural network whose input data output from the 1×1 convolutional layer has the same size as the output data, and outputting the restoration image from the first neural network.

15. The image restoration method according to claim 12, wherein: The training artificial intelligence model includes training the discriminant model to output a value above a preset first reference value for the correct image, and training the discriminant model to output a value below a preset second reference value for the repaired image.