A terminal integrated imaging method, electronic device and medium based on deep learning
By combining metal imaging system and image recovery architecture for adversarial learning, the limitations of metal lens imaging system and the loss of DNN model position information are solved, and efficient image reconstruction and resolution improvement are achieved.
Patent Information
- Application Number
- CN202510287440.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing metal lens imaging systems have limitations in terms of focus efficiency, lens diameter and spectral bandwidth, and the DNN-based image recovery model cannot effectively learn position changes, resulting in loss of position information.
The metal imaging system is combined with the image recovery architecture, and adversarial learning is performed through the image recovery module and the discriminator. The end-to-end image reconstruction is performed using metal superlenses and image recovery models, combined with random cropping and position embedding training, and the image recovery model is adjusted using adversarial loss feedback.
It significantly improves image quality, reduces position information loss, enhances the performance of the imaging system, and achieves higher resolution and image fidelity.
Smart Images

Figure CN119784623B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a terminal integrated imaging method, electronic device, and medium based on deep learning. Background Art
[0002] Metasurface lenses, also known as metalenses, are planar lens structures based on metasurface technology. They hold great potential for compact imaging, photography, light detection and ranging (LiDAR), and virtual reality / augmented reality (VR / AR) applications. They can replace multiple components in traditional lens systems, achieving the same optical performance while significantly reducing system size and weight.
[0003] The relentless pursuit of miniaturization and performance enhancements in optical imaging systems has driven researchers in the field of optics to continuously explore breakthroughs in conventional geometric lens technology. A high-throughput, large-aperture metal lens, produced by combining deep ultraviolet immersion lithography with wafer-level nanoimprint lithography, is poised to break through the limitations of traditional lenses and usher in an era of compact, efficient imaging systems. While metal lenses have achieved remarkable results in focusing efficiency, lens diameter, and spectral bandwidth, they are significantly affected by chromatic aberration due to the inherent angular dispersion in their meta-atom design. Furthermore, even the most ideal metal lens cannot simultaneously achieve broadband operation and large diameter due to physical limitations.
[0004] With the continuous development of modern optical technology, researchers can design Pancharatnam-Berry (PB) phase-based metal lenses composed of arrays of nanostructures with arbitrary rotation angles. These lenses can achieve diffraction-limited focusing, but are susceptible to angular aberrations. With the continuous development of computer technology and machine learning, computers have been able to enhance non-ideal images. Achieving higher resolution through linear deconvolution is the most classic traditional image restoration method. Compared with traditional image restoration methods, DNN-based image restoration models have excellent performance in image restoration tasks. However, since they cannot learn image degradation with position changes, using randomly cropped patches to train the model will result in the loss of position-related information. Summary of the Invention
[0005] The purpose of the present invention is to provide a terminal integrated imaging method, electronic device and medium based on deep learning to solve the above problems.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A deep learning-based end-integrated imaging method comprises the steps of: using a metal imaging system including a metal superlens to acquire an image of a real scene, and inputting the image into an image restoration architecture to obtain a restored output image;
[0008] The image restoration architecture includes at least an image restoration module and a discriminator; the image restoration module is configured to reconstruct the input image through the image restoration model to obtain a reconstructed image; the discriminator is configured to perform adversarial learning on the frequency domain image obtained after fast Fourier transform of the reconstructed image and the real image respectively, to obtain an adversarial loss between the discriminator and the image restoration model; the adversarial loss is fed back to the image restoration model to train it.
[0009] Furthermore, the metal imaging system includes a left-handed circular polarizer, a metal superlens, a right-handed circular polarizer, and an image sensor arranged in sequence; the left-handed circular polarizer and the right-handed circular polarizer are represented to only generate left-handed circularly polarized light or right-handed circularly polarized light from the light entering the metal imaging system for transmission to the metal superlens.
[0010] Furthermore, the left-handed circular polarizer and the right-handed circular polarizer are represented as stacks of a linear polarizer and a quarter-wave plate, respectively.
[0011] Furthermore, the metallic superlens uses nanoplates with arbitrary rotation angles to make it a superlens based on the PB phase principle.
[0012] Furthermore, the metallic superlens was fabricated by nanoimprint lithography and atomic layer deposition.
[0013] Furthermore, the point spread function of the metal superlens is measured to obtain the modulation transfer function, which is expressed as:
[0014] ;
[0015] Where x and y are the horizontal and vertical positions on the image sensor, is the measured point spread function; and are the spatial frequencies along the x and y axes, respectively.
[0016] Furthermore, the point spread function is measured by capturing images of red, green, and blue collimated beams using a metal imaging system.
[0017] Furthermore, when training the image restoration model, the output image of the metal imaging system is randomly cropped and positionally embedded and then input into the image restoration model. The peak signal-to-noise ratio loss between the reconstructed image and the real image is calculated and passed to the image restoration model.
[0018] The present invention also provides an electronic device, which includes a processor and a storage medium communicatively connected to the processor, the storage medium is suitable for storing multiple instructions, and the processor is suitable for calling the instructions in the storage medium to execute the steps of implementing the above-mentioned deep learning-based terminal integrated imaging method.
[0019] The present invention also provides a storage medium storing a plurality of instructions, which are suitable for being loaded by a processor and executing the steps of the above-mentioned deep learning-based terminal integrated imaging method.
[0020] Compared with the prior art, the present invention has at least the following beneficial effects:
[0021] (1) The traditional imaging system is reorganized into a deep learning-driven imaging method that combines a metal imaging system and an image restoration architecture. The metal imaging system is responsible for acquiring images, while the image restoration architecture is responsible for restoring the captured images. The metal imaging system uses a hyperplane lens to achieve stronger performance, while the image restoration architecture uses an end-to-end image restoration framework to address the shortcomings of traditional DNN methods that cause loss of position information.
[0022] (2) Using randomly cropped images as the input of the restoration model, the loss between the reconstructed image and the real image represents the image fidelity loss, and this loss is fed back to the image restoration model for adjustment to reduce the difference between the output of the restoration model and the real image;
[0023] (3) The reconstructed image and the real image are fast Fourier transformed and input into the auxiliary discriminator for adversarial learning. At the same time, the adversarial loss between the discriminator and the image restoration model is fed back to the image restoration model using a GAN training scheme with hinge loss. By training the image restoration model, the most reasonable loss function can be found, which can significantly improve the image quality produced by the metal imaging system. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0025] Figure 1 This is a flowchart of the overall implementation of the imaging method in the embodiment provided by the present invention;
[0026] Figure 2 is a schematic diagram of a metal imaging system according to an embodiment of the present invention;
[0027] Figure 3 is a schematic diagram of an optical device for point spread function (PSF) measurement in an embodiment provided by the present invention;
[0028] Figure 4 is a schematic diagram of an image restoration architecture in an embodiment provided by the present invention;
[0029] Figure 5 Schematic diagram of comparing a real image, an image captured by a metal lens, and a restored image in an embodiment provided by the present invention. DETAILED DESCRIPTION
[0030] It should be noted that, in the present invention, unless otherwise expressly specified or limited, the terms "connection" and "fixation" should be understood in a broad sense. For example, "fixation" can mean fixed connection, detachable connection, or integration; it can mean mechanical connection or electrical connection; it can mean direct connection or indirect connection through an intermediate medium; it can mean internal communication between two elements or interaction between two elements, unless otherwise expressly specified. Those skilled in the art will be able to understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0031] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that ordinary technicians in this field can implement it. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0032] The following are specific embodiments of the present invention, and the technical solutions of the present invention are further described in conjunction with the accompanying drawings, but the present invention is not limited to these embodiments.
[0033] like Figure 1 As shown, this embodiment provides a terminal integrated imaging method based on deep learning, including the steps of:
[0034] S1. Acquire images of real scenes using a metal imaging system including a metal superlens;
[0035] S2. Input the image into the image restoration architecture to obtain a restored output image.
[0036] In this embodiment, a traditional imaging system is restructured into a deep learning-driven imaging system. This system combines a metal imaging system with an image restoration architecture. The metal imaging system acquires images, while the image restoration architecture restores the captured images. The metal imaging system utilizes a hyperplane lens, which enables enhanced performance, while the image restoration architecture utilizes an end-to-end image restoration framework to address the positional information loss associated with traditional DNN approaches.
[0037] like Figure 2 As shown, in the metal imaging system, it includes a left-handed circular polarizer, a metal superlens, a right-handed circular polarizer and an image sensor arranged in sequence.
[0038] The polarizer can be represented as a stack of a linear polarizer (LP) and a quarter-wave plate (QWP). Light entering the metal imaging system first passes through a left-hand circular polarizer, which transmits only left-hand circularly polarized light (LCP) to the metal superlens. The metal lens focuses the light using the polarization-dependent transmittance of the meta-atom and transmits it to a right-hand circular polarizer. Finally, the linearly polarized light transmitted from the right-hand circular polarizer is focused on the image sensor. The right-hand circular polarizer transmits only right-hand circularly polarized light (RCP) and blocks the LCP light to remove unfocused light.
[0039] The metal superlens is a new type of super-flat lens. It uses a nanoplate with an arbitrary rotation angle to make it a superlens based on the PB phase principle. By measuring its point spread function (PSF), the modulation transfer function (MTF) passing through the metal superlens is calculated.
[0040] Furthermore, the metal superlens used in the metal imaging system is created by applying a conventional imprint resin (MINS-311RM) to a specific mold and covering it with a glass wafer. After curing the imprint resin under a predetermined pressure and UV irradiation, the mold is removed, forming a metal pattern on the glass wafer. Finally, a high-refractive-index titanium dioxide film is coated on the imprinted metal using atomic layer deposition to create the metal superlens required for the metal imaging system.
[0041] like Figure 3 As shown, when measuring its point spread function, the red light, green light or blue light emitted by the light-emitting diode with peak intensity of 640nm, 525nm and 457nm is spatially filtered and collimated by a 0.5-inch aspheric lens, a 20-micron pinhole and a 1-inch spherical lens arranged in sequence. When the optical axis of the spatial filter and the metal lens imaging system are matched, the PSF can be measured, which is expressed as .
[0042] The modulation transfer function describes the image quality in terms of resolution and contrast. The MTF can be calculated from the measured PSF:
[0043] ;
[0044] Where x and y are the horizontal and vertical positions on the image sensor, is the measured point spread function; and are the spatial frequencies along the x and y axes, respectively.
[0045] When using a metal imaging system, aberrations can occur due to the metal superlens being affected by chromatic and angular aberrations, and the manufacturing process being affected by surface defects. Therefore, this embodiment employs a deep learning-based image restoration architecture to address the aberration issues in images captured by the metal imaging system.
[0046] like Figure 4 As shown, the image restoration architecture includes at least an image restoration module and a discriminator. The image restoration module is configured to reconstruct the input image through the image restoration model to obtain a reconstructed image. The discriminator is configured to perform adversarial learning on the frequency domain images obtained by performing fast Fourier transforms (FFTs) on the reconstructed image and the real image, respectively, to obtain an adversarial loss between the discriminator and the image restoration model. The adversarial loss is fed back to the image restoration model to train it. By training the image restoration architecture, the most reasonable loss function can be found, significantly improving the image quality produced by the metal imaging system.
[0047] In this embodiment, the width of the initial layer of the image restoration model network is set to 32, and the width of each successive block is doubled as the network goes deeper. The reconstructed image is calculated based on the output of the restoration model. , compared with the ground truth image The difference between the two is used to calculate the peak signal-to-noise ratio (PSNR) loss of the restoration model. , and will eventually lose Passed to the image restoration model in the image restoration architecture to reduce the difference between the output of the restoration model and the true image.
[0048] At the same time, the output image of the metal imaging system is cropped, and the cropped coordinate information is used to apply random cropping and position embedding to the input data, so that the model network randomly crops patches from the full-resolution image for training.
[0049] In the image restoration stage, the reconstructed image Compared with the ground truth image The difference between the calculated peak signal-to-noise ratio loss for:
[0050] ;
[0051] in, 、 and R represent the reconstructed image, the true image, and the maximum value of the true image, respectively.
[0052] Then, when training the image restoration model, after expressing the reconstructed image and the real image in the spatial frequency domain through fast Fourier transform (FFT), the discriminator is used to perform adversarial learning to enhance the ability of the image restoration architecture to recover the lost information.
[0053] In this embodiment, the spatial data from the reconstructed image and the real image channels are converted into the frequency domain by fast Fourier transform (FFT), which can effectively solve the occurrence of pattern artifacts while restoring high-frequency information.
[0054] In the process of adversarial learning using the discriminator, the width of the discriminator is set to 64, and all layers are set to the same width. The FFT output is used as the input of the discriminator for adversarial training. Since the data transformed by FFT is very complex, including real and imaginary parts, it can be represented as a two-dimensional vector and then used as the input of the discriminator. Spectral normalization is also applied to improve the stability of training. In the adversarial learning stage, the adversarial loss between the discriminator (D) and the image restoration model (G) is :
[0055] ;
[0056] ;
[0057] Where F refers to FFT, and are operators representing the computation of the average of the true image value and the reconstructed image value in a given mini-batch.
[0058] To verify the effectiveness of this embodiment, the method provided in this embodiment was tested under five target conditions: PSNR, SSIM, LPIPS, MAE, and CS. The method was also compared with four state-of-the-art image restoration models: MIRNetv2, SFNet, HINet, and NAFNet, using the same image test set (n=70). The experimental simulation is as follows:
[0059] The number of iterations during the training of the deep learning image restoration architecture was set to 300,000. In the image restoration model, AdamW was used as the optimizer and the learning rate was initially set to 3×10 -4 , and gradually reduced to 10 according to the cosine annealing schedule -7 In the image restoration model, AdamW is used as the optimizer and the learning rate is initially set to 3×10 -4 , and then gradually decreases to 10 after cosine annealing -7, the coefficient of β is [0.9, 0.9]. For the discriminative optimizer, the learning rate of Adam is set to 3×10 -4 , the same model as recovery, but with the β coefficient set to [0.0, 0.9].
[0060] like Figure 5 As shown in Figure 3, the image captured of the metal is clearly corrupted by chromatic aberration, with a noticeable difference in the sharpness of the green compared to the red and blue components, resulting in significant blurring. In contrast, the image reconstructed by our proposed framework exhibits significant fidelity to the ground truth, demonstrating the framework's proficiency in recovering details obliterated by chromatic aberration.
[0061] The following table shows a qualitative comparison of the four mainstream models used in this example's comparative experiments with the aforementioned methods. Our framework achieves a 35.6% reduction in LPIPS, indicating a significant increase in the perceptual similarity between the reconstructed image and its original counterpart. Furthermore, it significantly outperforms the state-of-the-art models in terms of PSNR, SSIM, MAE, and CS.
[0062]
[0063] This embodiment also provides an electronic device, which includes a processor and a storage medium communicatively connected to the processor, the storage medium is suitable for storing multiple instructions, and the processor is suitable for calling the instructions in the storage medium to execute the method provided by this embodiment.
[0064] This embodiment also provides a storage medium, which stores a plurality of instructions. The instructions are suitable for being loaded and executed by a processor to enable the above-mentioned electronic device to execute the method provided by this embodiment.
[0065] The computer programs for implementing the methods of the embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer programs are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0066] In the context of the present embodiment, a storage medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0067] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.
Claims
1. A terminal integrated imaging method based on deep learning, characterized in that: The method comprises the steps of: acquiring an image of a real scene using a metal imaging system including a metal superlens, and inputting the image into an image restoration architecture to obtain a restored output image; The image restoration architecture includes at least an image restoration module and a discriminator; the image restoration module is configured to reconstruct an input image through an image restoration model to obtain a reconstructed image; the discriminator is configured to perform adversarial learning on a frequency domain image obtained by performing fast Fourier transform on the reconstructed image and the real image, respectively, to obtain an adversarial loss between the discriminator and the image restoration model; the adversarial loss is fed back to the image restoration model to train it; The adversarial loss is expressed as: ; Where F refers to FFT, and are operators representing the computation of the average of the true image value and the reconstructed image value in a given mini-batch.
2. The deep learning-based terminal integrated imaging method according to claim 1, characterized in that: The metal imaging system includes a left-handed circular polarizer, a metal superlens, a right-handed circular polarizer, and an image sensor arranged in sequence; the left-handed circular polarizer and the right-handed circular polarizer are shown to only generate left-handed circularly polarized light or right-handed circularly polarized light from light entering the metal imaging system for transmission to the metal superlens.
3. The terminal integrated imaging method based on deep learning according to claim 2, characterized in that: The left-handed circular polarizer and the right-handed circular polarizer are represented as stacks of a linear polarizer and a quarter-wave plate, respectively.
4. The method for terminal integrated imaging based on deep learning according to claim 2, characterized in that: The metal superlens adopts a nanoplate with an arbitrary rotation angle to become a superlens based on the PB phase principle.
5. The method for terminal integrated imaging based on deep learning according to claim 4, characterized in that: The metal superlens is manufactured by nanoimprint lithography and atomic layer deposition.
6. The method for terminal integrated imaging based on deep learning according to claim 2, characterized in that: The modulation transfer function of the metal superlens is obtained by measuring its point spread function, and the modulation transfer function is expressed as: ; Where x and y are the horizontal and vertical positions on the image sensor, is the measured point spread function; and are the spatial frequencies along the x and y axes, respectively.
7. The terminal integrated imaging method based on deep learning according to claim 6, characterized in that: The point spread function is measured by capturing images of red, green, and blue collimated light beams by the metal imaging system.
8. The terminal integrated imaging method based on deep learning according to claim 1, characterized in that: When training the image restoration model, random cropping and position embedding are applied to the output image of the metal imaging system and then input into the image restoration model, and the peak signal-to-noise ratio loss between the reconstructed image and the real image is calculated to pass it to the image restoration model.
9. An electronic device, characterized in that: The electronic device includes a processor and a storage medium communicatively connected to the processor, the storage medium is suitable for storing multiple instructions, and the processor is suitable for calling the instructions in the storage medium to execute the steps of implementing the deep learning-based terminal integrated imaging method described in any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium stores a plurality of instructions, which are suitable for being loaded by a processor and executing the steps of the deep learning-based terminal integrated imaging method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Ultra-compact spectral light field camera system based on super-structured lens array
CN111426381A
Assembly for detecting the intensity distribution of components of the electromagnetic field in beams of radiation
CN112789506A
Near-to-eye display optical system for AR glasses and AR glasses
CN115308903A
Unmanned aerial vehicle image blind super-resolution reconstruction method based on frequency domain residual error
CN116563101A