Image fusion method based on mixed multi-order diffraction lens, medium and program product

By employing an image fusion method based on hybrid multi-order diffraction lenses, the sharpness of infrared and visible light images generated by HMODL is restored using a restoration network. Furthermore, the diffraction characteristics are simulated in the fusion network to perform image fusion, thus solving the problems of blurring and loss of detail in HMODL images and achieving high-quality dual-band image fusion results.

CN121837040APending Publication Date: 2026-04-10INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional compact dual-band imaging optical systems are complex, heavy, and costly. Images generated by hybrid multi-order diffraction lenses (HMODL) suffer from blurring and loss of detail, affecting image clarity and cross-band information utilization efficiency, and limiting the reliability of image fusion results.

Method used

An image fusion method based on hybrid multi-order diffraction lenses is adopted. By acquiring infrared and visible light image data, a restoration network is used to recover clear images. The diffraction characteristics are simulated in the fusion network to perform image fusion, and training samples that conform to the diffraction imaging characteristics are constructed to improve image quality and information integrity.

Benefits of technology

While retaining the compact imaging advantages of HMODL, it significantly improves the detail and information integrity of dual-band imaging, enhancing the robustness and application value of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837040A_ABST
    Figure CN121837040A_ABST
Patent Text Reader

Abstract

The invention discloses an image fusion method based on a mixed multi-order diffraction lens, a medium and a program product, and relates to the technical field of image processing and optical engineering. First diffraction imaging data output by a mixed multi-order diffraction lens are obtained and input into a restoration network to obtain restored second diffraction imaging data, so that the defects of diffraction imaging are overcome. Furthermore, the image fusion network fully considers the scene of inherent degradation characteristics of the diffraction image during model training, and constructs a training sample more conforming to diffraction imaging characteristics, so that a stable and high-quality fusion result can be obtained by inputting the second diffraction imaging data into the fusion network. According to the two-stage fusion strategy provided by the invention, while the advantage of HMODL compact imaging is reserved, the detail representation and information integrity of dual-band imaging can be remarkably improved, and meanwhile, the robustness and application value of the system under the actual diffraction imaging condition are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of image processing and optical engineering technology, specifically to the field of infrared and visible light image fusion methods, and particularly to an image fusion method, medium, and program product based on hybrid multi-order diffraction lenses. Background Technology

[0002] With the continuous expansion of photoelectric detection demands, modern imaging systems are facing increasingly stringent requirements in terms of size, weight, and imaging bands. Traditional compact dual-band imaging optical systems typically employ a front-end Cassegrain structure to compress the optical path and mitigate chromatic aberration, then use a beam splitter to guide the two beams of different bands to independent channels for aberration correction. However, this architecture places a high complexity on the design of the rear mirror group and is highly dependent on precision assembly and adjustment. It is not only structurally complex and heavy, but also has high manufacturing and maintenance costs, making it difficult to meet the requirements of next-generation equipment for lightweight design and multi-band imaging performance.

[0003] To overcome the aforementioned limitations, diffractive lenses have attracted widespread attention due to their thin and light profile, ease of fabrication, and tunable dispersion. Traditional end-to-end diffractive lens designs often rely on single-order or a few-order diffraction structures, sacrificing focusing efficiency for some wavelengths in exchange for multi-wavelength confocal imaging capabilities. While this approach extends the imaging band to some extent, it also leads to problems such as enhanced sidelobe energy and limited effective imaging bandwidth, hindering further improvements in image quality. Hybrid multi-order diffractive lenses (HMODLs), on the other hand, achieve multi-wavelength confocal imaging by constructing multi-order diffraction structures on the front surface of the lens, and designing a corrective diffraction structure on the rear surface to compensate for aberrations and fabrication errors, ultimately achieving monolithic visible and infrared dual-band imaging. HMODLs exhibit superior multi-band imaging performance while maintaining the advantages of a compact and lightweight optical system.

[0004] However, due to diffraction effects and limitations in actual manufacturing processes, images generated by HMODL still exhibit a certain degree of blurring and loss of detail. This degradation not only affects the sharpness of the dual-band images themselves but also reduces the efficiency of subsequent cross-band information utilization. In particular, in dual-band image fusion tasks, information loss or insufficient detail representation frequently occurs, thus limiting the reliability of the fusion results. Summary of the Invention

[0005] To address the technical problem of information loss due to the degradation of output image features under diffraction imaging conditions, which limits the stability and reliability of fusion results, this invention proposes an image fusion method, medium, and program product based on hybrid multi-order diffraction lenses.

[0006] The first aspect discloses an image fusion method based on hybrid multi-order diffraction lenses, the method comprising:

[0007] Acquire first diffraction imaging data output by a hybrid multi-order diffraction lens, the first diffraction imaging data including infrared image and visible light image data pairs;

[0008] The first diffraction imaging data is input into the restoration network for image restoration processing to obtain the restored second diffraction imaging data.

[0009] The second diffraction imaging data is input into the fusion network for image fusion processing to obtain a fused image. The fusion network is trained on a training dataset obtained by simulating the diffraction characteristics of the output image of a multi-order diffraction lens, performing degradation processing, and then restoring it through the restoration network.

[0010] As one possible implementation, the method includes: the restoration network is trained through the following steps:

[0011] Obtain the first training dataset constructed from multiple pairs of infrared and visible light images;

[0012] The first training dataset is input into a preset image degradation model to obtain degraded image data, wherein the image degradation model is constructed based on the response characteristics of a hybrid multi-order diffraction lens to a spatial point source;

[0013] The obtained degraded image data is input into the restoration network, and the restoration network is trained with the target loss function as a condition until the model converges.

[0014] The second aspect discloses an electronic device including a processor and a memory, the memory storing a computer program that, when the computer program is executed, implements the image fusion method based on hybrid multi-order diffraction lenses as disclosed in the first aspect or any possible implementation of the first aspect.

[0015] The third aspect discloses a computer-readable storage medium storing a computer program or computer instructions that, when executed, implement the image fusion method based on hybrid multi-order diffraction lenses as disclosed in the first aspect or any possible implementation thereof.

[0016] The fourth aspect discloses a computer program product that, when run on a computer, causes the computer to perform the image fusion method based on hybrid multi-order diffraction lenses disclosed in the first aspect or any possible implementation of the first aspect.

[0017] As can be seen from the above technical solutions, the beneficial effects of this invention are as follows: by acquiring the first diffraction imaging data output by a hybrid multi-order diffraction lens, the first diffraction imaging data includes infrared image and visible light image data pairs; and inputting the first diffraction imaging data into a restoration network for image restoration processing, a restored second diffraction imaging data is obtained, thereby compensating for the defects of diffraction imaging. Furthermore, the image fusion network of this invention fully considers the inherent degradation characteristics of diffraction images during model training, constructing training samples that better conform to the characteristics of diffraction imaging. Therefore, after inputting the second diffraction imaging data into the fusion network for image fusion processing, a stable and high-quality fusion result can be obtained. The two-stage fusion strategy proposed in this invention, while retaining the compact imaging advantages of HMODL, can significantly improve the detail performance and information integrity of dual-band imaging, while enhancing the robustness and application value of the system under actual diffraction imaging conditions. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the HMODL dual-band imaging system provided by the present invention.

[0019] Figure 2 The flowchart shows the image fusion method based on hybrid multi-order diffraction lenses provided by the present invention.

[0020] Figure 3 The flowchart of the overall network model provided by this invention.

[0021] Figure 4 The images show the original infrared images captured by the HMODL system provided in this invention, as well as the restored infrared images.

[0022] Figure 5 The original visible light image and the restored visible light image captured by the HMODL system provided by this invention.

[0023] Figure 6 The result of fusing infrared and visible light images captured by the HMODL system provided in this invention.

[0024] Figure 7 This invention provides quantitative results based on datasets captured using the HMODL system. Detailed Implementation

[0025] To make the objectives, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although several embodiments of the present invention have been given, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of the present invention will be thorough and complete.

[0026] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0027] To better understand the image fusion method based on hybrid multi-order diffraction lenses disclosed in the embodiments of the present invention, the system architecture used in the embodiments of the present invention will be described below. Figure 1 This is a schematic diagram of the HMODL dual-band imaging system disclosed in an embodiment of the present invention. Figure 1 As shown, the dual-band images to be fused are infrared and visible light images. It should be noted that for the infrared and visible light bands output by HMODL, the infrared image relies on thermal radiation imaging of objects, highlighting the thermal features of prominent targets such as people, buildings, and vehicles. However, it has lower overall resolution, insufficient detail, and is easily affected by system noise. In contrast, the visible light image captures reflected light from objects, resulting in rich detail and texture, high spatial resolution, and strong scene representation capabilities. Therefore, effectively fusing infrared and visible light images can enhance overall detail and visual readability while preserving the thermal features of prominent targets, thereby improving the system's detection performance and application range, and providing a new technical approach for Earth observation.

[0028] In addition, the present invention employs a hybrid multi-order diffraction lens (HMODL), which enables monolithic visible light and infrared dual-band imaging, meeting the equipment requirements for lightweight design and multi-band imaging performance.

[0029] However, due to diffraction effects and limitations in actual manufacturing processes, images generated by HMODL still exhibit a certain degree of blurring and detail loss. This degradation not only affects the clarity of the dual-band image itself but also reduces the efficiency of subsequent cross-band information utilization. Particularly in dual-band image fusion tasks, information loss or insufficient detail representation frequently occurs, thus limiting the reliability of the fusion result. To address this issue, in one embodiment, such as... Figure 2 As shown, this invention discloses an image fusion method based on hybrid multi-order diffraction lenses, specifically including:

[0030] 101. Acquire first diffraction imaging data output through a hybrid multi-order diffraction lens, wherein the first diffraction imaging data includes infrared image and visible light image data pairs.

[0031] Specifically, this invention employs a hybrid multi-order diffraction lens (HMODL) to achieve monolithic visible and infrared dual-band imaging. The infrared and visible images output by the HMODL system are acquired, and the two images acquired simultaneously are treated as a single image data pair for subsequent fusion processing.

[0032] 102. Input the first diffraction imaging data into the restoration network for image restoration processing to obtain the restored second diffraction imaging data.

[0033] Specifically, due to diffraction effects and limitations in actual manufacturing processes, the output images of the HMODL system suffer from a certain degree of blurring and loss of detail. This information loss or insufficient detail limits the reliability of the fusion results. Therefore, to address the image degradation caused by diffraction effects and processing errors during dual-band imaging, as well as the insufficient utilization of cross-band information, this invention inputs the infrared and visible light images output by the HMODL system into a restoration network before fusion, thereby recovering a high-quality, clear image and ensuring the reliability of the subsequent fused image quality.

[0034] 103. Input the second diffraction imaging data into the fusion network for image fusion processing to obtain a fused image. The fusion network is trained on a training dataset obtained by simulating the diffraction characteristics of the output image of a multi-order diffraction lens, performing degradation processing, and then restoring it through the restoration network.

[0035] The image fusion method disclosed in this invention fully considers the diffraction degradation characteristics of the output images of the HMODL system. By restoring the degraded infrared and visible light images separately before fusion, the sharpness of the images before fusion is restored, thus compensating for the blurring and information loss caused by diffraction effects. Furthermore, the fusion network is trained on a training dataset obtained by simulating the degraded characteristics of the output images from a hybrid multi-order diffraction lens and then restoring them through the restoration network. This constructs training samples that better conform to the characteristics of diffraction imaging. Therefore, while retaining the compact imaging advantages of HMODL, this invention can significantly improve the detail and information integrity of dual-band imaging, and enhance the robustness and application value of the system under actual diffraction imaging conditions.

[0036] Furthermore, in another embodiment, the restoration network is trained through the following steps:

[0037] 201. Obtain the first training dataset constructed from multiple pairs of infrared and visible light images.

[0038] Specifically, the infrared and visible light images in the first training dataset are from the LLVIP dataset. Specifically, 12,025 pairs of infrared / visible light images from the LLVIP dataset are used as the training dataset for the restoration network. The image size is set to 256×256 and normalized to the [0, 1] interval.

[0039] 202. Input the first training dataset into a preset image degradation model to obtain degraded image data, wherein the image degradation model is constructed based on the response characteristics of a hybrid multi-order diffraction lens to a spatial point source.

[0040] Specifically, constructing a pre-defined image degradation model requires first calculating the point spread function (PSF) of the HMODL system. The PSF describes the system's response to spatial point sources, and its calculation expression is as follows:

[0041] (1)

[0042] In formula (1), Represents the coordinates of the sampling ray. The coordinates of the sampling points on the image plane. Here are the coordinates of the complex amplitude sampling points on the HMODL transmission surface. The distance from the image plane point to the sampling point on the transmission surface. = Wavelength, The wavelength of light The design focal length for HMODL, It is a natural exponential function. The imaginary unit, To express differentiation, The complex amplitude optical field on the HMODL transmission surface is expressed as:

[0043] (2)

[0044] In formula (2), This is the total optical path of the light ray from the incident surface to the transmitting surface.

[0045] Then, a pre-defined image degradation model is constructed based on the calculated point spread function (PSF), specifically:

[0046] (3)

[0047] In formula (3), For the original clear image, For the image degraded by mixing multiple diffraction lenses, For a hybrid multi-order diffraction lens system, the point spread function is... This is an additive noise term.

[0048] 203. Input the obtained degraded image data into the restoration network, and train the restoration network with the target loss function as a condition until the model converges.

[0049] Specifically, the restoration network adopts an encoder-decoder structure based on the U-Net framework, such as... Figure 3 As shown in the first stage,

[0050] The restoration network first passes through a 3×3 convolutional layer to extract shallow features. It then enters the encoder section, which consists of multiple stages, each containing several NAF (Nonlinear Activation Free) modules. At the end of each stage, downsampling is achieved through a convolution with a stride of 2. Correspondingly, the decoder structure is symmetrical to the encoder. Each decoding stage first restores the spatial resolution through upsampling, then performs skip connections and concatenates the features with the corresponding encoder output features. Next, it passes through several NAF modules to fuse high- and low-level feature information, thereby enhancing the network's ability to recognize local information and preserve image details. Finally, the network uses a 3×3 convolutional layer to restore the feature maps to the output image.

[0051] The core component of the restoration network is the NAF module, which enables efficient feature modeling and nonlinear representation under lightweight conditions. In this module, the input features are first normalized. The normalized features are then processed through a 1×1 convolution to achieve channel blending, followed by a 3×3 depthwise separable convolution to extract spatial features. Next, the features are passed through a gating unit to introduce nonlinear mapping, and then through a channel attention mechanism. Finally, the output is processed through a 1×1 convolution to restore the channel dimension, and the input features are added to form a residual structure. The computation process can be represented as follows:

[0052] (4)

[0053] In formula (4), As input features, It is a convolutional layer. For depthwise separable convolution, To simplify the gating function, the channel features are split into two along one dimension and multiplied element-wise to achieve efficient nonlinear mapping. For channel attention mechanisms, LN is layer normalization. It obtains channel importance through global pooling and performs adaptive weighting to highlight salient features.

[0054] Building upon this, the module further introduces channel interaction paths to supplement and enhance feature representation, as shown in the following formula:

[0055] (5)

[0056] This branch achieves efficient channel feature interaction through two layers of 1×1 convolutions and gating operations, and... By forming progressively progressive residual connections, the feature representation capability and information flow efficiency can be improved while maintaining the network's lightweight nature.

[0057] The recovery network training process aims to stabilize the loss function. Specifically, the target loss function is defined as:

[0058] (6)

[0059] In formula (6): k is the current iteration number, and K is the total number of iterations. For optical loss function, Let K be the image loss function. The target loss function is obtained by weighted summation of the optical loss function and the image loss function. For example, K is set to 200; by the 200th iteration, the system has essentially converged to its optimal state, exhibiting the best imaging quality and stability. Furthermore, it should be noted that the conditions for stopping training in the actual fusion model training process can be determined according to the specific circumstances. In one case, training can be stopped when the loss function has stabilized before the maximum number of iterations is reached.

[0060] Furthermore, the image loss function in the target loss function Defined as:

[0061] (7)

[0062] In formula (7), To restore images from the network, This is the ground truth image. The image loss function specifically includes pixel loss, gradient loss, and structure loss:

[0063] (8)

[0064] in, All of these are hyperparameters. For pixel loss, For gradient loss, For structural loss. Pixel loss. The calculation method is as follows:

[0065] (9)

[0066] in, , These are infrared images and visible light images, respectively. To output the image, Representation matrix Norm, and These represent the height and width of the image, respectively. It is a weighting parameter that controls the balance between the two projects.

[0067] Gradient loss The calculation method is as follows:

[0068] (10)

[0069] in, Representing a matrix Norm, This represents gradient operation. This indicates the maximum selection.

[0070] Structural loss The calculation method is as follows:

[0071] (11)

[0072] in, This represents the structural similarity loss.

[0073] Furthermore, the optical loss function of the target loss function. Specifically defined as:

[0074] (12)

[0075] In formula (12), For wavelength, For the Strelby, optimize the Strelby for the wavelength by constraint. This improves the diffraction efficiency of HMODL, helping the restoration network to converge correctly in the initial stage.

[0076] Furthermore, traditional image fusion methods typically assume high-quality input images and do not fully consider the inherent degradation characteristics of diffraction images. Therefore, using traditional fusion methods not only increases the learning burden on the network but also makes it difficult to obtain stable and high-quality fusion results when the input image is severely degraded. Therefore, to address the above problems, in another embodiment, this invention discloses a fusion network, specifically trained through the following steps:

[0077] 301. Obtain a second training dataset constructed from multiple pairs of infrared and visible light images; wherein the first training dataset and the second training dataset are different;

[0078] 302. Input the second training dataset into the preset image degradation model to obtain degraded image data;

[0079] 303. Input the degraded image data into the trained restoration network to obtain restored image data;

[0080] 304. Use the restored image data as a new training dataset to train the fusion network until the model converges, thus obtaining the trained fusion network.

[0081] Specifically, to ensure that the fusion stage accurately reflects the imaging characteristics of the HMODL system, this method uses a different dataset for training than the restoration stage. The reason for not directly using the dataset used for training the restoration network is that the restoration network has already been trained on that dataset. If the same data were used again for training the fusion network, the results generated by the restoration network would be highly correlated with its training distribution, easily leading to information leakage and weakening the model's independence and generalization ability. Therefore, this invention introduces a dataset from a different source in the fusion stage, performs PSF-simulated degradation on this dataset, and then inputs the degraded images into the restoration network for restoration, thereby constructing training samples that better match the diffraction imaging characteristics.

[0082] Specifically, the second training dataset of this invention is taken from the RoadScene dataset, and preferably 221 pairs of infrared / visible light images from the RoadScene dataset are selected as the initial training set for the fusion network. It should be noted that, to ensure the quantity and quality requirements of the training data, this invention requires preprocessing the original training set before network training. The preprocessing operation involves cropping the training samples into 120 × 120 pixel image patches with a stride of 20, ultimately extracting 43,504 pairs of image patches and normalizing them to the [0, 1] interval.

[0083] Furthermore, the restored image data is used as a new training dataset to train the fusion network under the guidance of the loss function. When the loss function stabilizes or reaches the required number of iterations, the model converges, resulting in the trained fusion network.

[0084] The loss function includes pixel loss, gradient loss, and structural loss. Specifically, it is defined as follows:

[0085] (13)

[0086] in, All of these are hyperparameters. For pixel loss, For gradient loss, This is structural loss.

[0087] Pixel loss The calculation method is as follows:

[0088] (14)

[0089] in, Denotes the Frobenius norm of a matrix. and These represent the height and width of the source image, respectively. It is a weighting parameter that controls the balance between the two projects.

[0090] Gradient loss The calculation method is as follows:

[0091] (15)

[0092] in, Representing a matrix Norm, This represents gradient operation. This indicates the maximum selection.

[0093] Structural loss The calculation method is as follows:

[0094] (16)

[0095] in, This represents the structural similarity loss.

[0096] The fusion network of this invention first processes the initial training dataset by inputting it into an image degradation model for image degradation, and then inputting it into a restoration network for restoration, thereby obtaining high-quality image data with diffraction imaging properties to construct a training set for a diffraction imaging fusion network. By constructing a fusion training dataset that combines diffraction characteristics with high fidelity, the fusion network is provided with training samples that more closely resemble actual diffraction imaging conditions, thereby improving the model's generalization ability in real-world scenarios.

[0097] Furthermore, in another embodiment, the step of inputting the second diffraction imaging data into a fusion network for image fusion processing to obtain a fused image specifically includes:

[0098] 401. Input the second diffraction imaging data into the shallow sensing module to obtain the shallow features of the image;

[0099] 402. Input the shallow features of the image into the local guidance module, and obtain the enhanced features of the second diffraction imaging data by fusing the residual features and the enhancement features. The residual features are the features extracted by inputting the shallow features of the image into the residual block, and the enhancement features are the output features enhanced by combining the shallow features of the image with spatial attention and channel attention mechanisms.

[0100] 403. The enhanced features are input into the global perception module to extract global context-dependent features, and the output results are input into the reconstruction module for image fusion to obtain the fused image.

[0101] Specifically, such as Figure 3 As shown in the second stage, the fusion network includes: a shallow feature block, a local guidance block, a global perception block, a TFM block, and a reconstruction block.

[0102] In the shallow perception module, the infrared and visible light images are first concatenated along the channel dimension to form a two-channel input tensor. This input tensor is then fed into the first convolutional module to extract shallow features. This module uses a 5×5 convolutional kernel and incorporates batch normalization after the convolution operation. The activation function yields shallow features, specifically:

[0103] (17)

[0104] in, This is the output of the shallow sensing module. These are infrared images and visible light images, respectively. Indicates the number of input channels Number of output channels Convolution operation with a kernel size of 5×5. and These correspond to batch normalization and nonlinear activation, respectively.

[0105] After outputting from the shallow perception module, the local guidance module further enhances local details and global saliency features. Specifically, for example... Figure 3 The network architecture shown is the HARM Block, which combines residual learning with a dual attention mechanism. Input features first pass through the residual block, which consists of two convolutional layers, supplemented by batch normalization and... Activation is performed to extract texture information from the visible light image and thermal features from the infrared image, i.e., residual features. The output of the residual block can be represented as:

[0106] (18)

[0107] in, It is the input of the local boot module. It is the output of the residual block. This indicates a 3×3 convolution operation with 16 input and 16 output channels.

[0108] Subsequently, two mechanisms, channel attention and spatial attention, are introduced in parallel on the residual features to significantly enhance their representation. In the channel attention module, the input feature map first undergoes global average pooling to generate a global description vector for each channel. Then, a dimensionality-reducing fully connected layer compresses the number of channels from 16 to 4 and performs ReLU activation. A dimensionality-increasing fully connected layer restores the original number of channels, and finally, a Sigmoid activation generates channel attention weights. These weights are multiplied channel-by-channel with the input feature map, thereby highlighting key channels and suppressing non-key channels while maintaining the spatial size of the input feature map. This allows for flexible embedding of residual structures to improve feature representation capabilities. The formula can be expressed as:

[0109] (19)

[0110] in, It is the output of the channel attention module. Represents global average pooling. This means reducing the number of channels from 16 to 4. This indicates that the number of channels will be increased from 4 to 16. (Sigmoid) represents the activation function.

[0111] The spatial attention module performs a 3×3 convolution on the residual block output, compressing the number of channels from 16 to 4, and then... The activation introduces non-linearity, then another 3×3 convolution layer compresses the number of channels from 4 to 1, and finally... The activation function generates a spatial attention map, which can be expressed by the following formula:

[0112] (20)

[0113] in, It is the output of the spatial attention module. and These represent 3×3 convolution operations, with the input and output channels decreasing from 16 to 4 and then to 1.

[0114] After feature enhancement through channel attention and spatial attention, to fully utilize the residual features and the complementary information extracted by the two types of attention, they are accumulated and fused through residual connections. Specifically, the output of the channel attention branch can be expressed as:

[0115] (twenty one)

[0116] in, This represents the output of the channel attention module. This represents the characteristic output of the residual block. This is the final output.

[0117] Similarly, the output of the spatial attention branch is:

[0118] (twenty two)

[0119] in, This represents the output of the spatial attention module. The feature output of the residual block is represented by the feature weighting in the spatial dimension to highlight the salient target region in the image.

[0120] Subsequently, these two outputs are compared with the input residual features. Adding them together creates a fusion feature. :

[0121] (twenty three)

[0122] Finally, through The activation function yields the final output of the module. The formula is:

[0123] (twenty four)

[0124] in, The final output of the local guidance module is the enhanced feature. This connection method can retain the input feature information to the greatest extent, while fusing the features enhanced by channel and spatial attention, to achieve comprehensive extraction of important textures and salient targets in infrared and visible light images.

[0125] Furthermore, such as Figure 3 As shown in the network architecture of the Global Awareness Module (TFM Block), the input features are first normalized layer by layer to stabilize training and standardize the features. Then, they are fed into the multi-head self-attention module to calculate the global dependency of each location on other locations. Finally, the input features are added to the self-attention output, as shown in the following formula:

[0126] (25)

[0127] in, This is the input feature map for the global perception module. It is layer normalization. Represents the bulls' self-attention. This is the output.

[0128] Subsequently, the module... Perform another round of layer normalization, then input the data into the multilayer perceptron for nonlinear feature transformation, as shown in the following formula:

[0129] (26)

[0130] in, This is the final output of the global perception module. Contains two layers and Activation. Through this design, the global awareness module can fully extract global contextual dependencies while preserving input feature information.

[0131] It should be noted that, in one embodiment, after completing local guidance and global perception modeling, the output results are directly reconstructed to obtain the fused image. In another embodiment, such as... Figure 3 As shown, after completing the local guidance module, the fusion network sequentially stacks two levels of global perception modules to gradually reconstruct richer texture details from low to high layers. Subsequently, the fusion network introduces a local guidance module again on high-level semantic features to further enhance key local features.

[0132] Finally, a convolutional module was designed to reduce dimensionality and generate the final fused image. The convolution operation is 1×1, and the formula is expressed as:

[0133] (27)

[0134] in, This is the final output fused image. Indicates the input channel is Output channels are 1×1 convolution, The input feature map for this convolution is... This represents the hyperbolic tangent activation function, used to map the fusion result to... scope.

[0135] To further illustrate the beneficial effects achieved by this invention, an engineering experiment was conducted on the HMODL system. This experiment used a system with a broadband ZnS substrate and a 40mm aperture. The visible light sensor had a pixel size of 3.45μm and a resolution of 2048×2448; the infrared sensor had a pixel size of 15μm and a resolution of 640×512. Due to the difference in resolution between visible light and infrared images, the Final2x method was used for super-resolution processing of the infrared images in the experiment. The infrared images and their restoration results are shown below. Figure 4 As shown, the visible light image and its restoration result are as follows: Figure 5 As shown, the fused image is as follows Figure 6 As shown.

[0136] In addition, to quantitatively evaluate the fusion effect, the experiment introduced structural content difference (SCD), visual information fidelity (VIF), information entropy (EN), and gradient-based metrics. The five objective indicators, including structural similarity (SSIM), are as follows: Figure 7 As shown in the experimental results, most evaluation index values ​​were improved. These experimental results demonstrate that the method proposed in this invention can effectively improve the dual-band image fusion performance of a hybrid multi-order diffraction lens imaging system.

[0137] This application also provides an electronic device, which includes a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded by the processor and executed by the processor to perform the image fusion method based on hybrid multi-order diffraction lenses provided in the above method embodiments.

[0138] Furthermore, an electronic device is provided for implementing the method provided in the embodiments of this application. This device can participate in constituting or including the apparatus or system provided in the embodiments of this application. The electronic device 21 may include one or more processors (processors may include, but are not limited to, processing devices such as microprocessors (MCUs) or programmable logic devices (FPGAs), a memory for storing data, and a transmission device for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera.

[0139] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits can be implemented wholly or partially as software, hardware, firmware, or any other combination. Furthermore, the data processing circuits can be a single, independent processing module, or wholly or partially integrated into any other element within a device (or mobile device). As involved in the embodiments of this application, the data processing circuit serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0140] The memory can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method described in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the above-mentioned data processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to electronic devices via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0141] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the device's communication provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0142] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of an electronic device (or mobile device).

[0143] This application also provides a computer storage medium storing at least one instruction or at least one program, which is loaded and executed by a processor to implement the image fusion method based on hybrid multi-order diffraction lenses provided in the above-described method embodiments.

[0144] Optionally, in this embodiment, the aforementioned computer storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the aforementioned storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0145] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer storage medium. The processor of an electronic device reads the computer instructions from the computer storage medium and executes the computer instructions, causing the electronic device to perform the image fusion method based on hybrid multi-order diffraction lenses provided in the above-described method embodiments.

[0146] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0147] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.

Claims

1. An image fusion method based on hybrid multi-order diffraction lenses, characterized in that, The method includes: Acquire first diffraction imaging data output by a hybrid multi-order diffraction lens, the first diffraction imaging data including infrared image and visible light image data pairs; The first diffraction imaging data is input into the restoration network for image restoration processing to obtain the restored second diffraction imaging data. The second diffraction imaging data is input into the fusion network for image fusion processing to obtain a fused image. The fusion network is trained on a training dataset obtained by simulating the diffraction characteristics of the output image of a multi-order diffraction lens, performing degradation processing, and then restoring it through the restoration network.

2. The method according to claim 1, characterized in that, The restoration network is trained through the following steps: Obtain the first training dataset constructed from multiple pairs of infrared and visible light image data; The first training dataset is input into a preset image degradation model to obtain degraded image data, wherein the image degradation model is constructed based on the response characteristics of a hybrid multi-order diffraction lens to a spatial point source; The obtained degraded image data is input into the restoration network, and the restoration network is trained with the target loss function as a condition until the restoration network converges.

3. The method according to claim 2, characterized in that, The preset image degradation model is specifically as follows: (3) in, For the original clear image, For the image degraded by mixing multiple diffraction lenses, For a hybrid multi-order diffraction lens system, the point spread function is... This is an additive noise term.

4. The method according to claim 3, characterized in that, The point spread function of the hybrid multi-order diffraction lens system is specifically: (1) in, Represents the coordinates of the sampling ray. The coordinates of the sampling points on the image plane. Here are the coordinates of the complex amplitude sampling points on the HMODL transmission surface. The distance from the image plane point to the sampling point on the transmission surface. = Wavelength, The wavelength of light The design focal length for HMODL, This represents the complex amplitude optical field on the HMODL transmission surface. It is a natural exponential function. The imaginary unit, This indicates differentiation.

5. The method according to claim 2, characterized in that, The target loss function specifically includes: (6) Where k is the current iteration number, and K is the total number of iterations. For optical loss function, This is the image loss function.

6. The method according to claim 2, characterized in that, The fusion network is trained through the following steps: A second training dataset is constructed from multiple pairs of infrared and visible light image data; wherein the first training dataset and the second training dataset are different. Input the second training dataset into the preset image degradation model to obtain degraded image data; The degraded image data is input into the restoration network to obtain restored image data; The restored image data is used as a new training dataset to train the fusion network until the fusion network converges, resulting in a fully trained fusion network.

7. The method according to claim 1, characterized in that, The step of inputting the second diffraction imaging data into a fusion network for image fusion processing to obtain a fused image includes: The second diffraction imaging data is input into the shallow sensing module to obtain the shallow features of the image; The shallow features of the image are input into the local guidance module. By fusing residual features and enhancement features, the enhanced features of the second diffraction imaging data are obtained. The residual features are the features extracted by inputting the shallow features of the image into the residual block, and the enhancement features are the output features enhanced by combining the shallow features of the image with spatial attention and channel attention mechanisms. The enhanced features of the second diffraction imaging data are input into the global perception module to extract global context-dependent features, and the output results are input into the reconstruction module for image fusion to obtain a fused image.

8. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the method as claimed in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium, which a processor reads from and executes to implement the method as described in any one of claims 1 to 7.