Multimodal image fusion method and imaging system
By employing a multimodal image fusion method, utilizing filtering and frequency domain transformation of feature spectral images, pseudo-color images, and polarization images, combined with loss function optimization, the problems of high difficulty in multimodal image registration and poor fusion effect are solved, achieving high-precision multimodal information fusion and target detection.
Patent Information
- Application Number
- CN202511272512.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-09-08
AI Technical Summary
In existing technologies, multimodal image fusion technology relies on time-sharing imaging systems or multiple devices to capture target scene images, which leads to difficulties in registering images of different modalities, low registration accuracy, inconsistent image edge details, and poor fusion results.
A multimodal image fusion method is adopted, which obtains the original image and spectral information of the target scene, performs filtering and frequency domain transformation on the feature spectral image, pseudo-color image and polarization degree image, and combines loss function optimization to achieve accurate fusion of multimodal information.
It improves the accuracy and detail preservation of multimodal image fusion, enhances the accuracy and comprehensiveness of target detection tasks, reduces the difficulty of image registration, and improves imaging quality.
Smart Images

Figure CN120808093B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to image processing methods for optical remote sensing imaging, specifically to multimodal image fusion methods and imaging systems. Background Technology
[0002] In the field of optical remote sensing imaging, multimodal information images play a crucial role in computer vision tasks such as target classification, detection, localization, and segmentation. Depending on the target's attributes, different modalities play different roles in algorithms for different tasks. Panchromatic images provide information such as target location and intensity for various tasks, spectral images provide spectral feature information such as target chemical composition, and polarization images provide information such as target surface features and material texture details. Due to the significant complementary properties of multimodal information, its fusion technology has gradually become a key means to overcome the limitations of traditional optical remote sensing imaging.
[0003] Current multimodal image fusion technology relies on time-sharing imaging systems or multiple devices to capture target scene images. This results in the acquisition of multimodal information images of the target scene not being simultaneous or from the same source. In subsequent image processing, the registration of different modal images is difficult, the registration accuracy is low, and the image edge details are inconsistent, leading to poor fusion effect and limited application. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies where target scene image capture relies on time-sharing imaging systems or multi-device acquisition, resulting in difficulties in registering images of different modalities, low registration accuracy, inconsistent image edge details, and poor fusion effects. This invention provides a multimodal image fusion method and imaging system.
[0005] To achieve the above objectives, the technical solution provided by this invention is as follows:
[0006] A multimodal image fusion method, characterized by the following steps:
[0007] S1, Raw Image and Feature Spectrum Acquisition:
[0008] Acquire the original image of the target scene, including the spectral image and polarization image of the target scene; simultaneously acquire the spectral information of the target scene and plot the spectral radiation curve;
[0009] S2, Multimodal Image Preprocessing:
[0010] Characteristic spectral images are obtained based on spectral radiometric curve analysis; pseudo-color images are obtained based on spectral images; polarization degree images are obtained based on polarization images.
[0011] S3, Multimodal Image Fusion:
[0012] S3.1 Filters the characteristic spectral image, pseudo-color image, and polarization degree image to obtain the corresponding base image;
[0013] S3.2 After performing frequency domain transformation on the characteristic spectral image, pseudo-color image, polarization degree image and their corresponding base map, subtract the corresponding images, and then transform them back to the spatial domain to obtain the corresponding detail map;
[0014] S3.3 After the base maps are superimposed in the spatial domain, a fused base map is obtained. After the detail maps are superimposed in the spatial domain, a fused detail map is obtained. The fused base map and the fused detail map are then subjected to frequency domain transformation and added together. Finally, they are inversely transformed to the spatial domain to obtain a multimodal information fused image.
[0015] Furthermore, in step 1, the original image also includes a panchromatic image of the target scene;
[0016] The multimodal image fusion method further includes:
[0017] Step S4, Image quality evaluation and optimization:
[0018] S4.1 uses panchromatic images as a benchmark, constructs a loss function using image quality evaluation metrics, and evaluates the quality of multimodal information fusion images through the loss function;
[0019] S4.2 Determine whether the loss function has reached its minimum value. If yes, proceed to S4.3; otherwise, adjust the corresponding parameters of filtering and frequency domain transformation in steps S3.1-S3.3 and return to step S3.1.
[0020] S4.3 outputs the current multimodal information fused image, completing the multimodal image fusion.
[0021] Furthermore, in step S3.1, the feature spectral image, pseudo-color image, and polarization degree image are filtered by initialization convolution kernels of different scales.
[0022] In step S4.1, the image quality evaluation metrics include image information entropy, edge preservation index, and structural similarity index;
[0023] Step S4.3 also includes: saving the corresponding parameters of filtering and frequency domain transformation in steps S3.1-S3.3 for the current multimodal information fusion image, and obtaining the optimal model parameters for multimodal image fusion of the target scene.
[0024] Further, in step S4.1, the loss function is specifically as follows:
[0025]
[0026] Where P represents the reference image, F represents the fused image, H(P) represents the information entropy of the reference image, H(F) represents the information entropy of the fused image, and Q... abf (P,F) represents the edge preservation index of the fused image F with reference to the reference image P, and SSIM(P,F) represents the structural similarity index of the fused image F with reference to the reference image P. This is the loss function.
[0027] Furthermore, in step S4.1, the image quality evaluation metrics also include image cross-entropy, image conditional entropy, peak signal-to-noise ratio, and / or mutual information.
[0028] Furthermore, step S2 specifically includes:
[0029] S2.1 Based on the spectral radiation curve, perform spectral feature analysis to find the spectral band with the largest difference in reflected irradiance of different target objects, use this as the characteristic spectrum, and use the spectral images of the corresponding spectral bands to subtract to obtain the characteristic spectral image;
[0030] S2.2 Based on spectral images, pseudo-color images are obtained by using PCA principal component analysis, weighted mapping of four-channel spectral information, or mapping of three-channel spectral information;
[0031] S2.3 uses Stokes vectors to solve for the polarization information of the target scene based on the polarization image and outputs a polarization degree image.
[0032] Meanwhile, this invention provides a multimodal image imaging system for acquiring original images and spectral information of a target scene to realize the aforementioned multimodal image fusion method. Its special feature is that it includes a sub-aperture array compound eye unit and a spectrometer. The sub-aperture array compound eye unit includes a lens array and a multimodal focal plane camera. The lens array includes a central sub-aperture lens and multiple edge sub-aperture lenses arranged along its outer periphery. The fields of view of the central sub-aperture lens and the multiple edge sub-aperture lenses overlap, and the fields of view of the multiple edge sub-aperture lenses completely cover the field of view of the central sub-aperture lens. The central sub-aperture lens and each edge sub-aperture lens are spaced apart to prevent image aliasing. The optical channels corresponding to the central sub-aperture lens and each edge sub-aperture lens are set as panchromatic channels, spectral channels, or polarization channels to form corresponding sub-aperture images, which are panchromatic images, spectral images, and / or polarization images. The multimodal focal plane camera includes an image sensor and a data transmission unit. The data transmission unit is used to transmit the images captured by the image sensor to an external industrial control computer. The spectrometer is used to collect spectral information of the target scene.
[0033] Furthermore, the spectral channel is provided with a filter corresponding to the imaging area of the focal plane of the image sensor, and the polarization channel is provided with a polarizer corresponding to the imaging area of the focal plane of the image sensor.
[0034] Neutral density filters are provided on the surface of the polarizer and on the imaging area of the panchromatic channel to attenuate the incident light energy, so as to ensure that the light energy of the panchromatic channel, spectral channel and polarization channel imaging is consistent or matched.
[0035] Furthermore, the optical channel corresponding to the central sub-aperture lens is set as a panchromatic channel; the filter set in the spectral channel is a narrowband filter, which is set on the focal plane of the image sensor by physical vapor deposition; the polarizer is a linear polarizer or a circular polarizer, which is transferred to the focal plane of the image sensor by nanoimprint lithography; the neutral density filter is an absorption bandpass filter, and its working spectrum is consistent with the working spectrum of the image sensor.
[0036] Furthermore, the original image is a compound eye image. The panchromatic image, spectral image, and polarization image of the target scene are obtained through image segmentation. The image segmentation specifically involves: imaging the positioning calibration plate to obtain a calibration image; obtaining the center coordinates and radius parameters of the positioning calibration plate in the sub-aperture images formed by the sub-aperture lenses at each edge of the calibration image; the radius parameter being the minimum distance between the center coordinates of the positioning calibration plate and the edge of the sub-aperture image; and then segmenting the original image according to the center coordinates and radius parameters of the positioning calibration plate in each sub-aperture image to obtain the panchromatic image, spectral image, and / or polarization image.
[0037] The beneficial effects of this invention are:
[0038] 1. The multimodal image fusion method of the present invention acquires the spectral radiation curve of the target scene, analyzes the spectral features, reconstructs the characteristic spectral image of the target scene, and fuses it with the pseudo-color image. Compared with existing image fusion methods, it can more accurately calculate and preserve image edge details, and improve the multimodal information image fusion effect.
[0039] 2. In the multimodal image fusion method of the present invention, feature spectral images, pseudo-color images, and polarization degree images are used as source images, and panchromatic images are used as reference images. Multi-band spectral information, polarization information, and spatial information of the target scene are fused to preserve detailed texture information to the maximum extent. The multimodal source images make up for the defects of existing polarization detection technology, which has a small ambient light polarization component and poor detection effect, and effectively improve the accuracy and comprehensiveness of target detection tasks.
[0040] 3. The multimodal image fusion method of the present invention also uses a variety of image evaluation indicators to evaluate the quality of the multimodal information fused image, and compares it with the source image and reference image to comprehensively evaluate the fusion effect. By dynamically adjusting the initial convolution kernel scale through the loss function, the optimal fusion effect is achieved with a small number of parameters.
[0041] 4. In the sub-aperture array compound eye unit of the present invention, the central sub-aperture lens and each edge sub-aperture lens correspond to different information images, and their design parameters are completely consistent. When imaging the target scene, it can achieve simultaneous acquisition of different modal information from the same source. The hardware level reduces the difficulty of image registration and preserves the image edge detail information to the greatest extent.
[0042] 5. The sub-aperture array compound eye unit of the present invention also sets a neutral density filter in the imaging area of the panchromatic channel and on the polarizer integrated on the focal plane of the image sensor to attenuate the incident light energy, thereby adjusting the inconsistency of light energy in the spectrum and the panchromatic and polarized channels in the same exposure time, effectively improving the imaging quality of simultaneous acquisition of different modal information from the same source.
[0043] 6. The sub-aperture array compound eye unit of this invention employs physical vapor deposition and nanoimprint lithography processes to integrate filters and polarizers with different parameters onto a large-area sensor. Depending on the spectral response of the specific target scene, filter structures with different center wavelengths and different full width at half maximum (FWHM) can be selected. Furthermore, the spectral bands or polarization angles on the large-area sensor can be modified according to actual application needs, greatly enhancing the dynamic adaptability of the sub-aperture array compound eye unit. Attached Figure Description
[0044] Figure 1 This is a schematic diagram of the sub-aperture array compound eye unit in an embodiment of the multimodal image imaging system of the present invention;
[0045] Figure 2 This is a schematic diagram of the field of view of the lens array in the compound eye unit of the sub-aperture array in an embodiment of the multimodal image imaging system of the present invention;
[0046] Figure 3 This is a schematic diagram of the focal plane of the multimodal focal plane camera in an embodiment of the multimodal image imaging system of the present invention;
[0047] Figure 4 This is a schematic diagram of the positioning and calibration plate imaging in step 1 of the embodiment of the multimodal image fusion method of the present invention;
[0048] Figure 5 This is a schematic diagram of obtaining the radius parameter in step 1 of the embodiment of the multimodal image fusion method of the present invention;
[0049] Figure 6 This is a flowchart of step 2 in an embodiment of the multimodal image fusion method of the present invention;
[0050] Figure 7 This is a flowchart of step 3 in an embodiment of the multimodal image fusion method of the present invention;
[0051] Figure 8 This is a flowchart of step 4 in an embodiment of the multimodal image fusion method of the present invention.
[0052] Explanation of reference numerals in the attached figures:
[0053] 1-Lens array, 101-Center sub-aperture lens, 102-Edge sub-aperture lens, 2-Multimodal focal plane camera. Detailed Implementation
[0054] This invention provides a multimodal image imaging system, including a sub-aperture array compound eye unit and a spectrometer. The spectrometer is used to acquire spectral information of the target scene. The structure of the sub-aperture array compound eye unit is as follows: Figure 1 As shown, it mainly includes a lens array 1 and a multimodal focal plane camera 2.
[0055] Lens array 1 includes a central sub-aperture lens 101 and a plurality of edge sub-aperture lenses 102 arranged along its outer periphery; the design parameters (optical parameters and structural parameters) of the central sub-aperture lens 101 and each edge sub-aperture lens 102 are completely identical. In this embodiment, the plurality of edge sub-aperture lenses 102 are arranged in a rectangle, located at the top corner and four sides of the rectangle, respectively. Figure 2 As shown, each edge sub-aperture lens 102 has a large field of view overlap with the central sub-aperture lens 101. With the central sub-aperture lens 101 as a reference, the overlapping fields of view of the edge sub-aperture lenses 102 completely cover the field of view of the central sub-aperture lens 101. A certain distance exists between the central sub-aperture lens 101 and each edge sub-aperture lens 102, ensuring that the images formed on the image sensor have spacing and do not overlap. In other embodiments of the present invention, the multiple edge sub-aperture lenses 102 can also be arranged in the form of a regular hexagon, regular octagon, circle, etc., respectively located on the sides and / or vertices of the regular hexagon, regular octagon, etc.
[0056] The multimodal focal plane camera 2 includes a data transmission unit and an image sensor. The data transmission unit transmits images captured by the image sensor to an industrial control computer. The data transmission unit provides at least two data transmission methods; in this embodiment, a USB 3.0 interface or a GigE interface is selected for data transmission. In other embodiments of the invention, a CameraLink interface, a CoaXPress interface, or other data transmission interfaces can also be selected. In other embodiments of the invention, the multimodal focal plane camera 2 also includes a power supply system. The power supply system uses a common power supply interface, which can be a DC power I / O interface, or a USB 3.0 or CoaXPress interface, to supply power to the image sensor and the data transmission unit.
[0057] like Figure 3As shown, in the multimodal focal plane camera 2, the optical channels corresponding to the central sub-aperture lens 101 and each edge sub-aperture lens 102 are set as panchromatic channels, spectral channels, or polarization channels. A filter is set on the focal plane of the image sensor corresponding to the spectral channel, and a polarizer is set on the focal plane of the image sensor corresponding to the polarization channel. The filter or polarizer can completely cover the imaging area of the corresponding single central sub-aperture lens 101 or edge sub-aperture lens 102, and is used to acquire spectral images and polarization images, respectively.
[0058] The filter is integrated using a physical vapor deposition process. In this embodiment, the filter is specifically fabricated by plasma sputtering, where multiple layers are stacked on the focal plane of the image sensor. The filter is configured as a narrowband filter, and the center wavelength and corresponding full width at half maximum (FWHM) can be selected according to the actual target scene.
[0059] The polarizer is fabricated using nanoimprint lithography, where a nanostructured grating is mechanically imprinted onto the focal plane. After mechanically imprinting the nanostructured grating, a neutral density filter is integrated onto its surface using physical vapor deposition to attenuate the incident light energy, ensuring that the spectral and polarization information have consistent or matched light energy during the same exposure time in the imaging process.
[0060] like Figure 3 As shown, in this embodiment, filters are respectively set in the imaging areas located at the upper left, upper right, lower left, and lower right on the focal plane. The center wavelengths are selected as red light (R: 650 nm), green light (G: 532 nm), blue light (B: 472 nm), and near-infrared light (NIR: 730 nm), respectively. These four spectral bands are used as the center wavelengths of the spectral information, and the full width at half maximum (FWHM) is uniformly 20 nm. Simultaneously, four linear polarizers are set to acquire polarization information, located in the upper, right, lower, and left imaging areas of the focal plane, respectively, with polarization angles of 0°, 45°, 90°, and 135°. The optical channels corresponding to the edge sub-aperture lens 102 are polarization channels. In other embodiments of the present invention, other center wavelengths and corresponding FWHMs can be selected according to the focal plane response and task requirements. Linear polarizers with other polarization angles (30°, 60°) can also be selected, or circular polarizers can be used to acquire polarization information.
[0061] The optical channel corresponding to the central sub-aperture lens 101 is set as a panchromatic channel (PAN). A neutral density filter is integrated into the imaging area of the central sub-aperture lens 101 corresponding to the focal plane to attenuate the incident light energy. However, its transmittance should be less than that of the neutral density filter on the polarizer surface to ensure that the light energy of the panchromatic channel, spectral channel, and polarization channel imaging is consistent. All of the above-mentioned neutral density filters are set as absorption-type bandpass filters, with their working spectrum consistent with that of the image sensor. The transmittance can be adjusted according to the number of deposited film layers.
[0062] The multimodal image fusion method of this invention fuses original images to output a single multimodal information fused image. The original image is a multimodal image acquired by the compound eye unit of the sub-aperture array, including multiple sub-aperture images corresponding to the central sub-aperture lens 101 and the edge sub-aperture lenses 102, respectively. Each sub-aperture image contains different modal information. The multimodal image fusion method specifically includes five stages: multimodal image registration, feature spectrum acquisition, multimodal image preprocessing, multimodal image fusion, and fused image quality evaluation, as shown below:
[0063] 1. Multimodal image registration;
[0064] 1.1 such as Figure 4 As shown, a positioning calibration plate is drawn, with a circular marker at the center and right-angled lines equidistant from the center of the marker. A sub-aperture array compound eye unit is used to image the plate. The distance between the sub-aperture array compound eye unit and the positioning calibration plate is adjusted so that the central sub-aperture image precisely encompasses the entire positioning calibration plate, and all four right-angled lines are located at the edge of the central sub-aperture image.
[0065] The positioning calibration plate is used to confirm the overlapping area of the field of view of the edge sub-aperture lens 102 and the center sub-aperture lens 101, and can be set to any form or pattern.
[0066] 1.2 such as Figure 5 As shown, calculate and record the center coordinates of the positioning calibration plate in each sub-aperture image. Take the center of the positioning calibration plate in the spectral image (upper left, upper right, lower left, lower right) as the center, and the minimum distance of the plate from the edge of the sub-aperture image as the radius, and record the radius parameter.
[0067] 1.3 Save the center coordinates and radius parameters of the positioning calibration plate in each of the above sub-aperture images. Similarly, obtain the center coordinates and radius parameters of the positioning calibration plate in the sub-aperture image corresponding to the polarization image.
[0068] 2. Acquisition of original images and characteristic spectra; such as... Figure 6 As shown, it includes the following steps:
[0069] 2.1 The original image of the target scene is acquired using a sub-aperture array compound eye unit. Simultaneously, during the imaging process, a spectrometer is used to collect spectral information of the target scene, and the acquired spectral information is uploaded to the industrial control computer at the processing end.
[0070] In this embodiment, the target scene is an outdoor scene, and the target objects include plants, artificial plants, and metal products of the same color. Other embodiments of the present invention can collect spectral information for different target scenes and different target objects.
[0071] 2.2 The industrial control computer plots the spectral radiation curve based on the collected spectral information, where the horizontal axis is the wavelength, the wavelength range is the working spectral band of the multimodal focal plane camera 2, and the vertical axis is the reflected irradiance of the target object in the scene collected by the spectrometer.
[0072] 3. Multimodal image preprocessing; such as Figure 7 As shown, the specific steps include:
[0073] 3.1 Using the center coordinates and radius parameters of the positioning calibration plate in each sub-aperture image from step 1, the original image is segmented to obtain the overlapping field of view in the sub-aperture images. In this embodiment, the original image is segmented to obtain the panchromatic image (PAN) corresponding to the central sub-aperture lens 101 and the four-band spectral images (R, G, B, NIR) and four-angle polarization images (0°, 45°, 90°, 135°) corresponding to the edge sub-aperture lenses 102.
[0074] 3.2 Based on the above spectral radiation curves, spectral feature analysis is performed to identify the spectral bands with the greatest difference in reflected irradiance among different target objects, which are then used as characteristic spectra. The characteristic spectral images are obtained by subtracting the spectral images of the corresponding spectral bands.
[0075] In this embodiment, the spectral radiation curves of plants, artificial plants, and metal products of the same color show the greatest differences in the near-infrared (NIR) and green (G) spectral bands. Therefore, the characteristic spectral image of the target scene is obtained by subtracting the near-infrared image and the green spectral image. In other embodiments of the present invention, spectral images of other spectral bands can be selected according to different target objects.
[0076] 3.3 Based on the four-band spectral images, principal component analysis (PCA) is used to calculate the covariance matrix of the four-band spectral images. Singular value decomposition is then performed to obtain the eigenvalues and eigenvectors corresponding to the four spectral bands. The three principal components containing the most eigenvalues are selected, and their corresponding eigenvectors are mutually orthogonal. The four-band spectral images are projected onto the eigenvectors corresponding to the selected mutually orthogonal principal components to obtain three principal component images. These three principal component images are then mapped to the RGB color space to obtain pseudo-color images.
[0077] In this embodiment, the principal component images corresponding to the first three principal components containing most of the feature values are defined as the first principal component image, the second principal component image, and the third principal component image, respectively. The first principal component image is mapped to the R channel, the second principal component image is mapped to the G channel, and the third principal component image is mapped to the B channel to obtain a pseudo-color image.
[0078] In other embodiments of the present invention, in addition to the PCA principal component analysis method, other methods such as existing four-channel spectral information weighted mapping and three-channel spectral information mapping can also be used to obtain pseudo-color images; when using the PCA principal component analysis method to obtain pseudo-color images, the above three principal component images can also be mapped to other channels of the RGB color space; in addition to the RGB color space, other color space systems such as HSV and HIS can also be selected for mapping.
[0079] 3.4 Based on the four-angle polarization image, the Stokes vector is used to solve the polarization information of the target scene and output the polarization degree image.
[0080] In this embodiment, the four-angle polarization images are linearly polarized images at four angles: 0°, 45°, 90°, and 135°. In other embodiments of the present invention, linearly polarized images at 30° and 60° or circularly polarized images can also be selected to solve for polarization information.
[0081] The aforementioned characteristic spectral image, pseudo-color image, and polarization degree image are used as the source images for the multimodal image fusion model in step 4.
[0082] 4. Multimodal image fusion, such as Figure 8 As shown, the source image is processed using a multimodal image fusion model, specifically including the following steps:
[0083] 4.1 Different scales of initialization convolution kernels are used to filter the feature spectral image, pseudo-color image, and polarization degree image respectively, making full use of the different scale feature information of different modal images to obtain corresponding base maps, which are denoted as the first base map, the second base map, and the third base map respectively. The base maps retain most of the low-frequency information in the source image, namely the overall structure and brightness distribution information, and can effectively extract the structural edges of the target scene;
[0084] In this embodiment, the initial convolution kernel type is a bilateral filter convolution kernel, and the initial convolution kernel scales corresponding to the feature spectral image, pseudo-color image, and polarization degree image are 3×3, 9×9, and 15×15, respectively. In other embodiments of the present invention, other types of convolution kernels can be selected, such as Gaussian filters, low-pass filters, and high-pass filters, and other sizes of convolution kernels can also be selected, such as 5×5, 7×7, and 11×11.
[0085] 4.2 The characteristic spectral image, pseudo-color image, polarization degree image, and corresponding base map are transformed in the frequency domain, i.e., transformed from the spatial domain to the frequency domain. In the frequency domain, the corresponding base map is subtracted from the characteristic spectral image, pseudo-color image, and polarization degree image, and then transformed back to the spatial domain to obtain the corresponding detail maps, which are denoted as the first detail map, the second detail map, and the third detail map, respectively. The detail maps retain the high-frequency information in the source image, i.e., the detail texture information.
[0086] In this embodiment, the Fast Fourier Transform (FFT) method is used when transforming the image from the spatial domain to the frequency domain or from the frequency domain to the spatial domain. In other embodiments of the present invention, other methods such as Discrete Fourier Transform, Inverse Discrete Fourier Transform, Wavelet Transform, and Laplace Transform may also be used.
[0087] 4.3 The first, second, and third base maps are superimposed in the spatial domain to obtain a fused base map, and the first, second, and third detail maps are superimposed in the spatial domain to obtain a fused detail map. The fused base map and the fused detail map are then subjected to frequency domain transformations and added together, followed by an inverse transformation to the spatial domain to obtain a multimodal information fused image.
[0088] 5. Image quality evaluation and optimization;
[0089] 5.1 Establish a loss function based on image quality evaluation metrics. The loss function is calculated using image quality evaluation metrics to evaluate the quality of multimodal information fusion images. The smaller the loss function value, the higher the image fusion quality.
[0090] In this embodiment, the loss function includes image information entropy (Entropy), edge preservation index (Qabf), and structural similarity index (SSIM), as detailed below:
[0091] Image information entropy, a no-reference image quality evaluation metric, is used to evaluate the richness of information content in multimodal information fusion images;
[0092] Edge Preservation Index (EPI) is a reference image quality evaluation metric used to evaluate the ability of a multimodal information fusion image to preserve important edge information in the source image.
[0093] The structural similarity index is a reference image quality evaluation metric used to evaluate the local similarity of multimodal information fusion images in terms of brightness, contrast, and structure.
[0094] The required reference image is a panchromatic image (PAN) obtained from the compound eye cells of the sub-aperture array.
[0095] loss function Specifically:
[0096]
[0097] Right now,
[0098]
[0099] Where P represents the reference image, F represents the fused image, H(P) represents the information entropy of the reference image, H(F) represents the information entropy of the fused image, and Q... abf(P,F) represents the edge preservation index of the fused image F with reference to the reference image P, and SSIM(P,F) represents the structural similarity index of the fused image F with reference to the reference image P.
[0100] 5.2 Determine if the loss function has reached its minimum value. If yes, proceed to step 5.3. If no, adjust the steps by modifying the scale of the initial convolution kernel in step 4.1 through parameter feedback, dynamically adjust the relevant parameters of the frequency domain transformation in steps 4.2 and 4.3, and return to step 4.1.
[0101] 5.3 Output the current multimodal information fusion image to complete the multimodal image fusion. At the same time, save the initial convolution kernel scale and the relevant parameters of the mid-frequency domain transformation in steps 4.2 and 4.3 corresponding to the current multimodal information fusion image, and use them as the optimal model parameters for the target scene.
[0102] In the quality evaluation of fused images, in addition to the image quality evaluation metrics used in the loss function mentioned above, more image quality evaluation metrics can be used to evaluate the quality of the fused image. Other embodiments of the present invention may also include:
[0103] Image cross-entropy: A reference image quality evaluation metric used to evaluate the difference between the fused image and the reference image;
[0104] Image conditional entropy: A reference image quality evaluation metric used to evaluate the complementarity of fused images to information from different modal source images under the condition of a known reference image;
[0105] Peak signal-to-noise ratio (PSNR): A reference image quality metric used to evaluate the distortion of the fused image relative to the reference image;
[0106] Mutual information: A reference image quality evaluation metric used to evaluate the amount of shared information between the fused image and the source image.
[0107] The aforementioned multimodal image fusion method images the target scene by combining aperture and focal plane. Different modal information images are captured simultaneously by the same image sensor, which greatly reduces the difficulty of image registration. By fusing spectral and polarization information, it can effectively distinguish different types of targets, achieve simultaneous acquisition of panchromatic, spectral, and polarization information from the same source, and effectively complete the fusion of multimodal information images. It provides an important technical means for optical remote sensing imaging detection and has great application potential in target detection and classification tasks.
Claims
1. A multi-modal image fusion method, characterized in that, The method comprises the following steps: S1, original image and characteristic spectrum acquisition: Obtain an original image of a target scene, the original image comprising a spectral image, a panchromatic image and a polarization image of the target scene; simultaneously, acquire spectral information of the target scene, and draw a spectral radiation curve; S2, multi-modal image preprocessing: Obtain a characteristic spectral image based on spectral radiation curve analysis; Obtain a pseudo-color image based on the spectral image, and obtain a degree of polarization image based on the polarization image; S3, multi-modal image fusion: S3.1, filter the characteristic spectral image, the pseudo-color image and the degree of polarization image to obtain corresponding base images respectively; S3.2, subtract the characteristic spectral image, the pseudo-color image and the degree of polarization image after frequency domain transformation, and then re-transform to the spatial domain to obtain corresponding detail images; S3.3, superimpose the base images in the spatial domain to obtain a fusion base image, superimpose the detail images in the spatial domain to obtain a fusion detail image, add the fusion base image and the fusion detail image after frequency domain transformation, and then inverse transform to the spatial domain to obtain a multi-modal information fusion image; S4, fusion image quality evaluation and optimization: S4.1, use the panchromatic image as a reference, construct a loss function using an image quality evaluation index, and evaluate the quality of the multi-modal information fusion image through the loss function; S4.2, determine whether the loss function reaches a minimum value, if yes, execute S4.3, if not, adjust the corresponding parameters of the filtering processing and the frequency domain transformation in steps S3.1-S3.3, and return to step S3.1; S4.3, output the current multi-modal information fusion image, and complete the multi-modal image fusion.
2. The multi-modal image fusion method according to claim 1, wherein: In step S3.1, the characteristic spectral image, the pseudo-color image and the degree of polarization image are filtered through different scale initial convolution kernels respectively; In step S4.1, the image quality evaluation index comprises image information entropy, edge preservation index and structural similarity index; Step S4.3 further comprises: saving the corresponding parameters of the filtering processing and the frequency domain transformation of steps S3.1-S3.3 of the current multi-modal information fusion image to obtain optimal model parameters for multi-modal image fusion of the target scene.
3. The multi-modal image fusion method of claim 2, wherein, In step S4.1, the loss function is specifically: wherein P represents a reference image, F represents a fused image, H(P) represents an information entropy of the reference image, H(F) represents an information entropy of the fused image, Q abf (P, F) represents an edge preservation index of the fused image F with reference to the reference image P, SSIM(P, F) represents a structural similarity index of the fused image F with reference to the reference image P, and Loss(P, F) is a loss function.
4. The multi-modal image fusion method according to claim 2, wherein: In step S4.1, the image quality evaluation index further comprises image cross-entropy, image conditional entropy, peak signal-to-noise ratio and / or mutual information.
5. The multi-modal image fusion method according to any one of claims 1-4, characterized in that, Step S2 is specifically: S2.1, perform spectral feature analysis based on the spectral radiation curve, find a maximum difference spectrum of reflected irradiance of different target objects, use the spectrum as a characteristic spectrum, and use the spectral image of the corresponding spectrum to obtain a characteristic spectral image; S2.2, based on the spectral image, obtain a pseudo-color image by using PCA principal component analysis or four-channel spectral information weighted mapping or three-channel spectral information mapping; S2.3, based on the polarization image, use Stokes vector to solve polarization information of the target scene, and output a degree of polarization image.
6. A multi-modal image imaging system for acquiring original images and spectral information of a target scene to implement the multi-modal image fusion method of claim 1, comprising: a sub-aperture array compound eye unit and a spectrometer, the sub-aperture array compound eye unit comprising a lens array (1) and a multi-modal split focal plane camera (2); the lens array (1) comprising a central sub-aperture lens (101) and a plurality of edge sub-aperture lenses (102) arranged along the outer periphery thereof; the central sub-aperture lens (101) and the plurality of edge sub-aperture lenses (102) have overlapping fields of view, and the fields of view of the plurality of edge sub-aperture lenses (102) completely cover the field of view of the central sub-aperture lens (101); the central sub-aperture lens (101) and each edge sub-aperture lens (102) are arranged with a spacing to prevent aliasing of the images formed thereby; the corresponding light channels of the central sub-aperture lens (101) and each edge sub-aperture lens (102) are arranged as a panchromatic channel, a spectral channel or a polarization channel to form corresponding sub-aperture images, which are panchromatic images, spectral images and / or polarization images; the multi-modal split focal plane camera (2) comprises an image sensor and a data transmission unit; the data transmission unit is configured to transmit the images captured by the image sensor to an external industrial computer; and the spectrometer is configured to acquire spectral information of the target scene.
7. The multi-modal image imaging system of claim 6, wherein: the spectral channel is provided with a filter corresponding to the imaging area of the focal plane of the image sensor, and the polarization channel is provided with a polarizer corresponding to the imaging area of the focal plane of the image sensor; a neutral filter is arranged on the surface of the polarizer and the imaging area of the panchromatic channel to attenuate the incident light energy, so as to ensure that the light energy of the images formed by the panchromatic channel, the spectral channel and the polarization channel is consistent or matched.
8. The multi-modal image imaging system of claim 7, wherein: the corresponding light channel of the central sub-aperture lens (101) is arranged as a panchromatic channel; the filter of the spectral channel is a narrowband filter, which is arranged on the focal plane of the image sensor by a physical vapor deposition process; the polarizer is a linear polarizer or a circular polarizer, which is transferred to the focal plane of the image sensor by a nanoimprint lithography process; and the neutral filter is an absorption type bandpass filter, which has a working spectral range consistent with that of the image sensor.
9. The multi-modal image imaging system of claim 8, wherein: the original image is a compound eye image, and the panchromatic image, the spectral image and / or the polarization image of the target scene are obtained by image segmentation, which specifically comprises: imaging a positioning calibration board to obtain a calibration image, acquiring the center coordinates and radius parameters of the positioning calibration board in the sub-aperture image formed by each edge sub-aperture lens (102), the radius parameter being the minimum distance from the center coordinates of the positioning calibration board to the edge of the sub-aperture image, and then segmenting the original image according to the center coordinates and radius parameters of the positioning calibration board in each sub-aperture image to obtain the panchromatic image, the spectral image and / or the polarization image.
Citation Information
Patent Citations
Multi-mode fusion imaging method and system based on micro-polarization array
CN115797239A
Intelligent driving target identification method based on polarization visual image fusion
CN120544165A