Multi-modal image fusion method and imaging system
Patent Information
- Application Number
- CN202511272512.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-09-08
AI Technical Summary
In the existing technology, multimodal information image fusion technology has the problems of non-simultaneous and non-homologous image acquisition, which leads to great difficulty in registration, low registration accuracy, inconsistent image edge details and poor fusion effect.
A multimodal image fusion method is adopted to obtain the original image and spectral information of the target scene, perform filtering processing and frequency domain transformation of the characteristic spectral image, pseudo-color image and polarization image, and optimize the loss function of the image quality evaluation index to achieve accurate fusion of multimodal information.
It improves the accuracy and detail retention ability of multimodal information image fusion, enhances the accuracy and comprehensiveness of target detection tasks, reduces the difficulty of image registration, and improves imaging quality.
Smart Images

Figure CN120808093A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the image processing method of optical remote sensing imaging detection, and in particular to a multi-modal image fusion method and an imaging system. BACKGROUND
[0002] In the field of optical remote sensing imaging detection, multi-modal information images play an important role in computer vision tasks such as target classification, detection, positioning, segmentation, etc. According to the target attributes, different modal information has different effects on the algorithms of different tasks. Panchromatic images provide target position, intensity and other information for various tasks, spectral images provide spectral feature information such as target chemical composition, and polarization images provide information such as target surface features and material details. Due to the significant complementary characteristics of multi-modal information, its fusion technology has gradually become a key means to break through the limitations of traditional optical remote sensing imaging detection.
[0003] Current multi-modal information image fusion technology relies on time-sharing imaging systems or multi-device acquisition for target scene image capture, resulting in non-simultaneous and non-simultaneous acquisition of multi-modal information images of the target scene. In subsequent image processing, different modal images are difficult to register, have low registration accuracy, and have inconsistent image edge details, resulting in poor fusion effect and limited application. SUMMARY
[0004] The purpose of the present application is to solve the problems in the prior art that target scene image capture relies on time-sharing imaging systems or multi-device acquisition, resulting in large registration difficulty, low registration accuracy, inconsistent image edge details, and poor fusion effect of different modal images, and to provide a multi-modal image fusion method and imaging system.
[0005] To achieve the above purpose, the technical solution provided by the present application is as follows: A multi-modal image fusion method, characterized in that it comprises the following steps: S1, original image and characteristic spectrum acquisition: Obtain the original image of the target scene, which includes the spectral image and the polarization image of the target scene; simultaneously acquire the spectral information of the target scene and draw the spectral radiation curve; S2, multi-modal image preprocessing: Obtain the characteristic spectral image based on the spectral radiation curve analysis; obtain the pseudo-color image based on the spectral image; and obtain the polarization degree image based on the polarization image; S3, multi-modal image fusion: S3.1, filter processing is performed on the characteristic spectral image, the pseudo-color image and the polarization degree image to obtain corresponding base images, respectively; S3.2 subtracting the corresponding base maps after frequency domain transformation of the characteristic spectral image, the pseudo-color image, the polarization degree image and the corresponding base maps, and then re-transforming to the spatial domain to obtain the corresponding detail maps; S3.3 obtaining a fusion base map by superimposing the base maps in the spatial domain, obtaining a fusion detail map by superimposing the detail maps in the spatial domain, adding the fusion base map and the fusion detail map after frequency domain transformation, and then inverse transforming to the spatial domain to obtain a multi-modal information fusion image.
[0006] Further, in step 1, the original image further includes a panchromatic image of the target scene. The multi-modal image fusion method further includes: Step S4, image quality evaluation and optimization: S4.1 constructing a loss function by using an image quality evaluation index based on the panchromatic image, and evaluating the quality of the multi-modal information fusion image by the loss function; S4.2 determining whether the loss function reaches a minimum value, if yes, executing S4.3, and if no, adjusting the corresponding parameters of the filtering processing and the frequency domain transformation in steps S3.1-S3.3, and returning to step S3.1; S4.3 outputting the current multi-modal information fusion image to complete the multi-modal image fusion.
[0007] Further, in step S3.1, the characteristic spectral image, the pseudo-color image and the polarization degree image are respectively filtered by different scale initialization convolution kernels; In step S4.1, the image quality evaluation index includes an image information entropy, an edge preservation index and a structural similarity index. Step S4.3 further includes: saving the corresponding parameters of the filtering processing and the frequency domain transformation in steps S3.1-S3.3 of the current multi-modal information fusion image to obtain optimal model parameters for multi-modal image fusion of the target scene.
[0008] Further, in step S4.1, the loss function is specifically:
[0009] Wherein, P represents a reference image, F represents a fusion image, H(P) represents an information entropy of the reference image, H(F) represents an information entropy of the fusion image, Q abf (P,F) represents an edge preservation index of the fusion image F with the reference image P as the reference, SSIM(P,F) represents a structural similarity index of the fusion image F with the reference image P as the reference, is the loss function.
[0010] Further, in step S4.1, the image quality evaluation index further includes image cross entropy, image conditional entropy, peak signal-to-noise ratio and / or mutual information.
[0011] Further, step S2 is specifically: S2.1, performing spectral feature analysis based on the spectral radiation curve, finding a maximum difference spectrum of reflected irradiance of different target objects, taking the maximum difference spectrum as a characteristic spectrum, and using a spectral image corresponding to the spectrum to obtain a characteristic spectral image by difference; S2.2, based on the spectral image, using PCA principal component analysis or four-channel spectral information weighted mapping or three-channel spectral information mapping to obtain a pseudo-color image; S2.3, based on the polarization image, using Stokes vector to solve the polarization information of the target scene, and outputting a polarization degree image.
[0012] Meanwhile, the application provides a multi-modal image imaging system for collecting original images and spectral information of a target scene to realize the multi-modal image fusion method, and the special feature is that the system comprises a sub-aperture array compound eye unit and a spectrometer, the sub-aperture array compound eye unit comprises a lens array and a multi-modal focal plane camera; the lens array comprises a central sub-aperture lens and a plurality of edge sub-aperture lenses arranged along the outer periphery; the field of view of the central sub-aperture lens and the plurality of edge sub-aperture lenses overlap, and the field of view of the plurality of edge sub-aperture lenses completely covers the field of view of the central sub-aperture lens; a spacing is arranged between the central sub-aperture lens and each edge sub-aperture lens to prevent aliasing of the images formed thereby; the light channels corresponding to the central sub-aperture lens and each edge sub-aperture lens are arranged as panchromatic channels or spectral channels or polarization channels to form corresponding sub-aperture images, and the sub-aperture images are panchromatic images, spectral images and / or polarization images; the multi-modal focal plane camera comprises an image sensor and a data transmission unit; the data transmission unit is used to transmit the images captured by the image sensor to an external industrial computer; and the spectrometer is used to collect spectral information of the target scene.
[0013] Further, a filter is arranged on the imaging area of the focal plane of the image sensor corresponding to the spectral channel, and a polarizer is arranged on the imaging area of the focal plane of the image sensor corresponding to the polarization channel; A neutral density filter is arranged on the surface of the polarizer and the imaging area of the panchromatic channel to attenuate the incident light energy, so as to ensure that the light energy of the panchromatic channel, the spectral channel and the polarization channel is consistent or matched.
[0014] Further, the light channel corresponding to the center sub-aperture lens is set as a full-color channel; the filter of the spectral channel is a narrow-band filter, which is set on the focal plane of the image sensor through a physical vapor deposition process; the polarizer is a linear polarizer or a circular polarizer, which is transferred to the focal plane of the image sensor through a nano-imprint lithography process; and the neutral filter is an absorption type band-pass filter, whose working spectral range is consistent with that of the image sensor.
[0015] Further, the original image is a compound eye image, and the full-color image, the spectral image and the polarization image of the target scene are obtained through image segmentation, which is specifically: imaging the positioning calibration board to obtain a calibration image, obtaining the center coordinates and the radius parameter of the positioning calibration board in the sub-aperture image formed by each edge sub-aperture lens, the radius parameter being the minimum value of the distance between the center coordinates of the positioning calibration board and the edge of the sub-aperture image, and then segmenting the original image according to the center coordinates and the radius parameter of the positioning calibration board in each sub-aperture image to obtain the full-color image, the spectral image and / or the polarization image.
[0016] The beneficial effects of the present application are: 1. The multi-modal image fusion method of the present application collects the spectral radiation curve of the target scene, analyzes the spectral characteristics, reconstructs the characteristic spectral image of the target scene, and fuses it with the pseudo-color image, which can more accurately calculate and retain image edge details compared with the existing image fusion method, and improves the multi-modal information image fusion effect.
[0017] 2. In the multi-modal image fusion method of the present application, the characteristic spectral image, the pseudo-color image and the polarization degree image are used as source images, and the full-color image is used as a reference image, and the multi-spectral spectral information, polarization information and spatial information of the target scene are fused to maximize the retention of detail texture information. The multi-modal source image compensates for the small environmental light polarization component and poor detection effect in the existing polarization detection technology, effectively improving the accuracy and comprehensiveness of the target detection task.
[0018] 3. The multi-modal image fusion method of the present application also uses a variety of image evaluation indexes to evaluate the quality of the multi-modal information fusion image, and compares it with the source image and the reference image, which can comprehensively evaluate the fusion effect, dynamically adjust the initialization convolution kernel scale through the loss function, and achieve the best fusion effect with small parameter quantity.
[0019] 4. In the sub-aperture array compound eye unit of the present application, the center sub-aperture lens and each edge sub-aperture lens correspond to different information images, and their design parameters are completely consistent, which can realize the simultaneous and homologous acquisition of different modal information when imaging the target scene, reduces the difficulty of image registration at the hardware level, and maximizes the retention of image edge detail information.
[0020] 5. The sub-aperture array compound eye unit of the present application also sets a neutral filter in the imaging area of the panchromatic channel and integrates a polarizer on the focal plane of the image sensor, for attenuating the incident light energy, thereby adjusting the inconsistency of the spectral and panchromatic channels in the light energy within the same exposure time, effectively improving the imaging quality of simultaneous and co-sourced acquisition of different modal information.
[0021] 6. The sub-aperture array compound eye unit of the present application adopts physical vapor deposition and nano-imprint lithography process to integrate filters and polarizers of different parameters on a large array sensor, and according to the different spectral responses of specific target scenes, filter structures with different center wavelengths and different half-widths can be selected. In addition, the spectral range or polarization angle on the large array sensor can be modified according to actual application needs, and the dynamic adaptability of the sub-aperture array compound eye unit is greatly improved. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is a structural schematic diagram of a sub-aperture array compound eye unit in an embodiment of a multi-modal image imaging system of the present application; Figure 2 is a field of view schematic diagram of a lens array in a sub-aperture array compound eye unit in an embodiment of a multi-modal image imaging system of the present application; Figure 3 is a focal plane schematic diagram of a multi-modal focal plane camera in an embodiment of a multi-modal image imaging system of the present application; Figure 4 is a positioning calibration plate imaging schematic diagram of step 1 in an embodiment of a multi-modal image fusion method of the present application; Figure 5 is a radius parameter acquisition schematic diagram of step 1 in an embodiment of a multi-modal image fusion method of the present application; Figure 6 is a flowchart of step 2 in an embodiment of a multi-modal image fusion method of the present application; Figure 7 is a flowchart of step 3 in an embodiment of a multi-modal image fusion method of the present application; Figure 8 is a flowchart of step 4 in an embodiment of a multi-modal image fusion method of the present application.
[0023] BRIEF DESCRIPTION OF DRAWINGS 1-lens array, 101-center sub-aperture lens, 102-edge sub-aperture lens, 2-multi-modal focal plane camera. DETAILED DESCRIPTION
[0024] The present application provides a multi-modal image imaging system, which comprises a sub-aperture array compound eye unit and a spectrometer, the spectrometer is used for collecting spectral information of a target scene, and the structure of the sub-aperture array compound eye unit is as shown in Figure 1 the drawing, mainly comprising a lens array 1 and a multi-modal focal plane camera 2.
[0025] The lens array 1 includes a central sub-aperture lens 101 and a plurality of edge sub-aperture lenses 102 arranged along its periphery. The design parameters (optical parameters and structural parameters) of the central sub-aperture lens 101 and each edge sub-aperture lens 102 are completely consistent. In this embodiment, the plurality of edge sub-aperture lenses 102 are arranged in a rectangular shape, and are respectively located at the top corners and four sides of the rectangle. Figure 2 As shown, each edge sub-aperture lens 102 has a maximum field of view overlap with the central sub-aperture lens 101. With the central sub-aperture lens 101 as a reference, the overlapping fields of view of the edge sub-aperture lenses 102 can completely cover the field of view of the central sub-aperture lens 101. A certain distance exists between the central sub-aperture lens 101 and each edge sub-aperture lens 102, so that the images formed on the image sensor are spaced apart and aliasing is prevented. In other embodiments of the present invention, multiple edge sub-aperture lenses 102 can also be arranged in a regular hexagon, regular octagon, circle, or the like, and positioned on the sides and / or corners of the regular hexagon, regular octagon, or the like.
[0026] The multimodal focal plane camera 2 includes a data transmission unit and an image sensor. The data transmission unit is used to transmit the image captured by the image sensor to the industrial computer. The data transmission unit provides at least two data transmission methods. In this embodiment, a USB 3.0 interface or a GigE interface is selected for data transmission. In other embodiments of the present invention, a data transmission interface such as a CameraLink interface or a CoaXPress interface can also be selected. In other embodiments of the present invention, the multimodal focal plane camera 2 also includes a power supply system. The power supply system uses a common power supply interface, which can be a DC power I / O interface, or can be powered by a USB 3.0 or CoaXPress interface to power the image sensor and the data transmission unit.
[0027] like Figure 3 As shown, in the multimodal focal plane camera 2, the optical channels corresponding to the central sub-aperture lens 101 and each edge sub-aperture lens 102 are set to be panchromatic channels, spectral channels, or polarization channels. Filters are set on the focal planes of the image sensors corresponding to the spectral channels, and polarizers are set on the focal planes of the image sensors corresponding to the polarization channels. The filters or polarizers can completely cover the imaging area of the corresponding single central sub-aperture lens 101 or edge sub-aperture lens 102, and are used to obtain spectral images and polarization images, respectively.
[0028] The filter is integrated using a physical vapor deposition process. In this embodiment, the filter is deposited on the focal plane of the image sensor through plasma sputtering, forming a multilayer film stack. The filter is a narrowband filter, with the center wavelength and corresponding half-width (FWHM) selected based on the target scene.
[0029] The polarizer adopts a nano-imprint lithography process to transfer the nano-structure grating of the polarizer to the focal plane by mechanical imprinting. After mechanical imprinting of the nano-structure grating, a neutral filter should also be integrated on the surface thereof by a physical vapor deposition process, for attenuating the incident light energy, so that the light energy of the spectral information and the polarization information is consistent or matched in the same exposure time during imaging.
[0030] As shown in Figure 3 In this embodiment, the imaging areas located at the upper left, upper right, lower left and lower right of the focal plane are provided with filters, the center wavelengths are respectively selected as red light (R: 650 nm), green light (G: 532 nm), blue light (B: 472 nm) and near-infrared light (NIR: 730 nm), and the four spectral bands are selected as the center wavelengths of the spectral information, and the half-height width is uniform at 20 nm. At the same time, four linear polarizers are arranged to obtain polarization information, which are respectively located in the imaging areas on the upper, right, lower and left of the focal plane, and the polarization angles are respectively 0°, 45°, 90° and 135°, and the light channels of the edge sub-aperture lens 102 correspond to the polarization channels. In other embodiments of the present application, other center wavelengths and corresponding half-height widths can be selected according to the focal plane response and task requirements, or linear polarizers with other polarization angles (30°, 60°) can be selected, or circular polarizers can be used to obtain polarization information.
[0031] The light channel corresponding to the center sub-aperture lens 101 is set as a panchromatic channel (PAN), and the imaging area of the focal plane corresponding to the center sub-aperture lens 101 is integrated with a neutral filter for attenuating the incident light energy, but the transmittance should be less than that of the neutral filter on the surface of the polarizer, so as to ensure that the light energy of the panchromatic channel, the spectral channel and the polarization channel is consistent. The above-mentioned neutral filters are all set as absorption type band-pass filters, and the working spectral band is consistent with the working spectral band of the image sensor, and the transmittance can be adjusted according to the number of deposited films.
[0032] The multi-modal image fusion method of the present application fuses the original images to output a single multi-modal information fusion image, wherein the original images are multi-modal images collected by the sub-aperture array compound eye unit, including a plurality of sub-aperture images corresponding to the center sub-aperture lens 101 and the edge sub-aperture lens 102 respectively, and the plurality of sub-aperture images respectively contain different modal information. The multi-modal image fusion method specifically includes five stages of multi-modal image registration, characteristic spectrum acquisition, multi-modal image preprocessing, multi-modal image fusion and fusion image quality evaluation, as shown below: 1. Multi-modal image registration; 1.1 As shown in Figure 4As shown, a positioning calibration board is drawn, and a circular mark point located in the center and a right-angle mark line equidistant from the center of the mark point are arranged on the positioning calibration board. The positioning calibration board is imaged by using the sub-aperture array compound eye unit, and the distance between the sub-aperture array compound eye unit and the positioning calibration board is adjusted, so that the center sub-aperture image contains the entire positioning calibration board, and the four right-angle mark lines are located at the edges of the center sub-aperture image.
[0033] The positioning calibration board is used to confirm the field of view overlap area of the edge sub-aperture lens 102 and the center sub-aperture lens 101, and can be arranged in any form or pattern.
[0034] 1.2As Figure 5 shown, the center coordinates of the positioning calibration board in each sub-aperture image are calculated and recorded, and the center of the positioning calibration board in the spectral image (top left, top right, bottom left, bottom right) is taken as the center of a circle, and the minimum distance between the center and the edge of the sub-aperture image is taken as the radius of the circle. The radius parameter is recorded.
[0035] 1.3The center coordinates and radius parameters of the positioning calibration board in each sub-aperture image are saved, and the center coordinates and radius parameters of the positioning calibration board in the sub-aperture image corresponding to the polarization image are obtained in the same way.
[0036] 2. Raw image and feature spectrum acquisition; as Figure 6 shown, the following steps are included: 2.1 The raw image of the target scene is obtained by using the sub-aperture array compound eye unit. At the same time, during the imaging process, the spectral information of the target scene is collected by using a spectrometer, and the spectral information is uploaded to the industrial computer at the processing end.
[0037] In this embodiment, the target scene is an outdoor scene, and the target objects in the scene include plants, artificial plants, and metal products of the same color and luster. Other embodiments of the present application can be used for spectral information collection for different target scenes and different target objects.
[0038] 2.2 The industrial computer draws a spectral radiation curve based on the collected spectral information, wherein the horizontal axis is the wavelength, the wavelength range is the working spectral range of the multi-modal focal plane camera 2, and the vertical axis is the reflectance of the target object in the scene collected by the spectrometer.
[0039] 3. Multi-modal image preprocessing; as Figure 7 shown, the following steps are included: 3.1 Using the center coordinates and radius parameters of the calibration board positioned in each sub-aperture image in step 1, the original image is segmented to obtain the overlapping part of the field of view in the sub-aperture image. In this embodiment, the original image is segmented to obtain a panchromatic image (PAN) corresponding to the center sub-aperture lens 101 and four spectral images (R, G, B, NIR) and four angle polarization images (0°, 45°, 90°, 135°) corresponding to the edge sub-aperture lens 102.
[0040] 3.2 Based on the spectral radiation curve, spectral feature analysis is performed to find the maximum difference spectrum of the reflectance of different target objects, which is used as the characteristic spectrum. The spectral image corresponding to the spectrum is used to obtain the characteristic spectral image.
[0041] In this embodiment, the spectral radiation curves of plants, artificial plants and metal products of the same color have the maximum difference in the near-infrared (NIR) and green (G) spectrum, so the near-infrared image and the green spectrum image are selected to obtain the characteristic spectral image of the target scene. In other embodiments of the present application, spectral images of other spectrum can be selected according to different target objects.
[0042] 3.3 Based on the four-spectrum spectral image, the PCA principal component analysis method is used to calculate the covariance matrix of the four-spectrum spectral image, and the singular value decomposition is performed to obtain the eigenvalues and eigenvectors corresponding to the four spectrum. The first three principal components containing the most eigenvalues are selected, the eigenvectors corresponding to the first three principal components are orthogonal to each other, the four-spectrum spectral image is projected onto the eigenvectors corresponding to the selected mutually orthogonal principal components, and three principal component images are obtained. The three principal component images are mapped to the RGB color space to obtain a pseudo-color image.
[0043] In this embodiment, the principal component images corresponding to the first three principal components containing most of the eigenvalues are defined as the first principal component image, the second principal component image and the third principal component image, respectively. The first principal component image is mapped to the R channel, the second principal component image is mapped to the G channel, and the third principal component image is mapped to the B channel to obtain a pseudo-color image.
[0044] In other embodiments of the present application, in addition to the PCA principal component analysis method, existing four-channel spectral information weighted mapping, three-channel spectral information mapping and other methods can be used to obtain a pseudo-color image. When the PCA principal component analysis method is used to obtain a pseudo-color image, the above three principal component images can be mapped to other channels of the RGB color space, respectively. In addition to the RGB color space, other color space systems such as HSV and HIS can also be selected for mapping.
[0045] 3.4 Based on the four-angle polarization image, the Stokes vector is used to solve the polarization information of the target scene, and a polarization degree image is output.
[0046] In this embodiment, the four-angle polarization image is a linear polarization image of 0°, 45°, 90°, and 135°. In other embodiments of the present application, linear polarization images of 30° and 60° or circular polarization images can also be selected to solve the polarization information.
[0047] The above characteristic spectral image, pseudo-color image, and polarization degree image are used as the source images of the multi-modal image fusion model in step 4.
[0048] 4. Multi-modal image fusion, as shown in Figure 8 The source images are processed using a multi-modal image fusion model, which includes the following steps: 4.1 The characteristic spectral image, pseudo-color image, and polarization degree image are respectively filtered using different sizes of initial convolution kernels, so as to fully utilize the different scale feature information of different modal images, and obtain corresponding base maps, which are respectively denoted as the first base map, second base map, and third base map. The base map retains the low-frequency information in the source image, i.e., the overall structure and brightness distribution information, and can effectively extract the structure edge of the target scene. In this embodiment, the initial convolution kernel type is a bilateral filter convolution kernel, and the sizes of the initial convolution kernels corresponding to the characteristic spectral image, pseudo-color image, and polarization degree image are 3x3, 9x9, and 15x15, respectively. In other embodiments of the present application, other types of convolution kernels can be selected, such as Gaussian filter, low-pass filter, and high-pass filter, and other sizes of convolution kernels can also be selected, such as 5x5, 7x7, and 11x11.
[0049] 4.2 The characteristic spectral image, pseudo-color image, and polarization degree image and the corresponding base maps are subjected to frequency domain transformation, i.e., from spatial domain to frequency domain. The corresponding base maps are subtracted from the characteristic spectral image, pseudo-color image, and polarization degree image in the frequency domain, and the images are re-transformed to the spatial domain to obtain corresponding detail maps, which are respectively denoted as the first detail map, second detail map, and third detail map. The detail map retains the high-frequency information in the source image, i.e., the detail texture information.
[0050] In this embodiment, the fast Fourier transform method is used when the image is transformed from the spatial domain to the frequency domain or from the frequency domain to the spatial domain. In other embodiments of the present application, other methods such as discrete Fourier transform and inverse discrete Fourier transform, wavelet transform, and Laplace transform can also be used.
[0051] 4.3 superimpose the first base map, the second base map and the third base map in the spatial domain to obtain a fused base map, superimpose the first detail map, the second detail map and the third detail map in the spatial domain to obtain a fused detail map. Perform frequency domain transformation on the fused base map and the fused detail map respectively, add them together, and then perform inverse transformation to the spatial domain to obtain a multi-modal information fused image.
[0052] 5. Fused image quality evaluation and optimization; 5.1 Establish a loss function based on image quality evaluation indicators. Calculate the loss function using image quality evaluation indicators to evaluate the quality of the multi-modal information fused image. The smaller the loss function value is, the higher the image fusion quality is.
[0053] In this embodiment, the loss function includes image information entropy (Entropy), edge preservation index (Qabf) and structural similarity index (SSIM), which are specifically as follows: Image information entropy is a no-reference image quality evaluation indicator, which is used to evaluate the richness of the information amount of the multi-modal information fused image. Edge preservation index is a reference image quality evaluation indicator, which is used to evaluate the ability of the multi-modal information fused image to retain important edge information in the source image. Structural similarity index is a reference image quality evaluation indicator, which is used to evaluate the local similarity of the multi-modal information fused image in terms of brightness, contrast and structure. The required reference image is a panchromatic image (PAN) obtained by a sub-aperture array compound eye unit.
[0054] Loss function Specifically,
[0055] That is,
[0056] Wherein, P represents the reference image, F represents the fused image, H(P) represents the reference image information entropy, H(F) represents the fused image information entropy, Q abf (P, F) represents the edge preservation index of the fused image F with the reference image P as the reference, and SSIM(P, F) represents the structural similarity index of the fused image F with the reference image P as the reference.
[0057] 5.2 Determine whether the loss function reaches the minimum value. If yes, execute step 5.3; if no, adjust the step through parameter feedback, modify the scale size of the initialized convolution kernel in step 4.1, dynamically adjust the related parameters of the frequency domain transformation in steps 4.2 and 4.3, and return to step 4.1. 5.3 output the current multi-modal information fusion image, complete the multi-modal image fusion, at the same time, save the initialization convolution kernel scale corresponding to the current multi-modal information fusion image and the related parameters of the frequency domain transformation in steps 4.2 and 4.3, and take them as the optimal model parameters of the target scene.
[0058] In the fusion image quality evaluation, in addition to the image quality evaluation indexes used in the loss function, more image quality evaluation indexes can be used to evaluate the quality of the fusion image. Other embodiments of the application can also include: Image cross-entropy: a reference image quality evaluation index, used to evaluate the difference between the fusion image and the reference image; Image conditional entropy: a reference image quality evaluation index, used to evaluate the complementarity of different modal source image information under the condition of known reference image; Peak signal-to-noise ratio: a reference image quality evaluation index, used to evaluate the distortion of the fusion image relative to the reference image; Mutual information: a reference image quality evaluation index, used to evaluate the amount of shared information between the fusion image and the source image.
[0059] The above multi-modal image fusion method images the target scene by combining the aperture and the focal plane, and different modal information images are captured by the same image sensor at the same time, which greatly reduces the image registration difficulty. Through spectral and polarization information fusion, different types of targets can be effectively distinguished, full-color, spectral and polarization information can be acquired at the same time and from the same source, and the fusion of multi-modal information images can be effectively completed, which provides an important technical means for optical remote sensing imaging detection, and has great application potential in target detection and classification tasks.
Claims
1. A multimodal image fusion method, characterized in that: The following steps are involved: S1, original image and characteristic spectrum acquisition: Acquire the original image of the target scene, which includes the spectral image and polarization image of the target scene; at the same time, collect the spectral information of the target scene and draw the spectral radiation curve; S2, multimodal image preprocessing: Obtain characteristic spectrum images based on spectral radiation curve analysis; Obtain a pseudo-color image based on the spectral image; obtain a polarization degree image based on the polarization image; S3, multimodal image fusion: S3.1 performs filtering processing on the characteristic spectrum image, the pseudo-color image, and the polarization degree image to obtain corresponding base images respectively; S3.2 performs frequency domain transformation on the characteristic spectrum image, pseudo color image, polarization degree image and their corresponding base images, subtracts them accordingly, and then re-transforms them into the spatial domain to obtain the corresponding detail image; S3.3 superimposes each base image in the spatial domain to obtain a fused base image, and superimposes each detail image in the spatial domain to obtain a fused detail image. The fused base image and the fused detail image are then transformed in the frequency domain respectively and added together, and then inversely transformed to the spatial domain to obtain a multimodal information fusion image.
2. The multimodal image fusion method according to claim 1, characterized in that: In step 1, the original image also includes a full-color image of the target scene; The multimodal image fusion method further includes: Step S4, fusion image quality evaluation and optimization: S4.1 uses full-color images as a benchmark and image quality evaluation indicators to construct a loss function, and uses the loss function to evaluate the quality of multimodal information fusion images; S4.2 determines whether the loss function has reached a minimum value. If so, execute S4.3; if not, adjust the corresponding parameters of the filtering process and frequency domain transformation in steps S3.1-S3.3 and return to step S3.1; S4.3 outputs the current multimodal information fusion image to complete the multimodal image fusion.
3. The multimodal image fusion method according to claim 2, characterized in that: In step S3.1, the characteristic spectrum image, pseudo color image, and polarization degree image are filtered using initialization convolution kernels of different scales. In step S4.1, the image quality evaluation indicators include image information entropy, edge preservation index and structural similarity index; Step S4.3 also includes: saving the corresponding parameters of the filtering processing and frequency domain transformation of steps S3.1-S3.3 corresponding to the current multimodal information fusion image, and obtaining the optimal model parameters for multimodal image fusion of the target scene.
4. The multimodal image fusion method according to claim 3, characterized in that: In step S4.1, the loss function is specifically: ; Where P represents the reference image, F represents the fused image, H(P) represents the information entropy of the reference image, H(F) represents the information entropy of the fused image, and Q abf (P, F) represents the edge preservation index of the fused image F with reference to the reference image P, SSIM(P, F) represents the structural similarity index of the fused image F with reference to the reference image P, is the loss function.
5. The multimodal image fusion method according to claim 3, characterized in that: In step S4.1, the image quality evaluation index further includes image cross entropy, image conditional entropy, peak signal-to-noise ratio and / or mutual information.
6. The multimodal image fusion method according to any one of claims 1 to 5, characterized in that: Step S2 is specifically as follows: S2.1 performs spectral feature analysis based on the spectral radiation curve, finds the spectral band with the maximum difference in reflected irradiance of different target objects, uses this as the characteristic spectrum, and uses the spectral images of the corresponding spectral bands for subtraction to obtain the characteristic spectrum image; S2.2 is based on spectral images and uses PCA principal component analysis or four-channel spectral information weighted mapping or three-channel spectral information mapping to obtain pseudo-color images; S2.3 uses the Stokes vector to solve the polarization information of the target scene based on the polarization image and outputs a polarization degree image.
7. A multimodal imaging system for acquiring original images and spectral information of a target scene to implement the multimodal image fusion method of claim 1, characterized in that: It includes a sub-aperture array compound eye unit and a spectrometer, wherein the sub-aperture array compound eye unit includes a lens array (1) and a multi-modal focal plane camera (2); The lens array (1) comprises a central sub-aperture lens (101) and a plurality of edge sub-aperture lenses (102) arranged along the periphery of the central sub-aperture lens (101); The fields of view of the central sub-aperture lens (101) and the plurality of edge sub-aperture lenses (102) overlap, and the fields of view of the plurality of edge sub-aperture lenses (102) completely cover the field of view of the central sub-aperture lens (101); A spacing is provided between the central sub-aperture lens (101) and each edge sub-aperture lens (102) so that the images formed therebetween are free of aliasing; The optical channels corresponding to the central sub-aperture lens (101) and each edge sub-aperture lens (102) are set as full-color channels, spectral channels, or polarization channels, for forming corresponding sub-aperture images, where the sub-aperture images are full-color images, spectral images, and / or polarization images; The multimodal focal plane camera (2) comprises an image sensor and a data transmission unit; The data transmission unit is used to transmit the image captured by the image sensor to an external industrial computer; The spectrometer is used to collect spectral information of the target scene.
8. The multimodal imaging system according to claim 7, characterized in that: The spectral channel is provided with a filter in the imaging area of the focal plane of the image sensor, and the polarization channel is provided with a polarizer in the imaging area of the focal plane of the image sensor; Neutral density filters are provided on the surface of the polarizer and the imaging area of the full-color channel to attenuate the incident light energy to ensure that the light energy of the full-color channel, spectral channel and polarization channel imaging is consistent or matched.
9. The multimodal imaging system according to claim 8, characterized in that: The optical channel corresponding to the central sub-aperture lens (101) is set as a full-color channel; the filter set in the spectral channel is a narrow-band filter, which is set on the focal plane of the image sensor through a physical vapor deposition process; the polarizer is a linear polarizer or a circular polarizer, which is transferred to the focal plane of the image sensor through a nanoimprint lithography process; and the neutral filter is an absorption-type bandpass filter, and its working spectrum is consistent with the working spectrum of the image sensor.
10. The multimodal imaging system according to claim 9, characterized in that: The original image is a compound eye image, and a full-color image, a spectral image, and a polarization image of the target scene are obtained by image segmentation. The image segmentation is specifically as follows: imaging the positioning calibration plate to obtain a calibration image, obtaining the center coordinates and radius parameters of the positioning calibration plate in the sub-aperture images formed by each edge sub-aperture lens (102) on the calibration image, wherein the radius parameter is the minimum value of the distance between the center coordinate of the positioning calibration plate and the edge of the sub-aperture image, and then segmenting the original image according to the center coordinates and radius parameters of the positioning calibration plate in each sub-aperture image to obtain a full-color image, a spectral image, and / or a polarization image.
Citation Information
Patent Citations
Multi-mode fusion imaging method and system based on micro-polarization array
CN115797239A
Intelligent driving target identification method based on polarization visual image fusion
CN120544165A
3D scene engineering simulation and real-life scene fusion system
US10769843B1