A real-time enhanced display system for multispectral image fusion and a control method thereof
Patent Information
- Application Number
- CN202610728996.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-26
- Publication Date
- 2026-09-15
AI Technical Summary
[0003]但是,现有多光谱图像融合方法大多侧重于对多源图像进行直接配准、简单加权或统一特征拼接处理,对不同谱段图像之间在采集时间、成像噪声、镜头畸变、空间视场差异和局部质量波动方面的差异考虑不足,容易导致输入数据之间存在时序不一致、空间映射偏差和局部信息失真等问题
本发明围绕多光谱图像从原始采集、同步重排、校正映射、矩阵建模、特征融合到重建显示的完整处理链路进行统一设计,相较于现有技术中多源图像直接配准、统一拼接或整图增强的处理方式,能够在多谱段图像输入一致性、局部质量表达能力、跨谱段融合深度以及最终增强显示稳定性方面形成更完整的技术支撑。通过对可见光图像、近红外图像、中波红外图像和长波红外图像进行基于时间偏差的帧号重排,并剔除超过预设容差范围的无效图像帧,可以使进入后续处理流程的多谱段图像在时序上保持一致,避免不同采集时刻图像混入融合过程。通过对各谱段图像依次执行坏点替换、暗电流校正、固定模式噪声去除、镜头畸变校正以及统一空间基准映射,可以减少由于成像器件差异、镜头成像偏差和视场不统一所带来的图像失真,使后续分析对象建立在统一视场基础上。
Smart Images

Figure CN122761111A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a real-time enhanced display system and its control method for multispectral image fusion. Background Technology
[0002] Multispectral image fusion enhancement technology is an important research direction in image processing, intelligent sensing, and display control, and is widely used in scenarios such as complex environment perception, target recognition, assisted observation, and real-time display. Existing technologies typically improve the comprehensive representation of brightness, texture, and thermal response information in a single spectral band image through joint processing of visible light and infrared images, enabling more complete scene display results under conditions of low illumination, occlusion interference, complex backgrounds, or significant target radiation characteristics. With the increasing application of multi-source imaging devices, the collaborative fusion of multispectral band images, including visible light, near-infrared, mid-wave infrared, and long-wave infrared, has become an important development direction for real-time enhancement display technology.
[0003] However, most existing multispectral image fusion methods focus on direct registration, simple weighting, or uniform feature stitching of multi-source images, failing to adequately consider the differences between images in different spectral bands in terms of acquisition time, imaging noise, lens distortion, spatial field of view differences, and local quality fluctuations. This easily leads to problems such as temporal inconsistencies, spatial mapping deviations, and local information distortion among input data. Especially when the synchronization control of multispectral images is insufficient, invalid frames are often mixed into the fusion process, thus affecting the stability of subsequent fusion results. At the same time, after image preprocessing, existing technologies mostly adopt a unified analysis method at the whole image level, lacking a local quality modeling mechanism based on a fixed grid. It is difficult to describe and organize the image quality status at different locations in a targeted manner, resulting in the inability to fully distinguish and utilize the effective local information in images of different spectral bands. In addition, the implementation methods of multispectral image feature fusion in existing technologies mainly rely on ordinary convolution extraction, simple feature stitching, or conventional attention mechanisms, failing to adequately explore the hierarchical correlation, window correlation, and cross-spectral matrix transfer relationships between multispectral grid quality matrices, making it difficult to form a unified, regular fusion feature expression suitable for subsequent reconstruction. Meanwhile, existing enhancement display methods typically display the fused features or fused image directly as a whole during the fusion result output stage. This lacks reconstruction mesh plane mapping and neighborhood unfolding processing oriented towards a common reference field of view, which can easily lead to problems such as discontinuous transitions of local details, unstable spatial correspondences, and insufficient enhancement of display results.
[0004] Therefore, how to provide a real-time enhanced display system and its control method for multispectral image fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a real-time enhanced display system and its control method for multispectral image fusion. This invention achieves regular fusion and real-time enhanced display of multispectral images under a unified field of view by constructing a collaborative control mechanism for multispectral image synchronous rearrangement, correction mapping, grid quality modeling, CrossFormer fusion processing and reconstruction display, thereby improving the continuity, stability and display integrity of the fusion results.
[0006] A real-time enhanced display control method for multispectral image fusion according to an embodiment of the present invention includes the following steps: Step 1: Collect multispectral raw observation units in the same scene, and rearrange the frame numbers according to the time deviation between timestamps to form a multispectral synchronous raw frame group; Step 2: Perform bad pixel replacement, dark current correction and fixed pattern noise removal on each spectral band image frame in the multispectral synchronous original frame group, and perform lens distortion correction and unified spatial reference mapping to obtain a unified field of view image group. Step 3: Divide the unified field-of-view image group into blocks according to a fixed grid, and generate a corresponding grid quality matrix for each spectral band image; Step 4: Input various mesh quality matrices into the improved CrossFormer model, and perform fusion processing through the matrix mapping module, hierarchical analysis module, cross fusion module and tensor generation module to generate a unified fused feature tensor; Step 5: Perform tensor splitting on the unified fusion feature tensor, map it to the reconstructed mesh plane corresponding to the common reference field of view, and perform neighborhood expansion processing to obtain the enhanced display fusion map.
[0007] Optionally, step one specifically includes: The multispectral raw observation unit under the same scene is collected. The multispectral raw observation unit includes visible light image, near infrared image, mid-wave infrared image, long-wave infrared image, as well as timestamp, intrinsic parameter matrix, extrinsic parameter matrix and distortion parameter corresponding to each spectral band image. A unified master clock is selected as the time reference. The frame numbers of each spectral segment image are rearranged according to the time deviation between the timestamp of each spectral segment image and the unified master clock. The frame number rearrangement involves determining the spectral segment image frames whose time deviation exceeds the preset tolerance range as invalid frames and removing them. The retained image frames of each spectral band, along with their corresponding timestamps, intrinsic parameter matrices, extrinsic parameter matrices, and distortion parameters, form a multispectral synchronous original frame group.
[0008] Optionally, step two specifically involves: Bad pixel detection is performed on each spectral band image frame in the multispectral synchronous original frame group. The bad pixel detection is to count the abnormal pixel positions where the gray value of the pixel exceeds the preset deviation range of the gray median value of the preset neighborhood size, mark the abnormal pixel positions as bad pixel positions, and replace the bad pixel positions with the median value of the neighborhood of the bad pixel positions. Subtract the dark field reference value of the corresponding spectral imaging device pixel by pixel from each spectral image frame that has completed bad pixel replacement, and count the fixed stripe component and fixed background component in the row direction and column direction respectively for each spectral image frame that has completed dark current correction. The fixed-pattern noise removal involves subtracting fixed stripe components and fixed background components pixel by pixel from the corresponding spectral band image to obtain an initial cleaned image group. Based on the intrinsic parameter matrix and distortion parameters, lens distortion correction is performed on each spectral band image in the initial cleaned image group to obtain the corrected image for each spectral band. The unified spatial reference mapping is to select the field of view of the corrected image corresponding to the visible light image as a common reference field of view; Based on the extrinsic parameter matrices corresponding to the near-infrared, mid-infrared, and long-infrared images, the near-infrared, mid-infrared, and long-infrared images are transformed to their initial mapping positions under a common reference field of view to obtain a unified field of view image group.
[0009] Optionally, step three specifically includes: The images of each spectral band in the unified field of view image group are divided according to the preset number of grid rows and preset number of grid columns to obtain multiple fixed grids that correspond one-to-one in the position of each spectral band image, and each fixed grid is numbered according to the row number and column number of the grid. For each fixed grid in each spectral band image, extract all pixel values within the grid and calculate the corresponding sharpness value, local gradient value, noise amplitude, local contrast value, local brightness mean, and saturation pixel ratio. For each fixed grid in the mid-wave infrared and long-wave infrared images, the thermal response pixel regions within the fixed grid that are higher than the background response benchmark are extracted, and the thermal area value, thermal boundary density value, and local temperature difference value are calculated. The hot zone area value is the percentage of pixels in the thermal response pixel region; the hot zone boundary density value is the ratio between the number of pixels at the boundary of the thermal response pixel region and the area of the corresponding fixed grid; the local temperature difference value is the difference between the average response value of the thermal response pixel region and the average response value of the corresponding fixed grid background region. Establish corresponding grid quality matrices for visible light images, near-infrared images, mid-wave infrared images, and long-wave infrared images.
[0010] Optionally, the step of establishing corresponding grid quality matrices for visible light images, near-infrared images, mid-wave infrared images, and long-wave infrared images specifically involves: The total number of grids in each spectral band image is determined according to the fixed grid division results, and a single fixed grid is used as a matrix recording unit. Matrix row numbers are generated sequentially according to the order of grid position numbers. In each row of the matrix, the position number of the current fixed grid, the sharpness value, the local gradient value, the noise amplitude, the local contrast value, the local brightness mean, and the saturation pixel ratio are written in the order of the preset fields to generate the grid quality matrix corresponding to the visible light image and the near-infrared image. For mid-wave infrared and long-wave infrared images, the area value of the hot zone, the density value of the hot zone boundary, and the local temperature difference value are written into each matrix row to generate the corresponding mid-wave infrared grid quality matrix and long-wave infrared grid quality matrix. The matrix rows corresponding to all fixed grids within the same spectral band image are combined in order of position number to form the grid quality matrix of that spectral band image. Then, the grid quality matrices corresponding to the visible light image, the near-infrared image, the mid-wave infrared image, and the long-wave infrared image are associated and stored according to the spectral band category to obtain the grid quality matrices stored separately by spectral band.
[0011] Optionally, the improved CrossFormer model is specifically as follows: Input various grid quality matrices into the matrix mapping module, and retrieve the matrix row records of each grid quality matrix according to the grid number; The records of each matrix row are extracted sequentially according to spectral band category and spatial coordinate order to form the corresponding quality record sequence; Perform row vector expansion, position alignment and column vector fixed-length processing on the quality record sequence to generate corresponding quality vector groups, and write each quality vector group into a unified input matrix in the order of spectral segment number and spatial coordinate number to obtain the quality mapping matrix; The quality mapping matrix is input into the hierarchical analysis module. The matrix rows of adjacent positions are divided into multiple basic windows according to the spatial coordinate numbers in the unified input matrix. The matrix rows in each basic window are arranged side by side according to the spectral band category to form a basic window matrix. Multiple adjacent basic window matrices are merged step by step in order from the preset minimum range set to the preset maximum range position set to form a multi-level window matrix group; The window matrix group is input into the cross-fusion module, and intra-layer correlation processing and inter-layer transfer processing are performed on each window matrix in hierarchical order to form a cross-spectral hierarchical fusion matrix. The cross-spectral fusion matrix is input into the tensor generation module. According to the spatial position number in the common reference field of view, all matrix records in the cross-spectral fusion matrix are read in groups, and matrix records with the same spatial position number are grouped into the same position set matrix group. The matrix records in each position aggregation matrix group are arranged vertically according to a preset hierarchical order and horizontally spliced according to a preset spectral segment order to generate a position combination matrix. Perform matrix column filtering on the position combination matrix, which involves removing duplicate and empty columns to generate a position arrangement matrix. Then, expand the position arrangement matrix row by row to obtain a unified fusion feature tensor.
[0012] Optionally, the step of performing intra-layer correlation processing and inter-layer transfer processing on each window matrix in hierarchical order to form a cross-spectral hierarchical fusion matrix is as follows: The intra-layer correlation processing includes: reading the matrix records of each spectral segment in the same window separately, and treating each matrix record as a row vector to be compared; For all row vectors to be compared in the same window, first align each pair of row vectors according to column position, then compare the direction and magnitude of change at corresponding positions column by column, and accumulate the comparison results at all column positions to obtain the inter-row correlation results between each pair of row vectors; Write the inter-row correlation results corresponding to all paired row vectors in the same window into the matrix cell according to the original row position and original column position of the paired row vectors in the window matrix to generate the spectral segment correlation matrix corresponding to the current window. The correlation matrix of the spectral bands is filtered row by row. Matrix cells with a value greater than the preset correlation threshold are retained and the rest of the matrix cells are set to zero to generate the correlation filtering matrix. The association filtering matrix is sorted row by row according to its size to generate an association sequence table. The current matrix record in the original window matrix is retrieved row by row according to the association sequence table. Each matrix record is updated column by column and the updated matrix records are written back to the corresponding window matrix according to their original row positions to obtain the intra-layer fusion matrix. The inter-layer transfer processing includes: matching the window matrices in two adjacent layers according to their spatial inclusion relationship, using the intra-layer fusion matrix in the lower-level window matrix as the local transfer matrix, and using the intra-layer fusion matrix in the higher-level window matrix as the region transfer matrix. Read the position records in the local transfer matrix and write them into the corresponding regional transfer matrix according to their respective regional positions. Then read the regional combination records in the regional transfer matrix and write them back to the corresponding local transfer matrix according to the covered local position range to obtain the inter-layer transfer matrix. The intra-layer fusion matrix and inter-layer transfer matrix of each level are combined and arranged in hierarchical order to form a cross-spectral hierarchical fusion matrix.
[0013] Optionally, step five specifically includes: The unified fused feature tensor is split into tensors according to the row coordinate order, column coordinate order and channel order in the common reference field of view, generating position feature units that correspond one-to-one with each spatial location. Based on the spatial position number corresponding to each position feature unit in the unified fusion feature tensor, the position feature units are written one by one into the reconstruction mesh plane corresponding to the common reference field of view to form the initial reconstruction feature map; The feature units at adjacent spatial locations in the initial reconstructed feature map are subjected to neighborhood expansion processing, and the feature units within a preset neighborhood range around each spatial location are arranged into a local reconstruction matrix according to the spatial adjacency relationship. Column combination processing and row combination processing are performed on each local reconstruction matrix to convert the multi-channel feature records in the local reconstruction matrix into pixel reconstruction records at the corresponding spatial locations; All pixel reconstruction records are written back to the reconstruction mesh plane according to their corresponding spatial positions, and the pixels are arranged in the order of row and column coordinates in the common reference field of view to obtain the enhanced display fusion map.
[0014] A real-time enhanced display control system for multispectral image fusion according to an embodiment of the present invention includes the following modules: The observation synchronization module is used to acquire multispectral raw observation units in the same scene and rearrange the frame numbers according to the time deviation between timestamps to form a multispectral synchronized raw frame group. The correction mapping module is used to perform bad pixel replacement, dark current correction and fixed pattern noise removal on each spectral band image frame in the multispectral synchronous original frame group, and to perform lens distortion correction and unified spatial reference mapping to obtain a unified field of view image group. The matrix generation module is used to divide the unified field-of-view image group into blocks according to a fixed grid, and generate a corresponding grid quality matrix for each spectral band image. The feature fusion module is used to input various mesh quality matrices into the improved CrossFormer model, and then perform fusion processing through the matrix mapping module, hierarchical analysis module, cross fusion module and tensor generation module to generate a unified fused feature tensor. The reconstruction enhancement module is used to perform tensor decomposition on the unified fusion feature tensor, map it to the reconstruction mesh plane corresponding to the common reference field of view, and perform neighborhood expansion processing to obtain the enhanced display fusion map.
[0015] The beneficial effects of this invention are: This invention provides a unified design for the entire processing chain of multispectral images, from raw acquisition, synchronous rearrangement, correction mapping, matrix modeling, feature fusion to reconstruction and display. Compared to existing technologies that rely on direct registration, unified stitching, or whole-image enhancement of multi-source images, this invention offers more comprehensive technical support in terms of input consistency across multiple spectral bands, local quality representation, cross-spectral fusion depth, and ultimately, enhanced display stability. By rearranging the frame numbers of visible light, near-infrared, mid-infrared, and long-infrared images based on time deviations and removing invalid image frames exceeding a preset tolerance range, the multispectral images entering subsequent processing flows maintain temporal consistency, preventing images from different acquisition times from being mixed into the fusion process. By sequentially performing bad pixel replacement, dark current correction, fixed-mode noise removal, lens distortion correction, and unified spatial reference mapping on each spectral band image, image distortion caused by differences in imaging devices, lens imaging deviations, and inconsistent fields of view can be reduced, ensuring that subsequent analysis is based on a unified field of view.
[0016] By dividing a unified field-of-view image group into a fixed grid and generating corresponding grid quality matrices, the local states in multi-spectral images can be expressed in matrix form, giving image information from different spectral bands and locations a unified organizational structure and providing a stable input foundation for subsequent matrix-level processing in the fusion model. By inputting various grid quality matrices into the improved CrossFormer model and processing them sequentially through the matrix mapping module, hierarchical analysis module, cross-fusion module, and tensor generation module, hierarchical analysis and cross-spectral fusion of multi-spectral matrix records can be achieved based on grid position correspondences, forming a unified fusion feature tensor. This avoids the problem of insufficient information utilization caused by simple stitching in existing technologies. By performing tensor decomposition, reconstructed grid plane mapping, and neighborhood expansion on the unified fusion feature tensor, the fusion result can be reconstructed at the pixel level according to a common reference field of view, improving the performance of the enhanced display fusion image in terms of spatial correspondence, local continuity, and display integrity. Therefore, the present invention can effectively solve the technical problems of insufficient synchronization of multispectral images, insufficient connection of correction mapping, insufficient utilization of local quality information, irregular expression of fusion features, and insufficient continuity of enhanced display results in the prior art, thereby obtaining multispectral image fusion results with complete structure, clear processing link and suitable for real-time enhanced display control. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of a real-time enhanced display control method for multispectral image fusion proposed in this invention; Figure 2 This is a schematic diagram of the structure of a real-time enhanced display control system for multispectral image fusion proposed in this invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0019] refer to Figure 1 A real-time enhanced display control method for multispectral image fusion, comprising: Step 1: Collect multispectral raw observation units in the same scene, and rearrange the frame numbers according to the time deviation between timestamps to form a multispectral synchronous raw frame group; Step 2: Perform bad pixel replacement, dark current correction and fixed pattern noise removal on each spectral band image frame in the multispectral synchronous original frame group, and perform lens distortion correction and unified spatial reference mapping to obtain a unified field of view image group. Step 3: Divide the unified field-of-view image group into blocks according to a fixed grid, and generate a corresponding grid quality matrix for each spectral band image; Step 4: Input various mesh quality matrices into the improved CrossFormer model, and perform fusion processing through the matrix mapping module, hierarchical analysis module, cross fusion module and tensor generation module to generate a unified fused feature tensor; Step 5: Perform tensor splitting on the unified fusion feature tensor, map it to the reconstructed mesh plane corresponding to the common reference field of view, and perform neighborhood expansion processing to obtain the enhanced display fusion map.
[0020] This step establishes a complete technical chain from multispectral raw observation acquisition, synchronous rearrangement, image cleaning and correction, grid quality matrix generation, improved CrossFormer fusion processing, to reconstruction and display output. This ensures that multispectral images have a unified temporal basis and spatial reference before entering the fusion stage, reducing information offset issues caused by acquisition time differences, imaging noise, lens distortion, and field-of-view differences between different spectral images. By dividing the unified field-of-view image group into a fixed grid and establishing a corresponding grid quality matrix, a regularized expression of local information in multiple spectral bands is achieved, transforming the subsequent fusion process from traditional coarse processing of the entire image to refined processing focused on local quality features. Through hierarchical analysis and cross-fusion of various grid quality matrices using the improved CrossFormer model, the correlation between different spectral bands at corresponding spatial locations is enhanced, improving the integrity and consistency of the fusion feature organization. By mapping the unified fusion feature tensor to the reconstructed mesh plane corresponding to the common reference field of view, and combining it with neighborhood expansion processing to generate an enhanced display fusion map, the performance of the final display result in terms of spatial continuity, detail preservation and overall display stability can be further improved. This effectively improves the problems of unstable multispectral image fusion results, insufficient utilization of local information and discontinuous enhanced display effects in the existing technology.
[0021] In this embodiment, step one specifically includes: The multispectral raw observation unit under the same scene is collected. The multispectral raw observation unit includes visible light image, near infrared image, mid-wave infrared image, long-wave infrared image, as well as timestamp, intrinsic parameter matrix, extrinsic parameter matrix and distortion parameter corresponding to each spectral band image. The intrinsic parameter matrix is a parameter matrix that characterizes the internal mapping relationship of the imaging coordinate system of the imaging device for the corresponding spectral band. It includes focal length parameters and principal point coordinate parameters, and is used to characterize the correspondence between image pixel coordinates and imaging optical center. The extrinsic parameter matrix is a parameter matrix that characterizes the spatial pose relationship of the imaging device in the corresponding spectral band relative to the common reference coordinate system. It includes rotation parameters and translation parameters and is used to characterize the coordinate transformation relationship between the imaging device in the corresponding spectral band and the common reference coordinate system. The distortion parameters are a set of parameters characterizing the lens imaging deviation of the imaging device for the corresponding spectral band, including radial distortion parameters and tangential distortion parameters, which are used to characterize the offset relationship of the image edge region relative to the ideal imaging position. A unified master clock is selected as the time reference. The frame numbers of each spectral segment image are rearranged according to the time deviation between the timestamp of each spectral segment image and the unified master clock. Spectral segment images with time deviations exceeding the preset tolerance range are judged as invalid frames and removed from the current pairing sequence. The image frames, along with their corresponding timestamps, intrinsic parameter matrices, extrinsic parameter matrices, and distortion parameters, are retained to form a multispectral synchronized original frame group.
[0022] In this embodiment, step two specifically includes: Bad pixel detection is performed on each spectral band image frame in the multispectral synchronous original frame group. The bad pixel detection is to count the abnormal pixel positions where the gray value of the pixel exceeds the preset deviation range of the gray median value of the preset neighborhood size, mark the abnormal pixel positions as bad pixel positions, and replace the bad pixel positions with the median value of the neighborhood of the bad pixel positions. Subtract the dark field reference value of the corresponding spectral imaging device pixel by pixel from each spectral image frame that has completed bad pixel replacement, and count the fixed stripe component and fixed background component in the row direction and column direction respectively for each spectral image frame that has completed dark current correction. The fixed-pattern noise removal involves subtracting fixed stripe components and fixed background components pixel by pixel from the corresponding spectral band image to obtain an initial cleaned image group. Based on the intrinsic parameter matrix and distortion parameters, lens distortion correction is performed on each spectral band image in the initial cleaned image group to obtain the corrected image for each spectral band. The lens distortion correction involves establishing original pixel coordinates for each pixel position in each spectral band image, and converting each original pixel coordinate into a normalized imaging coordinate based on the focal length parameter and principal point coordinate parameter. The radial and tangential distortion parameters of each normalized imaging coordinate point are summed according to a preset weight to obtain the distortion offset corresponding to each normalized imaging coordinate point. Subtract the corresponding distortion offset from each normalized imaging coordinate point to obtain the distortion-free normalized coordinate points; By using the focal length parameter and principal point coordinate parameter, the distortion-reduced normalized coordinate points are inversely calculated to the corrected pixel coordinate positions to obtain the target mapping position; Pixel resampling is performed on the target mapping location to generate corrected images for each spectral band; The unified spatial reference mapping is to select the field of view of the corrected image corresponding to the visible light image as a common reference field of view; Based on the extrinsic parameter matrices corresponding to the near-infrared, mid-infrared, and long-infrared images, the near-infrared, mid-infrared, and long-infrared images are transformed to their initial mapping positions under a common reference field of view to obtain a unified field of view image group.
[0023] This step, by sequentially performing bad pixel detection and replacement, dark current correction, fixed pattern noise removal, lens distortion correction, and unified spatial reference mapping on each spectral band image frame in the multispectral synchronous original frame group, can systematically eliminate local anomalies, background bias, stripe interference, lens imaging deviation, and cross-spectral spatial inconsistency issues in the original imaging data before the multispectral images enter subsequent quality modeling and fusion processing, thereby significantly improving the consistency and usability of the input data. Compared to existing technologies that only perform simple preprocessing on single-spectral-band images or lack a unified correction step before fusion, this invention reduces the damage to local texture and edge structure caused by abnormal pixels through bad pixel location identification and neighborhood median replacement. By subtracting the dark field reference value pixel by pixel and deducting fixed stripe components and fixed background components, it weakens the continuous interference of the imaging device's inherent noise on the image response value, making the response distribution of each spectral-band image more stable. By performing pixel-level lens distortion correction based on intrinsic parameter matrices and distortion parameters, it improves the geometric offset problem between the image edge region and the center region, and improves the accuracy of the imaging position within each spectral-band image. By using the field of view of the visible light image as a common reference field of view and combining it with the extrinsic parameter matrices corresponding to each infrared spectral band for unified spatial reference mapping, it establishes a clear correspondence between near-infrared, mid-wave infrared, and long-wave infrared images under common spatial coordinates, avoiding positional misalignment of different spectral-band images during subsequent segmentation, modeling, and fusion. It can provide a more stable, regular and reliable data foundation for the generation of unified field-of-view image groups, which is conducive to improving the accuracy of subsequent grid quality matrix construction and the spatial consistency and enhancement effect of multispectral fusion display results.
[0024] In this embodiment, step three specifically includes: The images of each spectral band in the unified field of view image group are divided according to the preset number of grid rows and preset number of grid columns to obtain multiple fixed grids that correspond one-to-one in the position of each spectral band image, and each fixed grid is numbered according to the row number and column number of the grid. For each fixed grid in each spectral band image, extract all pixel values within the grid and calculate the corresponding sharpness value, local gradient value, noise amplitude, local contrast value, local brightness mean, and saturation pixel ratio. The method for calculating the sharpness value is as follows: read the response values of all pixels within a fixed grid; The response change between adjacent pixels is counted row by row and column by column, and the response change at each position is accumulated. Pixel regions with response changes greater than a preset threshold are identified as edge change regions. The intensity of changes in all edge change regions is then summarized to obtain the sharpness value of the fixed grid. The more concentrated the edge changes within the fixed grid and the more obvious the abrupt changes in response between adjacent pixels, the greater the corresponding sharpness value. The method for calculating the local gradient value is as follows: taking each pixel in the fixed grid as the center, calculate the response difference between the pixel and its horizontally adjacent pixels, and the response difference between the pixel and its vertically adjacent pixels. The lateral and longitudinal response differences are combined according to position to obtain the local gradient components of each pixel. The local gradient components corresponding to all pixels within a fixed grid are statistically summarized to obtain the local gradient value of the fixed grid. The more pronounced the texture undulations and the steeper the edge transitions within a fixed mesh, the larger the local gradient value. The noise amplitude calculation method is as follows: perform local smoothing processing on all pixels within a fixed grid to obtain the corresponding smoothing response result; The original response value of each pixel within the fixed grid is compared point by point with the smoothed response result at the corresponding position, and the deviation of each pixel is calculated. The deviation of all pixels is statistically analyzed to obtain the overall amplitude of the random fluctuation component in the fixed grid, and the overall amplitude is used as the noise amplitude. The greater the irregular deviation of the pixel response within a fixed grid from the smoothed result, the greater the noise amplitude. The local contrast value is the difference between the maximum and minimum response values within a fixed grid. The average local brightness is the average value of all pixel values within a fixed grid. The saturated pixel ratio is determined by statistically analyzing the proportion of pixels within a fixed grid that reach a preset upper limit response value or a preset lower limit response value to the total number of pixels in that fixed grid. For each fixed grid in the mid-wave infrared and long-wave infrared images, after calculating the sharpness value, local gradient value, noise amplitude, local contrast value, local brightness mean, and saturation pixel ratio, the thermal response pixel regions within the fixed grid that are higher than the background response benchmark are extracted, and the thermal area value, thermal boundary density value, and local temperature difference value are calculated. The hot zone area value is the percentage of pixels in the thermal response pixel region; the hot zone boundary density value is the ratio between the number of pixels at the boundary of the thermal response pixel region and the area of the corresponding fixed grid; the local temperature difference value is the difference between the average response value of the thermal response pixel region and the average response value of the corresponding fixed grid background region. Establish corresponding grid quality matrices for visible light images, near-infrared images, mid-infrared images, and long-infrared images, specifically including: The total number of grids in each spectral band image is determined according to the fixed grid division results, and a single fixed grid is used as a matrix recording unit. Matrix row numbers are generated sequentially according to the order of grid position numbers. In each row of the matrix, the position number of the current fixed grid, the sharpness value, the local gradient value, the noise amplitude, the local contrast value, the local brightness mean, and the saturation pixel ratio are written in the order of the preset fields to generate the grid quality matrix corresponding to the visible light image and the near-infrared image. For mid-wave infrared and long-wave infrared images, the area value of the hot zone, the density value of the hot zone boundary, and the local temperature difference value are written into each matrix row to generate the corresponding mid-wave infrared grid quality matrix and long-wave infrared grid quality matrix. The matrix rows corresponding to all fixed grids within the same spectral band image are combined in order of position number to form the grid quality matrix of that spectral band image. Then, the grid quality matrices corresponding to the visible light image, the near-infrared image, the mid-wave infrared image, and the long-wave infrared image are associated and stored according to the spectral band category to obtain the grid quality matrices stored separately by spectral band.
[0025] This step involves dividing the unified field-of-view image group into a regular grid and extracting sharpness values, local gradient values, noise amplitude, local contrast values, local average brightness, saturation pixel ratio, and the corresponding thermal area, thermal boundary density, and local temperature difference values for each fixed grid. This transforms the local imaging states, originally scattered across different spectral bands, into matrix-based quality information with a one-to-one correspondence and unified field structure. This provides a regular, comparable, and correlated input foundation for subsequent cross-spectral fusion processing. Compared to existing technologies that directly perform unified analysis or coarse stitching of the entire image, this invention synchronously organizes the local sharpness, edge variation, random fluctuation, brightness difference, average brightness level, saturation distribution, and infrared thermal response characteristics of multi-spectral images. This not only improves the ability to identify local quality differences in different grid regions but also enhances the ability to express the correspondence between different spectral bands at the same spatial location.
[0026] By generating matrix row numbers according to fixed grid positions and writing various quality parameters in a preset field order, visible light images, near-infrared images, mid-wave infrared images, and long-wave infrared images can each form a grid quality matrix with a unified structure and clearly defined fields, thereby improving the callability of matrix data and the orderliness of subsequent processing. Especially for mid-wave infrared and long-wave infrared images, further introducing hot zone area values, hot zone boundary density values, and local temperature difference values allows for a more complete characterization of the spatial proportion, boundary distribution, and regional differences of the thermal response region, which is beneficial for enhancing the accuracy of expressing the local state of thermal targets. By storing the grid quality matrices of each spectral band in association according to spectral band category, this invention can also maintain the correspondence between multi-spectral band quality data under a unified management framework, reducing information misalignment problems caused by inconsistent input structures during subsequent fusion. This provides stable data support for matrix mapping, hierarchical analysis, cross-fusion, and unified tensor generation in the improved CrossFormer model, ultimately improving the regularity, local sensitivity, and display quality of multispectral image fusion results.
[0027] In this embodiment, the improved CrossFormer model is specifically as follows: Input various grid quality matrices into the matrix mapping module, and retrieve the matrix row records of each grid quality matrix according to the grid number; The records of each matrix row are extracted sequentially according to spectral band category and spatial coordinate order to form the corresponding quality record sequence; Perform row vector expansion, position alignment and column vector fixed-length processing on the quality record sequence to generate corresponding quality vector groups, and write each quality vector group into a unified input matrix in the order of spectral segment number and spatial coordinate number to obtain the quality mapping matrix; The quality mapping matrix is input into the hierarchical analysis module. The matrix rows of adjacent positions are divided into multiple basic windows according to the spatial coordinate numbers in the unified input matrix. The matrix rows in each basic window are arranged side by side according to the spectral band category to form a basic window matrix. Multiple adjacent basic window matrices are merged step by step in order from the preset minimum range set to the preset maximum range position set to form a multi-level window matrix group; The window matrix group is input into the cross-fusion module. Intra-layer correlation processing and inter-layer transfer processing are performed on each window matrix in hierarchical order to form a cross-spectral hierarchical fusion matrix, specifically including: The intra-layer correlation processing includes: reading the matrix records of each spectral segment in the same window separately, and treating each matrix record as a row vector to be compared; For all row vectors to be compared in the same window, first align each pair of row vectors according to column position, then compare the direction and magnitude of change at corresponding positions column by column, and accumulate the comparison results at all column positions to obtain the inter-row correlation results between each pair of row vectors; Write the inter-row correlation results corresponding to all paired row vectors in the same window into the matrix cell according to the original row position and original column position of the paired row vectors in the window matrix to generate the spectral segment correlation matrix corresponding to the current window. The correlation matrix of the spectral bands is filtered row by row. Matrix cells with a value greater than the preset correlation threshold are retained and the rest of the matrix cells are set to zero to generate the correlation filtering matrix. The association filtering matrix is sorted row by row according to its size to generate an association sequence table. The current matrix record in the original window matrix is retrieved row by row according to the association sequence table. Each matrix record is updated column by column and the updated matrix records are written back to the corresponding window matrix according to their original row positions to obtain the intra-layer fusion matrix. The inter-layer transfer processing includes: matching the window matrices in two adjacent layers according to their spatial inclusion relationship, using the intra-layer fusion matrix in the lower-level window matrix as the local transfer matrix, and using the intra-layer fusion matrix in the higher-level window matrix as the region transfer matrix. Read the position records in the local transfer matrix and write them into the corresponding regional transfer matrix according to their respective regional positions. Then read the regional combination records in the regional transfer matrix and write them back to the corresponding local transfer matrix according to the covered local position range to obtain the inter-layer transfer matrix. The intra-layer fusion matrix and inter-layer transfer matrix of each level are combined and arranged in hierarchical order to form a cross-spectral hierarchical fusion matrix; The cross-spectral fusion matrix is input into the tensor generation module. According to the spatial position number in the common reference field of view, all matrix records in the cross-spectral fusion matrix are read in groups, and matrix records with the same spatial position number are grouped into the same position set matrix group. The matrix records in each position aggregation matrix group are arranged vertically according to a preset hierarchical order and horizontally spliced according to a preset spectral segment order to generate a position combination matrix. Perform matrix column filtering on the position combination matrix, which involves removing duplicate and empty columns to generate a position arrangement matrix. Then, expand the position arrangement matrix row by row to obtain a unified fusion feature tensor.
[0028] The improved CrossFormer model proposed in this step shares similarities with the traditional CrossFormer model in that both employ hierarchical window organization and cross-scale feature interaction as core processing approaches. Neither performs a one-time overall calculation on all input data; instead, it first divides the input object into multiple local windows based on spatial proximity, then gradually expands the analysis scope at different levels to balance local detail representation with global correlation modeling. Both retain the hierarchical processing approach that progresses from small to large windows, establishing relationships between different spatial locations through multi-layered window structures and passing intermediate processing results between levels. This allows the model to reflect both differences within local regions and combinations over a larger scope. Furthermore, both possess an overall processing chain of "mapping—hierarchy—interaction—output," meaning that the input data is first uniformly represented, then interactive analysis is conducted within hierarchical windows, and finally, output results suitable for subsequent tasks are generated. From a structural perspective, the improved model in this step still follows the technical approach of CrossFormer, which emphasizes cross-window, cross-level, and cross-scope modeling. It still has the basic characteristics of hierarchical analysis, step-by-step merging, cross-processing, and unified output. Therefore, in terms of overall framework logic, hierarchical organization, and multi-level interaction ideas, it remains consistent with the traditional CrossFormer model.
[0029] The difference lies in the fact that the improved CrossFormer model proposed in this step does not use conventional image patch features or general feature sequences as input, but instead uses the grid quality matrix corresponding to each spectral band as the direct input object, and constructs a complete processing chain around the organization, alignment, filtering, transfer, and expansion of matrix records. At the input end, the traditional CrossFormer usually focuses on cross-scale representation of the original image patch or basic feature block, while this step first retrieves the matrix row records in various grid quality matrices according to the grid number, then constructs a quality record sequence according to the spectral band category and spatial coordinate order, and obtains the quality mapping matrix through row vector expansion, position alignment, and column vector fixed-lengthization, so that the input basis of the model is transformed from image patch representation to matrix-based quality representation. In hierarchical processing, this step explicitly provides intermediate objects such as the basic window matrix, window matrix group, intra-layer fusion matrix, inter-layer transfer matrix, and cross-spectral hierarchical fusion matrix, and completes dual operations within the window and between adjacent layers through intra-layer correlation processing and inter-layer transfer processing. In particular, it adds special processing procedures for matrix relationship filtering and updating, such as spectral band correlation matrix, correlation filtering matrix, and correlation order list. At the output end, this step does not directly output general fusion features, but further performs position aggregation, vertical arrangement, horizontal splicing, matrix column filtering and row expansion on the cross-spectral level fusion matrix to finally generate a unified fusion feature tensor. Therefore, its core difference is that the traditional model is biased towards conventional cross-scale feature extraction, while the model in this step is biased towards the normalization of the grid quality matrix, multi-level, cross-spectral matrix fusion processing.
[0030] The beneficial effects of the improvements are that by first converting the images of each spectral band into corresponding grid quality matrices before inputting them into the improved CrossFormer model, the model already possesses a data foundation organized according to spatial location and spectral band category before entering the fusion stage. This helps reduce the impact of inconsistent input formats of different spectral bands on the subsequent fusion process. The matrix mapping module performs unified extraction, sequential arrangement, and fixed-length processing of the matrix row records of each grid quality matrix, enabling data from different spectral bands and locations to form stable correspondences within a unified input matrix, facilitating subsequent hierarchical processing. The hierarchical analysis module progressively constructs basic window matrices and multi-level window matrix groups according to spatial location, allowing for the organization and preservation of matrix relationships within local and larger areas. Through intra-layer correlation processing in the cross-fusion module, the degree of correspondence between matrix records of different spectral bands can be identified within the same window. Correlation filtering and column-by-column combination updates improve the utilization of effective matrix information within the window. Inter-layer transfer processing enables bidirectional writing between local and regional records between adjacent layers, allowing matrix information from different analysis scopes to complement each other. By using a tensor generation module to perform position aggregation, concatenation, and column filtering on the cross-spectral hierarchical fusion matrix, a more structurally unified fusion feature tensor can be output. This provides a regularized input for subsequent tensor splitting, reconstruction of the grid plane mapping, and generation of enhanced display fusion graphs. Therefore, this improvement can enhance the organization, hierarchical correlation, and cross-spectral fusion integrity of multi-spectral matrix data.
[0031] In this embodiment, step five specifically includes: The unified fused feature tensor is split into tensors according to the row coordinate order, column coordinate order and channel order in the common reference field of view, generating position feature units that correspond one-to-one with each spatial location. Based on the spatial position number corresponding to each position feature unit in the unified fusion feature tensor, the position feature units are written one by one into the reconstruction mesh plane corresponding to the common reference field of view to form the initial reconstruction feature map; The feature units at adjacent spatial locations in the initial reconstructed feature map are subjected to neighborhood expansion processing, and the feature units within a preset neighborhood range around each spatial location are arranged into a local reconstruction matrix according to the spatial adjacency relationship. Column combination processing and row combination processing are performed on each local reconstruction matrix to convert the multi-channel feature records in the local reconstruction matrix into pixel reconstruction records at the corresponding spatial locations; All pixel reconstruction records are written back to the reconstruction mesh plane according to their corresponding spatial positions, and the pixels are arranged in the order of row and column coordinates in the common reference field of view to obtain the enhanced display fusion map.
[0032] refer to Figure 2A real-time augmented display control system for multispectral image fusion includes the following modules: The observation synchronization module is used to acquire multispectral raw observation units in the same scene and rearrange the frame numbers according to the time deviation between timestamps to form a multispectral synchronized raw frame group. The correction mapping module is used to perform bad pixel replacement, dark current correction and fixed pattern noise removal on each spectral band image frame in the multispectral synchronous original frame group, and to perform lens distortion correction and unified spatial reference mapping to obtain a unified field of view image group. The matrix generation module is used to divide the unified field-of-view image group into blocks according to a fixed grid, and generate a corresponding grid quality matrix for each spectral band image. The feature fusion module is used to input various mesh quality matrices into the improved CrossFormer model, and then perform fusion processing through the matrix mapping module, hierarchical analysis module, cross fusion module and tensor generation module to generate a unified fused feature tensor. The reconstruction enhancement module is used to perform tensor decomposition on the unified fusion feature tensor, map it to the reconstruction mesh plane corresponding to the common reference field of view, and perform neighborhood expansion processing to obtain the enhanced display fusion map.
[0033] Example 1: To verify the feasibility of this invention in practice, it was applied to a hazardous chemical storage and loading / unloading linkage area in a coastal port area. During nighttime and early morning hours, the area consistently suffers from insufficient illumination, interference from damp fog, strong reflectivity of metal facilities, significant thermal disturbance from vehicle exhaust, and frequent cross-operations between personnel and equipment. One side of the area is a container yard, and the other side is a tank inspection channel. In the middle are conveyor corridors, valve platforms, waiting areas for loading / unloading vehicles, and perimeter fencing. Conventional visible light monitoring is prone to losing detail in dark areas after dusk, and close-range searchlights easily cause localized overexposure. While infrared imaging alone can display thermal targets, equipment supports, residual ground heat, vehicle engine compartments, and steam vents all create significant thermal responses, making it difficult for monitoring personnel to quickly distinguish the actual target of interest from background heat sources. Especially when inspectors walk slowly along the outside of the tank area, forklifts travel laterally through the loading and unloading channels, and there is intermittent hot fog spreading under the pipe gallery, traditional fusion display screens often exhibit problems such as blurred target edges, inconsistent flicker enhancement in local areas, slight misalignment of different spectral segments, and significant differences in display effects between the center and the edges of the screen. On-duty personnel need to frequently zoom in, zoom out, and switch channels to complete a full judgment, which increases the observation burden and reduces the efficiency of on-site handling.
[0034] In this embodiment, a multispectral real-time enhanced display control system is deployed on the fixed monitoring tower and inspection vehicle pan-tilt-zoom (PTZ) of the hazardous chemical storage linkage area on the east side of the port area. The system synchronously accesses visible light, near-infrared, mid-wave infrared, and long-wave infrared images, acquiring multispectral raw observation units under the same scene. It compares the timestamps of each spectral band image based on a unified master clock, directly determining and discarding image frames with time deviations exceeding the tolerance range, retaining only image frames that meet the synchronization conditions to form a multispectral synchronized raw frame group. Subsequently, bad pixel replacement, dark current correction, and fixed-mode noise removal are sequentially performed on each spectral band image. Lens distortion correction and unified spatial reference mapping are completed by combining intrinsic and extrinsic parameter matrices and distortion parameters, ensuring that all visible light, near-infrared, mid-wave infrared, and long-wave infrared images are aligned to the same common reference field of view. After unifying the field of view, the system divides the image according to preset grid rows and columns, extracting sharpness, local gradient, noise amplitude, local contrast, local brightness mean, saturation pixel ratio, and hot zone area, hot zone boundary density, and local temperature difference values from the infrared grid for each grid. It then establishes grid quality matrices for visible light, near-infrared, mid-wave infrared, and long-wave infrared. These grid quality matrices are fed into an improved CrossFormer model, where matrix mapping, hierarchical analysis, cross-fusion, and tensor generation processes form a unified fusion feature tensor. This tensor is then mapped onto the reconstructed grid plane corresponding to the common reference field of view, and a neighborhood expansion process is used to output an enhanced fused image. On the monitoring terminal, personnel see a continuous, stable, and detailed fused image that balances thermal response, allowing them to directly observe personnel activity, vehicle movement, thermal anomalies around valve groups, and disturbances at fence edges without manually switching channels repeatedly.
[0035] In this scenario, the problem solved by this invention is quite typical. Although the original system can also perform dual-spectrum superposition, in a humid environment at night, water stains on the container facade and ground will cause large areas of low contrast in the visible light spectrum, while the thermal infrared spectrum will be affected by the residual heat of the vehicle chassis and the heating pipes, forming bright spots. As a result, the outlines of people in traditional images are often wrapped in the thermal background, the boundaries of equipment and fence grids are unclear, and moving targets are more likely to appear at the edges of the image as "bright spots can be seen but the category cannot be identified". After adopting the method of this invention, the front end first removes frames with temporal misalignment, and then uniformly corrects the spatial reference of each spectral band image, avoiding the separation of the human body contour and background structure after fusion. Subsequently, a quality matrix is constructed with fixed grid units, so that local dark areas, low texture areas, strong heat source areas and high reflectivity areas have structured representations. The improved CrossFormer model no longer averages the entire image, but fuses the images based on the correspondence, hierarchical relationship and cross-spectral relationship between each grid matrix. Finally, the enhanced display result retains structural information such as fences, columns and valve box edges, and highlights human bodies, operating equipment and local abnormal heat areas, thus directly serving the observation and judgment of the on-duty personnel.
[0036] To verify the technical effectiveness of this invention, multiple sets of inspection and operation scenario data were continuously collected in the port area, covering typical working conditions such as low-light inspection, vehicle passage, thermal fog disturbance, personnel approaching the perimeter, loading and unloading equipment operation, and high-reflectivity background. Test locations included the outer ring inspection channel of the tank area, the container transition channel, below the valve group platform, the corner of the fence, and the waiting area for loading and unloading vehicles. Each test lasted 120 minutes, with a total collection time of 96 hours, resulting in 480 usable multi-spectral synchronized scene segments. A total of 3120 target events were compared, including 1380 personnel events, 920 vehicle events, 460 equipment thermal anomaly events, and 360 perimeter disturbance events. Using the traditional multi-spectral direct registration weighted display method as a comparison scheme and the method of this invention as the implementation scheme, statistical analysis was conducted on indicators such as effective synchronization frame rate, spatial registration deviation, target discernibility in dark areas, thermal background interference suppression rate, image edge continuity, enhancement of local contrast, display output latency, and overall recognition accuracy. The results are shown in the table below.
[0037] Table 1. Comparison of Multispectral Real-time Enhancement Display Effects under Complex Working Conditions in Port Areas
[0038] As shown in Table 1, the implementation scheme of this invention outperforms traditional methods in terms of input synchronization, spatial alignment, preservation of local details, and final display effect. The proportion of usable synchronized frames increased from 91.4% to 98.7%, indicating that frame number rearrangement based on time deviation can effectively prevent misaligned images from entering the subsequent fusion process. The average spatial registration deviation decreased from 1.84 pixels to 0.46 pixels, indicating that after bad pixel replacement, dark current correction, fixed pattern noise removal, lens distortion correction, and unified spatial reference mapping, images of different spectral bands have higher positional consistency under the common reference field of view. The discernibility rate of personnel outlines in dark areas, the continuous display rate of vehicle edges, and the preservation rate of fence details all increased by more than 18 percentage points, reflecting that the fixed grid quality matrix has a significantly enhanced effect on the organization of local structure and brightness differences, and the improved CrossFormer model can more effectively utilize local quality information to complete cross-spectral fusion. The thermal background interference suppression rate increased from 61.5% to 88.6%, and the false enhancement rate in high reflectivity areas decreased from 18.7% to 5.2%, indicating that this method has achieved a more stable balance between thermal target preservation and background suppression. The average display output latency per frame decreased from 83 milliseconds to 56 milliseconds, the overall event recognition accuracy increased from 79.6% to 94.3%, and the average time for manual review of the terminal was shortened to 2.1 seconds per event. This shows that the present invention not only improves the screen itself, but also effectively reduces the burden of secondary judgment on the screen for the on-duty personnel.
[0039] This embodiment demonstrates that the present invention can solve the problems of inconsistent timing, spatial misalignment, insufficient utilization of local information, and discontinuous enhanced display in traditional technologies in multispectral monitoring scenarios with complex low illumination, thermal disturbance, and high reflectivity in port areas. It significantly improves the continuity, stability, detail preservation, and target discernibility of the enhanced display fusion image, and has good engineering application value.
[0040] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A real-time enhanced display control method for multispectral image fusion, characterized in that, include: Step 1: Collect multispectral raw observation units in the same scene, and rearrange the frame numbers according to the time deviation between timestamps to form a multispectral synchronous raw frame group; Step 2: Perform bad pixel replacement, dark current correction and fixed pattern noise removal on each spectral band image frame in the multispectral synchronous original frame group, and perform lens distortion correction and unified spatial reference mapping to obtain a unified field of view image group. Step 3: Divide the unified field-of-view image group into blocks according to a fixed grid, and generate a corresponding grid quality matrix for each spectral band image; Step 4: Input various mesh quality matrices into the improved CrossFormer model, and perform fusion processing through the matrix mapping module, hierarchical analysis module, cross fusion module and tensor generation module to generate a unified fused feature tensor; Step 5: Perform tensor splitting on the unified fusion feature tensor, map it to the reconstructed mesh plane corresponding to the common reference field of view, and perform neighborhood expansion processing to obtain the enhanced display fusion map.
2. The real-time enhanced display control method for multispectral image fusion according to claim 1, characterized in that, Step one specifically involves: The multispectral raw observation unit under the same scene is collected. The multispectral raw observation unit includes visible light image, near infrared image, mid-wave infrared image, long-wave infrared image, as well as timestamp, intrinsic parameter matrix, extrinsic parameter matrix and distortion parameter corresponding to each spectral band image. A unified master clock is selected as the time reference. The frame numbers of each spectral segment image are rearranged according to the time deviation between the timestamp of each spectral segment image and the unified master clock. The frame number rearrangement involves determining the spectral segment image frames whose time deviation exceeds the preset tolerance range as invalid frames and removing them. The retained image frames of each spectral band, along with their corresponding timestamps, intrinsic parameter matrices, extrinsic parameter matrices, and distortion parameters, form a multispectral synchronous original frame group.
3. The real-time enhanced display control method for multispectral image fusion according to claim 1, characterized in that, Step two specifically involves: Bad pixel detection is performed on each spectral band image frame in the multispectral synchronous original frame group. The bad pixel detection is to count the abnormal pixel positions where the gray value of the pixel exceeds the preset deviation range of the gray median value of the preset neighborhood size, mark the abnormal pixel positions as bad pixel positions, and replace the bad pixel positions with the median value of the neighborhood of the bad pixel positions. Subtract the dark field reference value of the corresponding spectral imaging device pixel by pixel from each spectral image frame that has completed bad pixel replacement, and count the fixed stripe component and fixed background component in the row direction and column direction respectively for each spectral image frame that has completed dark current correction. The fixed-pattern noise removal involves subtracting the fixed stripe component and the fixed background component from the corresponding spectral band image pixel by pixel to obtain an initial cleaned image group. Based on the intrinsic parameter matrix and distortion parameters, lens distortion correction is performed on each spectral band image in the initial cleaned image group to obtain the corrected image for each spectral band. The unified spatial reference mapping is to select the field of view of the corrected image corresponding to the visible light image as a common reference field of view; Based on the extrinsic parameter matrices corresponding to the near-infrared, mid-infrared, and long-infrared images, the near-infrared, mid-infrared, and long-infrared images are transformed to their initial mapping positions under a common reference field of view to obtain a unified field of view image group.
4. The real-time enhanced display control method for multispectral image fusion according to claim 1, characterized in that, Step three specifically involves: The images of each spectral band in the unified field of view image group are divided according to the preset number of grid rows and preset number of grid columns to obtain multiple fixed grids that correspond one-to-one in the position of each spectral band image, and each fixed grid is numbered according to the row number and column number of the grid. For each fixed grid in each spectral band image, extract all pixel values within the grid and calculate the corresponding sharpness value, local gradient value, noise amplitude, local contrast value, local brightness mean, and saturation pixel ratio. For each fixed grid in the mid-wave infrared and long-wave infrared images, the thermal response pixel regions within the fixed grid that are higher than the background response benchmark are extracted, and the thermal area value, thermal boundary density value, and local temperature difference value are calculated. The hot zone area value is the percentage of pixels in the thermal response pixel region; the hot zone boundary density value is the ratio between the number of pixels at the boundary of the thermal response pixel region and the area of the corresponding fixed grid; the local temperature difference value is the difference between the average response value of the thermal response pixel region and the average response value of the corresponding fixed grid background region. Establish corresponding grid quality matrices for visible light images, near-infrared images, mid-wave infrared images, and long-wave infrared images.
5. A real-time enhanced display control method for multispectral image fusion according to claim 4, characterized in that, The establishment of corresponding grid quality matrices for visible light images, near-infrared images, mid-wave infrared images, and long-wave infrared images is specifically as follows: The total number of grids in each spectral band image is determined according to the fixed grid division results, and a single fixed grid is used as a matrix recording unit. Matrix row numbers are generated sequentially according to the order of grid position numbers. In each row of the matrix, the position number of the current fixed grid, the sharpness value, the local gradient value, the noise amplitude, the local contrast value, the local brightness mean, and the saturation pixel ratio are written in the order of the preset fields to generate the grid quality matrix corresponding to the visible light image and the near-infrared image. For mid-wave infrared and long-wave infrared images, the area value of the hot zone, the density value of the hot zone boundary, and the local temperature difference value are written into each matrix row to generate the corresponding mid-wave infrared grid quality matrix and long-wave infrared grid quality matrix. The matrix rows corresponding to all fixed grids within the same spectral band image are combined in order of position number to form the grid quality matrix of the spectral band image; The grid quality matrices corresponding to the visible light image, the near-infrared image, the mid-wave infrared image, and the long-wave infrared image are then associated and stored according to the spectral band category to obtain grid quality matrices stored separately by spectral band.
6. The real-time enhanced display control method for multispectral image fusion according to claim 1, characterized in that, The improved CrossFormer model is specifically as follows: Input various grid quality matrices into the matrix mapping module, and retrieve the matrix row records of each grid quality matrix according to the grid number; The records of each matrix row are extracted sequentially according to spectral band category and spatial coordinate order to form the corresponding quality record sequence; Perform row vector expansion, position alignment and column vector fixed-length processing on the quality record sequence to generate corresponding quality vector groups, and write each quality vector group into a unified input matrix in the order of spectral segment number and spatial coordinate number to obtain the quality mapping matrix; The quality mapping matrix is input into the hierarchical analysis module. The matrix rows of adjacent positions are divided into multiple basic windows according to the spatial coordinate numbers in the unified input matrix. The matrix rows in each basic window are arranged side by side according to the spectral band category to form a basic window matrix. Multiple adjacent basic window matrices are merged step by step in order from the preset minimum range set to the preset maximum range position set to form a multi-level window matrix group; The window matrix group is input into the cross-fusion module, and intra-layer correlation processing and inter-layer transfer processing are performed on each window matrix in hierarchical order to form a cross-spectral hierarchical fusion matrix. The cross-spectral fusion matrix is input into the tensor generation module. According to the spatial position number in the common reference field of view, all matrix records in the cross-spectral fusion matrix are read in groups, and matrix records with the same spatial position number are grouped into the same position set matrix group. The matrix records in each position aggregation matrix group are arranged vertically according to a preset hierarchical order and horizontally spliced according to a preset spectral segment order to generate a position combination matrix. Perform matrix column filtering on the position combination matrix, which involves removing duplicate and empty columns to generate a position arrangement matrix. Then, expand the position arrangement matrix row by row to obtain a unified fusion feature tensor.
7. A real-time enhanced display control method for multispectral image fusion according to claim 6, characterized in that, The process of performing intra-layer correlation processing and inter-layer transfer processing on each window matrix in hierarchical order to form a cross-spectral hierarchical fusion matrix is as follows: The intra-layer correlation processing includes: reading the matrix records of each spectral segment in the same window separately, and treating each matrix record as a row vector to be compared; For all row vectors to be compared in the same window, first align each pair of row vectors according to column position, then compare the direction and magnitude of change at corresponding positions column by column, and accumulate the comparison results at all column positions to obtain the inter-row correlation results between each pair of row vectors; Write the inter-row correlation results corresponding to all paired row vectors in the same window into the matrix cell according to the original row position and original column position of the paired row vectors in the window matrix to generate the spectral segment correlation matrix corresponding to the current window. The correlation matrix of the spectral bands is filtered row by row. Matrix cells with a value greater than the preset correlation threshold are retained and the rest of the matrix cells are set to zero to generate the correlation filtering matrix. The association filtering matrix is sorted row by row according to its size to generate an association sequence table. The current matrix record in the original window matrix is retrieved row by row according to the association sequence table. Each matrix record is updated column by column and the updated matrix records are written back to the corresponding window matrix according to their original row positions to obtain the intra-layer fusion matrix. The inter-layer transfer processing includes: matching the window matrices in two adjacent layers according to their spatial inclusion relationship, using the intra-layer fusion matrix in the lower-level window matrix as the local transfer matrix, and using the intra-layer fusion matrix in the higher-level window matrix as the region transfer matrix. Read the position records in the local transfer matrix and write them into the corresponding regional transfer matrix according to their respective regional positions. Then read the regional combination records in the regional transfer matrix and write them back to the corresponding local transfer matrix according to the covered local position range to obtain the inter-layer transfer matrix. The intra-layer fusion matrix and inter-layer transfer matrix of each level are combined and arranged in hierarchical order to form a cross-spectral hierarchical fusion matrix.
8. A real-time enhanced display control method for multispectral image fusion according to claim 1, characterized in that, Step five specifically involves: The unified fused feature tensor is split into tensors according to the row coordinate order, column coordinate order and channel order in the common reference field of view, generating position feature units that correspond one-to-one with each spatial location. Based on the spatial position number corresponding to each position feature unit in the unified fusion feature tensor, the position feature units are written one by one into the reconstruction mesh plane corresponding to the common reference field of view to form the initial reconstruction feature map; The feature units at adjacent spatial locations in the initial reconstructed feature map are subjected to neighborhood expansion processing, and the feature units within a preset neighborhood range around each spatial location are arranged into a local reconstruction matrix according to the spatial adjacency relationship. Column combination processing and row combination processing are performed on each local reconstruction matrix to convert the multi-channel feature records in the local reconstruction matrix into pixel reconstruction records at the corresponding spatial locations; All pixel reconstruction records are written back to the reconstruction mesh plane according to their corresponding spatial positions, and the pixels are arranged in the order of row and column coordinates in the common reference field of view to obtain the enhanced display fusion map.
9. A real-time enhanced display control system for multispectral image fusion, comprising executing the real-time enhanced display control method for multispectral image fusion as described in any one of claims 1 to 8, characterized in that, Includes the following modules: The observation synchronization module is used to acquire multispectral raw observation units in the same scene and rearrange the frame numbers according to the time deviation between timestamps to form a multispectral synchronized raw frame group. The correction mapping module is used to perform bad pixel replacement, dark current correction and fixed pattern noise removal on each spectral band image frame in the multispectral synchronous original frame group, and to perform lens distortion correction and unified spatial reference mapping to obtain a unified field of view image group. The matrix generation module is used to divide the unified field-of-view image group into blocks according to a fixed grid, and generate a corresponding grid quality matrix for each spectral band image. The feature fusion module is used to input various mesh quality matrices into the improved CrossFormer model, and then perform fusion processing through the matrix mapping module, hierarchical analysis module, cross fusion module and tensor generation module to generate a unified fused feature tensor. The reconstruction enhancement module is used to perform tensor decomposition on the unified fusion feature tensor, map it to the reconstruction mesh plane corresponding to the common reference field of view, and perform neighborhood expansion processing to obtain the enhanced display fusion map.