An unregistered image fusion method based on pseudo-supervised structural consistency
The pseudo-supervised structural consistency method based on Finsler–Laplace–Beltrami spectral analysis solves the problem of structural inconsistency in unregistered image fusion, generates clear fused images, and improves the accuracy and safety of tin smelting process monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KUNMING UNIV OF SCI & TECH
- Filing Date
- 2026-02-26
- Publication Date
- 2026-05-29
AI Technical Summary
Existing image fusion technology, under unregistered conditions, especially in tin smelting production sites, struggles to maintain structural consistency, leading to ghosting of contours and structural streaking, which affects the accuracy of molten pool morphology analysis and furnace lining erosion assessment.
A pseudo-supervised structural consistency method based on Finsler–Laplace–Beltrami spectral analysis is adopted. By constructing structural feature maps, spectral decomposition, and pseudo-supervised structural consistency maps, a structurally consistent fused image is generated, reducing contour artifacts and improving fusion quality.
Under unregistered conditions, the structure remains stable, reducing the risk of ghosting and structural fracture caused by registration errors, generating a clear fused image, and improving the accuracy of monitoring and safety management of the tin smelting process.
Smart Images

Figure CN122115232A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method for fusion of unregistered images based on pseudo-supervised structural consistency. Background Technology
[0002] In existing image fusion technologies, much research focuses on multi-source image fusion under spatial alignment conditions, such as infrared and visible light images, or multispectral and ordinary camera images in industrial inspection scenarios. Common practices rely on registration algorithms to first transform the source images to a unified coordinate system, and then perform fusion at the pixel level, feature level, or deep network feature space. Traditional methods include registration plus fusion based on feature point matching, grayscale registration plus fusion based on mutual information or correlation measurement, and using convolutional neural networks and attention networks for end-to-end learning of the registration field and fusion results. Structural information in many scenarios relies on explicit transformation to solve. Once there are complex occlusions, viewpoint changes, and local deformations in the scene, the registration accuracy decreases, and contour ghosting and structural distortion are prone to occur during the fusion stage.
[0003] For unregistered images, several solutions have emerged in recent years that directly align and fuse in the feature space. A common approach involves regressing pixel-level mapping relationships from the data using networks such as optical flow estimation and deformation field prediction, followed by weighted fusion on the transformed features. Other methods utilize image pyramids, multi-scale block matching, and local affine models to locate approximate displacements at a coarse scale, refine alignment at a fine scale, and then combine Laplacian pyramids, sparse representations, or attention mechanisms to generate a fused image. These methods often still treat geometric registration as an explicit subtask, making them sensitive to noise, weak texture regions, and structurally broken areas. Structural consistency constraints are mostly limited to linear filtering or loss functions, rarely incorporating strict constraints in the operator layer design. The existing structural modeling framework lacks the combination of anisotropic spectral operators based on Finsler–Laplace–Beltrami metric and pseudo-supervised structural consistency maps. It fails to utilize the structural embedding features obtained from spectral decomposition to stably characterize the common structure among unregistered multi-source images. In tin smelting production sites, visible light cameras, infrared thermal imaging equipment, and other monitoring devices are usually arranged around the furnace. Long-term high temperature and dust cause displacement of fixed calibration devices and mechanical structures, making it difficult to maintain strict alignment of acquired images. Existing fusion schemes based on explicit registration are often affected by registration errors in furnace deformation monitoring, furnace charge distribution observation, and cooling water pipe status identification, resulting in unstable structural information.
[0004] In industrial monitoring systems, graph structure and spectral analysis methods have been used for equipment condition modeling and anomaly detection, but these methods are mostly concentrated in signal networks, sensor networks, and simple shape domains. The utilization of image-level structure is mainly based on smoothing or clustering of ordinary graph Laplacian, lacking fine control over anisotropic diffusion direction, structural principal direction, and edge continuity. For unregistered multi-source images, existing methods typically do not combine pixel-level structural features, structural principal direction vectors, and Finsler–Laplace–Beltrami spectral operators to construct a graph model that more closely approximates the structural distribution. Nor do they perform partitioning and adjustment of the initial structural embedding features through a pseudo-supervised structural consistency map under a unified reference coordinate system, and calculate fusion weights accordingly, directly outputting a structurally consistent fused image using a weighted combination method. In the monitoring of tin smelting production processes, there is still considerable room for improvement in obtaining structurally clear and contour-stable fusion results from unregistered multi-source monitoring images without relying on high-precision geometric registration, for use in molten pool morphology analysis, furnace lining erosion assessment, and safety status judgment.
[0005] Therefore, how to provide a method for unregistered image fusion based on pseudo-supervised structural consistency is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose a method for unregistered image fusion based on pseudo-supervised structural consistency. This invention utilizes structural feature extraction, graph structure modeling, and Finsler–Laplace–Beltrami spectral analysis techniques to construct structurally consistent embedded features without the need for geometric registration, and generates structurally consistent fused images based on a fusion weight model. It has the advantages of maintaining structural stability in unaligned scenes, reducing contour artifacts, and improving fusion quality, and can be used for multi-source image fusion in tin smelting process monitoring.
[0007] An unregistered image fusion method based on pseudo-supervised structural consistency according to an embodiment of the present invention includes the following steps:
[0008] Acquire at least two unregistered source images, perform size normalization and noise suppression to obtain preprocessed source images;
[0009] Gradient features, edge features, and texture features are extracted from the preprocessed source images to generate a structural feature map, and a pseudo-supervised structural consistency map is generated based on the structural similarity between the preprocessed source images.
[0010] Under a unified reference coordinate system, pixel blocks in the structural feature map are used as graph nodes and structural similarity is used as edge weights to construct a structural graph. Anisotropic weights based on Finsler–Laplace–Beltrami metric are set according to the structural graph to construct the Finsler–Laplace–Beltrami spectral operator.
[0011] The Finsler–Laplace–Beltrami spectral operator is used to perform spectral decomposition on the structure graph to obtain the initial structure embedding features of each graph node, forming an initial structure embedding feature map;
[0012] The initial structure embedding feature map is adjusted using a pseudo-supervised structure consistency map, and a structure-consistent embedding feature map is obtained under a unified reference coordinate system.
[0013] The fusion weights of the preprocessed source image corresponding to each spatial location are calculated based on the structurally consistent embedding feature map, and a fusion weight map is generated.
[0014] Based on the fusion weight map, the preprocessed source images are weighted and combined in a unified reference coordinate system to obtain a fused image with consistent structure.
[0015] Optionally, the generation of the preprocessed source image specifically includes:
[0016] Acquire at least two unregistered source images from the image acquisition device, register the image number and acquisition time for each unregistered source image, and form an unregistered source image sequence;
[0017] Determine the target image width and target image height, and perform a size adjustment operation on each unregistered source image in the unregistered source image sequence based on the target image width and target image height. The size adjustment operation includes a combination of scaling, cropping and padding. Map the size-adjusted image to the pixel grid corresponding to the target image width and target image height to obtain a size-normalized image sequence.
[0018] Noise suppression is performed on each size-normalized image in the size-normalized image sequence. The noise suppression operation includes a combination of smoothing filtering, edge-preserving filtering and frequency domain filtering. The noise-suppressed image sequence is then output as the preprocessing source image.
[0019] Optionally, the generation of the structural feature map and the pseudo-supervised structural consistency map specifically includes:
[0020] For each preprocessed source image, determine whether the image contains color channels. For preprocessed source images containing color channels, perform color space conversion to obtain a grayscale image. For single-channel preprocessed source images, directly use them as grayscale images. Establish a coordinate system on the grayscale image, with the horizontal axis denoted as x and the vertical axis denoted as y. Record the grayscale value for each coordinate position.
[0021] In a grayscale image, the grayscale difference between adjacent pixels in the horizontal direction and the grayscale difference between adjacent pixels in the vertical direction are calculated according to their coordinate positions. The grayscale difference in the horizontal direction and the grayscale difference in the vertical direction are registered as two components of the gradient feature. At each coordinate position, the gradient magnitude is obtained by taking the square root of the sum of the squares of the two components and registered as the magnitude component of the gradient feature, thus generating a gradient feature map.
[0022] Non-maximum suppression is performed based on the magnitude relationship of the gradient magnitude in the gradient feature map within the neighborhood. The edge and non-edge pixels are marked by a double threshold segmentation rule on the non-maximum suppression result to obtain the edge feature map. The edge feature map and the gradient feature map are aligned according to their coordinate positions in a unified reference coordinate system to form an intermediate structure feature map containing gradient features and edge features.
[0023] A texture filter bank with different scales and orientations is constructed on the grayscale image used for structural analysis. Convolution operation is performed on the grayscale image to obtain multi-scale texture response. The multi-scale texture response is aggregated and registered as a texture feature map in the coordinate dimension. The texture feature map is then concatenated with the intermediate structural feature map according to the coordinate position in a unified reference coordinate system to generate a structural feature map.
[0024] Under a unified reference coordinate system, a similarity index is calculated for the structural feature vectors of each preprocessed source image in the structural feature map at each coordinate position to obtain a structural similarity map. The structural similarity map is then normalized and mapped to obtain the structural consistency confidence score, forming a pseudo-supervised structural consistency map at all coordinate positions.
[0025] Optionally, the construction of the Finsler–Laplace–Beltrami spectral operator specifically includes:
[0026] Under a unified reference coordinate system, block partitioning parameters are set for the structural feature map. The block partitioning parameters include block width and block height. Based on the block partitioning parameters, the structural feature map is divided into several pixel blocks. For each pixel block, the block number, the coordinate value of the block center in the unified reference coordinate system, the statistics of gradient features within the block, the statistics of edge features within the block, and the statistics of texture features within the block are registered. Each pixel block is registered as a graph node, forming a graph node set.
[0027] In the graph node set, the spatial distance between any two graph nodes is calculated based on the block center coordinates. In graph node pairs with a spatial distance less than the preset neighborhood radius, the structural similarity is calculated based on the gradient feature statistics, edge feature statistics, and texture feature statistics within the corresponding pixel block. The structural similarity of each pair of graph nodes is registered as the candidate edge weight, forming a candidate edge set.
[0028] For each candidate edge in the candidate edge set, determine whether to retain the candidate edge based on the relationship between the candidate edge weight and the preset similarity threshold. For candidate edges that meet the retention conditions, register the edge number, starting node number, ending node number, and the direction vector of the line connecting the center of the starting node to the center of the ending node. Form an edge set by the candidate edges that meet the retention conditions, and combine the graph node set and the edge set into a structural graph under a unified reference coordinate system.
[0029] In the structural graph, the structural main direction vector is registered for each graph node based on the gradient feature statistics within the block, and the connection direction vector is registered for each edge in the edge set. The diffusion weight along the structural main direction and the diffusion weight in the direction perpendicular to the structural main direction are calculated based on the angle between the structural main direction vector and the connection direction vector, structural similarity, and spatial distance. The two diffusion weight parameters are registered as anisotropic weights based on the Finsler–Laplace–Beltrami metric, and an index relationship is established between the anisotropic weights and the edge set.
[0030] Based on the graph node set, edge set, and anisotropic weights, a weighted adjacency matrix is constructed for the structural graph. Each row and column of the weighted adjacency matrix corresponds to a graph node, and each matrix element corresponds to the anisotropic weight of an edge. A node degree matrix is constructed based on the sum of the elements in each row of the weighted adjacency matrix. The weighted adjacency matrix and the node degree matrix are combined in Finsler–Laplace–Beltrami discrete form to form a Finsler–Laplace–Beltrami spectral operator matrix. The graph node number is used as the row and column index of the Finsler–Laplace–Beltrami spectral operator matrix to obtain the Finsler–Laplace–Beltrami spectral operator.
[0031] Optionally, the generation of the initial structure embedding feature map specifically includes:
[0032] Read the Finsler–Laplace–Beltrami spectral operator matrix constructed in a unified reference coordinate system, determine the number of rows and columns of the spectral operator matrix, register the number of rows and columns as matrix dimensions, set the target feature dimension according to the matrix dimension, register the target feature dimension as a preset feature dimension parameter, set the number of eigenvalues to be retained, and register the number of retained eigenvalues as a preset eigenvalue quantity parameter, thus forming the eigenvalue decomposition task of the Finsler–Laplace–Beltrami spectral operator;
[0033] Perform eigenvalue decomposition on the Finsler–Laplace–Beltrami spectral operator matrix to obtain a set of eigenvalues and a set of eigenvectors corresponding to each eigenvalue. Assemble the eigenvalues into an eigenvalue sequence and the eigenvectors into an eigenvector sequence. Register an index for each eigenvalue in the eigenvalue sequence and register the corresponding eigenvalue index and graph node number for each eigenvector in the eigenvector sequence. Sort the eigenvalue sequence according to the magnitude of the eigenvalues and adjust the eigenvector sequence according to the sorting result.
[0034] From the sorted feature value sequence, select several non-zero feature values according to the preset feature value quantity parameter. Extract the feature vector corresponding to the selected feature values from the feature vector sequence. For each graph node, read the components at the corresponding positions in the selected feature vector in a fixed order. Combine the components in order to form the graph node-level initial structure embedding feature vector, forming the graph node-level initial structure embedding feature set. Record the index relationship between each graph node and the corresponding initial structure embedding feature vector.
[0035] Under a unified reference coordinate system, based on the correspondence between graph node numbers and pixel blocks in the structural feature map, the graph node-level initial structural embedding feature vector is assigned to the spatial position of the corresponding pixel block. The initial structural embedding feature vector is registered for each pixel block within the coverage area of the structural feature map, generating an initial structural embedding feature map. Normalization is performed on each feature dimension in the initial structural embedding feature map, mapping the values of each feature dimension in all spatial positions to a preset numerical range, resulting in a normalized initial structural embedding feature map.
[0036] Optionally, the generation of the structurally consistent embedding feature map specifically includes:
[0037] Under a unified reference coordinate system, read the pseudo-supervised structural consistency map and the initial structural embedding feature map, register the coordinate index, structural consistency confidence and initial structural embedding feature vector for each spatial location, and establish the mapping relationship between spatial location and structural consistency confidence, as well as the mapping relationship between spatial location and initial structural embedding feature vector;
[0038] Based on the range of structural consistency confidence, a first threshold and a second threshold are set. If the first threshold is greater than the second threshold, spatial locations with structural consistency confidence not less than the first threshold are assigned to the high consistency location set, spatial locations with structural consistency confidence not greater than the second threshold are assigned to the low consistency location set, and spatial locations with structural consistency confidence between the first threshold and the second threshold are assigned to the medium consistency location set. A set identifier is registered for each set.
[0039] Register the first adjustment coefficient for the set of highly consistent locations, the second adjustment coefficient for the set of medium consistent locations, and the third adjustment coefficient for the set of low consistent locations. The first adjustment coefficient is greater than the second adjustment coefficient, and the second adjustment coefficient is greater than the third adjustment coefficient. Write the corresponding adjustment coefficient for each spatial location according to its set to form an adjustment coefficient map.
[0040] Under a unified reference coordinate system, a neighborhood window size and shape are set for each spatial location. Based on the neighborhood window, the set of initial structural embedding feature vectors within the neighborhood is extracted from the initial structural embedding feature map, and the set of adjustment coefficients within the neighborhood is extracted from the adjustment coefficient map. In the high consistency location set, the initial structural embedding feature vectors are updated by weighted summation according to the adjustment coefficients within the neighborhood. In the low consistency location set, the initial structural embedding feature vectors are updated by weighted difference according to the adjustment coefficients within the neighborhood. In the medium consistency location set, the initial structural embedding feature vectors are kept unchanged. The updated structural embedding feature vectors are written into a new feature map according to their spatial location. The new feature map is registered as a structurally consistent embedding feature map under the unified reference coordinate system.
[0041] Optionally, the generation of the fusion weight graph specifically includes:
[0042] Under a unified reference coordinate system, read the structure-consistent embedding feature map and the preprocessed source image. In the structure-consistent embedding feature map, register the coordinate index and structure-consistent embedding feature vector for each spatial location. In each preprocessed source image, register the image number for the entire image and register the pixel value of the corresponding preprocessed source image at each spatial location. Establish the index relationship between spatial location and image number.
[0043] Set the fusion weight generation parameters, which include the input feature dimension, output weight dimension, linear mapping coefficient group, offset coefficient group, non-negativity constraint parameters, and normalization constraint parameters. Assign fusion weight component indexes to each image number according to the output weight dimension, register a set of weight coefficients for each image number in the linear mapping coefficient group, and register an offset for each image number in the offset coefficient group.
[0044] Under a unified reference coordinate system, for each spatial location, the structure-consistent embedding feature vector and fusion weight generation parameters are read. For each image number, the response value is calculated according to a preset linear mapping. The linear mapping operation includes multiplying each dimension of the structure-consistent embedding feature vector by the corresponding weight coefficient and summing them, adding the corresponding offset, comparing the obtained response value with the non-negative constraint parameters, replacing response values less than zero with zero, registering the replaced response value as the initial fusion weight component corresponding to the spatial location and the image number, and combining all initial fusion weight components in the order of image numbers to form the initial fusion weight vector.
[0045] For each spatial location under a unified reference coordinate system, the sum of all components in the initial fusion weight vector is calculated. The sum is compared with the normalization constraint parameter. If the sum is zero, all components are reset to uniform distribution values. If the sum is greater than zero, each component is divided by the sum to obtain a fusion weight vector with a sum of one. The fusion weight vector is then registered for that spatial location.
[0046] Under a unified reference coordinate system, for each spatial location and each image number, the fusion weight components are read from the corresponding fusion weight vector, written into the fusion weight storage structure, and arranged on the two-dimensional index plane formed by the spatial location dimension and the image number dimension to generate a fusion weight map.
[0047] Optionally, the generation of the structurally consistent fused image specifically includes:
[0048] Read the preprocessed source image and the fusion weight map in the unified reference coordinate system, register the image number for each preprocessed source image, register the pixel value of the corresponding preprocessed source image at each spatial location, register the fusion weight value for each spatial location and each image number in the fusion weight map, establish a one-to-one correspondence between pixel value and fusion weight value according to the spatial location index and the image number index, and form a pixel value index table and a fusion weight value index table.
[0049] For each spatial location in the unified reference coordinate system, the pixel values of all preprocessed source images at that spatial location are read in the order of image number according to the pixel value index table. The fusion weight values are read in the same order according to the fusion weight value index table. The pixel value corresponding to each image number is multiplied by the fusion weight value to obtain a weighted pixel component. A set of weighted pixel components is registered for the current spatial location. A summation operation is performed on this set of weighted pixel components in the image number dimension. The summation result is registered as the fusion pixel value of the current spatial location.
[0050] In a unified reference coordinate system, a fusion image storage area is established that corresponds one-to-one with the spatial location of the fusion image. The fusion pixel values are written into the fusion image storage area according to the spatial location index. After the fusion pixel value writing operation is completed in all spatial locations, the fusion image is formed. The mapping relationship between the fusion image and the unified reference coordinate system is registered in the metadata, and the fusion image is registered as a fusion image with consistent structure.
[0051] The beneficial effects of this invention are:
[0052] This invention constructs a structure graph and spectral operator based on the Finsler–Laplace–Beltrami metric in a unified reference coordinate system. Gradient features, edge features, and texture features are jointly modeled as structural features on graph nodes. Initial structural embedding features are obtained through spectral decomposition, and partitioning adjustments are performed using a pseudo-supervised structural consistency graph to generate a structurally consistent embedding feature map. Compared with image fusion methods that rely on explicit registration and deformation field estimation, this invention introduces anisotropic diffusion weights and structural principal direction constraints at the operator level. Even under unregistered conditions, it can still stably characterize structural information such as contours, edges, and textures. This makes the fusion process rely more on geometric and structural relationships rather than precise pixel alignment, thereby reducing the risk of ghosting, streaking, and structural breakage introduced by registration errors.
[0053] This invention, based on a structurally consistent embedded feature map, calculates non-negative and normalized fusion weights for each spatial location and each preprocessed source image through a fusion weight generation model. The fusion weight map is then used to weight and combine the preprocessed source images to output a structurally consistent fused image. This process connects "structural modeling—spectral embedding—weight calculation—weighted combination" into a clear link, resulting in higher stability in contrast, detail preservation, and contour continuity. It also exhibits stronger robustness against noise interference, local viewpoint changes, and slight deformations. In tin smelting process monitoring scenarios, the images acquired by various monitoring devices exhibit viewpoint shifts and installation errors. This invention can generate structurally clear fused images from multiple sources, such as infrared monitoring images and visible light monitoring images, without relying on high-precision geometric registration. This provides a more reliable visual basis for observing molten pool morphology, assessing furnace lining erosion, and determining temperature distribution in key areas, thus helping to improve the monitoring accuracy and safety management level of the tin smelting process. Attached Figure Description
[0054] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0055] Figure 1 This is a flowchart of an unregistered image fusion method based on pseudo-supervised structural consistency proposed in this invention;
[0056] Figure 2 This is a schematic diagram illustrating the Finsler–Laplace–Beltrami spectral operator construction process for an unregistered image fusion method based on pseudo-supervised structural consistency proposed in this invention. Detailed Implementation
[0057] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0058] refer to Figure 1-2 A method for fusion of unregistered images based on pseudo-supervised structural consistency includes the following steps:
[0059] Acquire at least two unregistered source images, perform size normalization and noise suppression to obtain preprocessed source images;
[0060] Gradient features, edge features, and texture features are extracted from the preprocessed source images to generate a structural feature map, and a pseudo-supervised structural consistency map is generated based on the structural similarity between the preprocessed source images.
[0061] Under a unified reference coordinate system, pixel blocks in the structural feature map are used as graph nodes and structural similarity is used as edge weights to construct a structural graph. Anisotropic weights based on Finsler–Laplace–Beltrami metric are set according to the structural graph to construct the Finsler–Laplace–Beltrami spectral operator.
[0062] The Finsler–Laplace–Beltrami spectral operator is used to perform spectral decomposition on the structure graph to obtain the initial structure embedding features of each graph node, forming an initial structure embedding feature map;
[0063] The initial structure embedding feature map is adjusted using a pseudo-supervised structure consistency map, and a structure-consistent embedding feature map is obtained under a unified reference coordinate system.
[0064] The fusion weights of the preprocessed source image corresponding to each spatial location are calculated based on the structurally consistent embedding feature map, and a fusion weight map is generated.
[0065] Based on the fusion weight map, the preprocessed source images are weighted and combined in a unified reference coordinate system to obtain a fused image with consistent structure.
[0066] In this embodiment, the generation of the preprocessed source image specifically includes:
[0067] Acquire at least two unregistered source images from the image acquisition device, register the image number and acquisition time for each unregistered source image, and form an unregistered source image sequence;
[0068] Determine the target image width and target image height, and perform a size adjustment operation on each unregistered source image in the unregistered source image sequence based on the target image width and target image height. The size adjustment operation includes a combination of scaling, cropping and padding. Map the size-adjusted image to the pixel grid corresponding to the target image width and target image height to obtain a size-normalized image sequence.
[0069] Noise suppression is performed on each size-normalized image in the size-normalized image sequence. The noise suppression operation includes a combination of smoothing filtering, edge-preserving filtering and frequency domain filtering. The noise-suppressed image sequence is then output as the preprocessing source image.
[0070] In this embodiment, the generation of the structural feature map and the pseudo-supervised structural consistency map specifically includes:
[0071] For each preprocessed source image, determine whether the image contains color channels. For preprocessed source images containing color channels, perform color space conversion to obtain a grayscale image. For single-channel preprocessed source images, directly use them as grayscale images. Establish a coordinate system on the grayscale image, with the horizontal axis denoted as x and the vertical axis denoted as y. Record the grayscale value for each coordinate position.
[0072] In a grayscale image, the grayscale difference between adjacent pixels in the horizontal direction and the grayscale difference between adjacent pixels in the vertical direction are calculated according to their coordinate positions. The grayscale difference in the horizontal direction and the grayscale difference in the vertical direction are registered as two components of the gradient feature. At each coordinate position, the gradient magnitude is obtained by taking the square root of the sum of the squares of the two components and registered as the magnitude component of the gradient feature, thus generating a gradient feature map.
[0073] Non-maximum suppression is performed based on the magnitude relationship of the gradient magnitude in the gradient feature map within the neighborhood. The edge and non-edge pixels are marked by a double threshold segmentation rule on the non-maximum suppression result to obtain the edge feature map. The edge feature map and the gradient feature map are aligned according to their coordinate positions in a unified reference coordinate system to form an intermediate structure feature map containing gradient features and edge features.
[0074] A texture filter bank with different scales and orientations is constructed on the grayscale image used for structural analysis. Convolution operation is performed on the grayscale image to obtain multi-scale texture response. The multi-scale texture response is aggregated and registered as a texture feature map in the coordinate dimension. The texture feature map is then concatenated with the intermediate structural feature map according to the coordinate position in a unified reference coordinate system to generate a structural feature map.
[0075] Under a unified reference coordinate system, a similarity index is calculated for the structural feature vectors of each preprocessed source image in the structural feature map at each coordinate position to obtain a structural similarity map. The structural similarity map is then normalized and mapped to obtain the structural consistency confidence score, forming a pseudo-supervised structural consistency map at all coordinate positions.
[0076] In this embodiment, the construction of the Finsler–Laplace–Beltrami spectral operator specifically includes:
[0077] Under a unified reference coordinate system, block partitioning parameters are set for the structural feature map. The block partitioning parameters include block width and block height. Based on the block partitioning parameters, the structural feature map is divided into several pixel blocks. For each pixel block, the block number, the coordinate value of the block center in the unified reference coordinate system, the statistics of gradient features within the block, the statistics of edge features within the block, and the statistics of texture features within the block are registered. Each pixel block is registered as a graph node, forming a graph node set.
[0078] In the graph node set, the spatial distance between any two graph nodes is calculated based on the block center coordinates. In graph node pairs with a spatial distance less than the preset neighborhood radius, the structural similarity is calculated based on the gradient feature statistics, edge feature statistics, and texture feature statistics within the corresponding pixel block. The structural similarity of each pair of graph nodes is registered as the candidate edge weight, forming a candidate edge set.
[0079] For each candidate edge in the candidate edge set, determine whether to retain the candidate edge based on the relationship between the candidate edge weight and the preset similarity threshold. For candidate edges that meet the retention conditions, register the edge number, starting node number, ending node number, and the direction vector of the line connecting the center of the starting node to the center of the ending node. Form an edge set by the candidate edges that meet the retention conditions, and combine the graph node set and the edge set into a structural graph under a unified reference coordinate system.
[0080] In the structural graph, the structural main direction vector is registered for each graph node based on the gradient feature statistics within the block, and the connection direction vector is registered for each edge in the edge set. The diffusion weight along the structural main direction and the diffusion weight in the direction perpendicular to the structural main direction are calculated based on the angle between the structural main direction vector and the connection direction vector, structural similarity, and spatial distance. The two diffusion weight parameters are registered as anisotropic weights based on the Finsler–Laplace–Beltrami metric, and an index relationship is established between the anisotropic weights and the edge set.
[0081] Based on the graph node set, edge set, and anisotropic weights, a weighted adjacency matrix is constructed for the structural graph. Each row and column of the weighted adjacency matrix corresponds to a graph node, and each matrix element corresponds to the anisotropic weight of an edge. A node degree matrix is constructed based on the sum of the elements in each row of the weighted adjacency matrix. The weighted adjacency matrix and the node degree matrix are combined in Finsler–Laplace–Beltrami discrete form to form a Finsler–Laplace–Beltrami spectral operator matrix. The graph node number is used as the row and column index of the Finsler–Laplace–Beltrami spectral operator matrix to obtain the Finsler–Laplace–Beltrami spectral operator.
[0082] This invention divides the structural feature map into blocks under a unified reference coordinate system and constructs a graph node set, a candidate edge set, and an edge set. It introduces the structural principal direction vector and the connection direction vector to jointly determine the anisotropic diffusion weights along the structural direction and the lateral direction. Then, it constructs a weighted adjacency matrix and a node degree matrix with the anisotropic weights to form a Finsler–Laplace–Beltrami spectral operator matrix. This enables the graph structure to simultaneously encode spatial neighborhood relations, structural similarity, and principal direction information. The spectral operator generated in the unregistered scene has a stronger ability to characterize edge direction, texture extension, and local structural changes, providing a more stable structural constraint basis for subsequent structural embedding and fusion weight calculation.
[0083] In this embodiment, the generation of the initial structure embedding feature map specifically includes:
[0084] Read the Finsler–Laplace–Beltrami spectral operator matrix constructed in a unified reference coordinate system, determine the number of rows and columns of the spectral operator matrix, register the number of rows and columns as matrix dimensions, set the target feature dimension according to the matrix dimension, register the target feature dimension as a preset feature dimension parameter, set the number of eigenvalues to be retained, and register the number of retained eigenvalues as a preset eigenvalue quantity parameter, thus forming the eigenvalue decomposition task of the Finsler–Laplace–Beltrami spectral operator;
[0085] Perform eigenvalue decomposition on the Finsler–Laplace–Beltrami spectral operator matrix to obtain a set of eigenvalues and a set of eigenvectors corresponding to each eigenvalue. Assemble the eigenvalues into an eigenvalue sequence and the eigenvectors into an eigenvector sequence. Register an index for each eigenvalue in the eigenvalue sequence and register the corresponding eigenvalue index and graph node number for each eigenvector in the eigenvector sequence. Sort the eigenvalue sequence according to the magnitude of the eigenvalues and adjust the eigenvector sequence according to the sorting result.
[0086] From the sorted feature value sequence, select several non-zero feature values according to the preset feature value quantity parameter. Extract the feature vector corresponding to the selected feature values from the feature vector sequence. For each graph node, read the components at the corresponding positions in the selected feature vector in a fixed order. Combine the components in order to form the graph node-level initial structure embedding feature vector, forming the graph node-level initial structure embedding feature set. Record the index relationship between each graph node and the corresponding initial structure embedding feature vector.
[0087] Under a unified reference coordinate system, based on the correspondence between graph node numbers and pixel blocks in the structural feature map, the graph node-level initial structural embedding feature vector is assigned to the spatial position of the corresponding pixel block. The initial structural embedding feature vector is registered for each pixel block within the coverage area of the structural feature map, generating an initial structural embedding feature map. Normalization is performed on each feature dimension in the initial structural embedding feature map, mapping the values of each feature dimension in all spatial positions to a preset numerical range, resulting in a normalized initial structural embedding feature map.
[0088] This invention performs eigenvalue decomposition on the Finsler–Laplace–Beltrami spectral operator matrix, combines eigenvalue sorting and non-zero eigenvalue selection, embeds graph nodes into a low-dimensional initial structure embedding feature vector space, and maps them to pixel blocks in a unified reference coordinate system to generate a normalized initial structure embedding feature map. This elevates structural information from local gradients, edges, and textures to a smooth and measurable spectral embedding representation on the global graph structure, thereby providing a consistent structural alignment benchmark between unregistered images. This provides a stable, compact, and noise-insensitive structural encoding for subsequent adjustment and fusion weight calculation based on pseudo-supervised structural consistency maps.
[0089] In this embodiment, the generation of the structurally consistent embedding feature map specifically includes:
[0090] Under a unified reference coordinate system, read the pseudo-supervised structural consistency map and the initial structural embedding feature map, register the coordinate index, structural consistency confidence and initial structural embedding feature vector for each spatial location, and establish the mapping relationship between spatial location and structural consistency confidence, as well as the mapping relationship between spatial location and initial structural embedding feature vector;
[0091] Based on the range of structural consistency confidence, a first threshold and a second threshold are set. If the first threshold is greater than the second threshold, spatial locations with structural consistency confidence not less than the first threshold are assigned to the high consistency location set, spatial locations with structural consistency confidence not greater than the second threshold are assigned to the low consistency location set, and spatial locations with structural consistency confidence between the first threshold and the second threshold are assigned to the medium consistency location set. A set identifier is registered for each set.
[0092] Register the first adjustment coefficient for the set of highly consistent locations, the second adjustment coefficient for the set of medium consistent locations, and the third adjustment coefficient for the set of low consistent locations. The first adjustment coefficient is greater than the second adjustment coefficient, and the second adjustment coefficient is greater than the third adjustment coefficient. Write the corresponding adjustment coefficient for each spatial location according to its set to form an adjustment coefficient map.
[0093] Under a unified reference coordinate system, a neighborhood window size and shape are set for each spatial location. Based on the neighborhood window, the set of initial structural embedding feature vectors within the neighborhood is extracted from the initial structural embedding feature map, and the set of adjustment coefficients within the neighborhood is extracted from the adjustment coefficient map. In the high consistency location set, the initial structural embedding feature vectors are updated by weighted summation according to the adjustment coefficients within the neighborhood. In the low consistency location set, the initial structural embedding feature vectors are updated by weighted difference according to the adjustment coefficients within the neighborhood. In the medium consistency location set, the initial structural embedding feature vectors are kept unchanged. The updated structural embedding feature vectors are written into a new feature map according to their spatial location. The new feature map is registered as a structurally consistent embedding feature map under the unified reference coordinate system.
[0094] In this embodiment, the generation of the fusion weight graph specifically includes:
[0095] Under a unified reference coordinate system, read the structure-consistent embedding feature map and the preprocessed source image. In the structure-consistent embedding feature map, register the coordinate index and structure-consistent embedding feature vector for each spatial location. In each preprocessed source image, register the image number for the entire image and register the pixel value of the corresponding preprocessed source image at each spatial location. Establish the index relationship between spatial location and image number.
[0096] Set the fusion weight generation parameters, which include the input feature dimension, output weight dimension, linear mapping coefficient group, offset coefficient group, non-negativity constraint parameters, and normalization constraint parameters. Assign fusion weight component indexes to each image number according to the output weight dimension, register a set of weight coefficients for each image number in the linear mapping coefficient group, and register an offset for each image number in the offset coefficient group.
[0097] Under a unified reference coordinate system, for each spatial location, the structure-consistent embedding feature vector and fusion weight generation parameters are read. For each image number, the response value is calculated according to a preset linear mapping. The linear mapping operation includes multiplying each dimension of the structure-consistent embedding feature vector by the corresponding weight coefficient and summing them, adding the corresponding offset, comparing the obtained response value with the non-negative constraint parameters, replacing response values less than zero with zero, registering the replaced response value as the initial fusion weight component corresponding to the spatial location and the image number, and combining all initial fusion weight components in the order of image numbers to form the initial fusion weight vector.
[0098] For each spatial location under a unified reference coordinate system, the sum of all components in the initial fusion weight vector is calculated. The sum is compared with the normalization constraint parameter. If the sum is zero, all components are reset to uniform distribution values. If the sum is greater than zero, each component is divided by the sum to obtain a fusion weight vector with a sum of one. The fusion weight vector is then registered for that spatial location.
[0099] Under a unified reference coordinate system, for each spatial location and each image number, the fusion weight components are read from the corresponding fusion weight vector, written into the fusion weight storage structure, and arranged on the two-dimensional index plane formed by the spatial location dimension and the image number dimension to generate a fusion weight map.
[0100] In this embodiment, the generation of the structurally consistent fused image specifically includes:
[0101] Read the preprocessed source image and the fusion weight map in the unified reference coordinate system, register the image number for each preprocessed source image, register the pixel value of the corresponding preprocessed source image at each spatial location, register the fusion weight value for each spatial location and each image number in the fusion weight map, establish a one-to-one correspondence between pixel value and fusion weight value according to the spatial location index and the image number index, and form a pixel value index table and a fusion weight value index table.
[0102] For each spatial location in the unified reference coordinate system, the pixel values of all preprocessed source images at that spatial location are read in the order of image number according to the pixel value index table. The fusion weight values are read in the same order according to the fusion weight value index table. The pixel value corresponding to each image number is multiplied by the fusion weight value to obtain a weighted pixel component. A set of weighted pixel components is registered for the current spatial location. A summation operation is performed on this set of weighted pixel components in the image number dimension. The summation result is registered as the fusion pixel value of the current spatial location.
[0103] In a unified reference coordinate system, a fusion image storage area is established that corresponds one-to-one with the spatial location of the fusion image. The fusion pixel values are written into the fusion image storage area according to the spatial location index. After the fusion pixel value writing operation is completed in all spatial locations, the fusion image is formed. The mapping relationship between the fusion image and the unified reference coordinate system is registered in the metadata, and the fusion image is registered as a fusion image with consistent structure.
[0104] Example 1:
[0105] To verify the feasibility of this invention in practice, it was applied to a tin smelting furnace monitoring system. The monitored objects included the surface of the molten pool, the furnace opening baffle, the inner wall of the furnace lining, and the cooling water pipe area. A visible light camera and an infrared thermal imager were installed on site. The two devices had significant differences in installation height, pitch angle, and horizontal angle. Moreover, under the influence of high temperature vibration and maintenance adjustments, the support position slowly shifted, and the acquired images were always unregistered in spatial coordinates. In the traditional solution, operators mainly relied on the visible light image to observe the furnace opening and the turbulence of the molten pool, and then used the infrared image to estimate the distribution of high-temperature areas. The traditional solution attempted to introduce geometric registration weighted fusion based on affine transformation to superimpose the infrared intensity onto the visible light image. However, due to significant differences in viewing angle, complex furnace reflection, and insufficient stability of the registration algorithm, obvious ghosting was formed on the free surface of the molten pool and at the junction of the furnace lining and the furnace opening. False high-temperature stripes were generated near the cooling water pipes. The team members needed to frequently switch images and zoom in on local areas for repeated comparison. Problems such as misjudgment of the molten pool boundary, deviation in judging the extent of furnace lining erosion, and delayed alarms for abnormal cooling water pipes still existed.
[0106] The image fusion module of this invention is integrated into the same monitoring system. Visible light source images and infrared source images are directly used as unregistered source images as input. The fusion module first performs size normalization and noise suppression on each frame of source image to obtain a preprocessed source image. Then, gradient features, edge features, and texture features are extracted in grayscale space to construct a structural feature map. A pseudo-supervised structural consistency map is generated through structural feature similarity. Next, the structural feature map is divided into pixel blocks and registered as graph nodes. A structural graph is constructed using structural similarity as edge weights. Anisotropic weights based on Finsler–Laplace–Beltrami metric are introduced to form a Finsler–Laplace–Beltrami spectral operator. After spectral decomposition by the spectral operator, the initial values of each graph node are obtained. The initial structure embedding features are mapped to form an initial structure embedding feature map. Then, the fusion module uses the pseudo-supervised structure consistency map to divide the confidence intervals in spatial location. In the high confidence region, the structure consistency is enhanced by neighborhood weighted averaging. In the low confidence region, unreliable structures are suppressed by weighted difference. In the intermediate confidence region, the original embedding remains unchanged, thus obtaining a structure-consistent embedding feature map. Finally, based on the structure-consistent embedding feature map, non-negative fusion weights are calculated for each spatial location and each preprocessed source image through linear mapping and normalization constraints to form a fusion weight map. The preprocessed source images are then weighted and combined according to the fusion weights to output a structure-consistent fusion image. The fusion result is then integrated into the original monitoring interface for molten pool morphology recognition, furnace lining erosion monitoring, and cooling water pipe anomaly recognition.
[0107] After several smelting cycles, a large number of time slices were extracted from the monitoring system records, generating single visible light images, traditional registration and fusion images, and fusion images of this invention. Three engineers familiar with tin smelting conditions conducted blind evaluations and scoring of typical areas. Combined with meticulously drawn reference edges and infrared high-temperature region masks, the structural sharpness score, ghosting pixel ratio, average number of edge misalignment pixels, and high-temperature region detection rate were calculated. The structural sharpness score used a continuous range from zero to one, with higher values indicating clearer outlines and details. The ghosting pixel ratio was obtained by statistically analyzing the proportion of pixels responding to repeated edges within the molten pool boundary neighborhood. The average number of edge misalignment pixels was calculated using the average distance between the fusion image edge and the reference edge. The high-temperature region detection rate was calculated using the overlap between the reference high-temperature mask and the bright areas in the fusion image. Statistical results are shown in Table 1.
[0108] Table 1 Comparison of Image Quality and Detection Performance
[0109] index Single visible light image Traditional registration and fusion methods Method of the present invention Structural clarity score (0–1) 0.62 0.78 0.89 Ghosting pixel percentage (%) 0 7.4 1.3 Average number of pixels at edge misalignment 5.8 3.9 1.6 Detection rate in high-temperature areas (%) 81.2 88.5 94.7
[0110] As can be seen from the table, the method of this invention significantly outperforms traditional registration and fusion methods in terms of structural clarity score, while also significantly reducing the proportion of ghost pixels and the average number of pixels with edge misalignment. In the free surface area of the molten pool, traditional registration and fusion results often show two or even three misaligned contours, requiring operators to rely on experience to determine which one is real. In contrast, the fusion result of this invention presents a single continuous curve for the molten pool boundary, which remains stable across multiple smelting stages. The detection rate of high-temperature areas is close to 95%, accurately representing the local temperature rise range in the vicinity of cooling water pipes and reducing interference from false high-temperature stripes. When reviewing recorded alarm events, team members reported that after using the fusion result of this invention, the molten pool morphology and temperature distribution can be directly observed in a single image, reducing screen switching and subjective comparison processes. The assessment of furnace lining erosion depth and the judgment of cooling water pipe anomalies are more intuitive, significantly reducing the monitoring burden. This demonstrates that the technical solution of this invention, which achieves structural consistency fusion under unregistered conditions, effectively improves the image monitoring quality of the smelting process and provides support for the safety management of tin smelting production.
[0111] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for fusion of unregistered images based on pseudo-supervised structural consistency, characterized in that, Includes the following steps: Acquire at least two unregistered source images, perform size normalization and noise suppression to obtain preprocessed source images; Gradient features, edge features, and texture features are extracted from the preprocessed source images to generate a structural feature map, and a pseudo-supervised structural consistency map is generated based on the structural similarity between the preprocessed source images. Under a unified reference coordinate system, pixel blocks in the structural feature map are used as graph nodes and structural similarity is used as edge weights to construct a structural graph. Anisotropic weights based on Finsler–Laplace–Beltrami metric are set according to the structural graph to construct the Finsler–Laplace–Beltrami spectral operator. The Finsler–Laplace–Beltrami spectral operator is used to perform spectral decomposition on the structure graph to obtain the initial structure embedding features of each graph node, forming an initial structure embedding feature map; The initial structure embedding feature map is adjusted using a pseudo-supervised structure consistency map, and a structure-consistent embedding feature map is obtained under a unified reference coordinate system. The fusion weights of the preprocessed source image corresponding to each spatial location are calculated based on the structurally consistent embedding feature map, and a fusion weight map is generated. Based on the fusion weight map, the preprocessed source images are weighted and combined in a unified reference coordinate system to obtain a fused image with consistent structure.
2. The unregistered image fusion method based on pseudo-supervised structural consistency according to claim 1, characterized in that, The generation of the preprocessed source image specifically includes: Acquire at least two unregistered source images from the image acquisition device, register the image number and acquisition time for each unregistered source image, and form an unregistered source image sequence; Determine the target image width and target image height, and perform a size adjustment operation on each unregistered source image in the unregistered source image sequence based on the target image width and target image height. The size adjustment operation includes a combination of scaling, cropping and padding. Map the size-adjusted image to the pixel grid corresponding to the target image width and target image height to obtain a size-normalized image sequence. Noise suppression is performed on each size-normalized image in the size-normalized image sequence. The noise suppression operation includes a combination of smoothing filtering, edge-preserving filtering and frequency domain filtering. The noise-suppressed image sequence is then output as the preprocessing source image.
3. The unregistered image fusion method based on pseudo-supervised structural consistency according to claim 1, characterized in that, The generation of the structural feature map and the pseudo-supervised structural consistency map specifically includes: Perform color space conversion and grayscale normalization on each preprocessed source image to obtain an image for structural analysis. In the image used for structural analysis, the horizontal and vertical gray-level differences are calculated according to the pixel position, and the horizontal and vertical gray-level differences are registered as gradient features. Edge detection is performed based on gradient features to obtain an edge feature map. The edge feature map and gradient features are aligned according to a unified reference coordinate system to form an intermediate structure feature map containing gradient features and edge features. A multi-scale texture filter bank is applied to the image used for structural analysis to generate a texture feature map. The texture feature map is then merged with the intermediate structural feature map according to pixel position to generate a structural feature map. Under a unified reference coordinate system, the structural features at each pixel location are similar to calculate the structural similarity map. The structural similarity map is then normalized and mapped to form a pseudo-supervised structural consistency map.
4. The unregistered image fusion method based on pseudo-supervised structural consistency according to claim 1, characterized in that, The construction of the Finsler–Laplace–Beltrami spectral operator specifically includes: Under a unified reference coordinate system, the structural feature map is divided according to the preset block size to obtain several pixel blocks. The block number, block center coordinates and structural feature statistics within the block are registered for each pixel block. Each pixel block is registered as a graph node to form a graph node set. Select graph node pairs whose spatial distance is within a preset neighborhood from the graph node set, calculate the structural similarity based on the structural feature statistics within the corresponding pixel block, register the structural similarity as candidate edge weights, and form a candidate edge set. Valid edges are selected based on the relationship between the structural similarity in the candidate edge set and the preset threshold. For each pair of graph nodes that meet the conditions, the edge number, the starting node number, the ending node number, and the connection direction vector are registered to form an edge set. The graph node set and the edge set are then combined into a structural graph under a unified reference coordinate system. In the structural graph, the main structural direction vector is registered for each graph node, and the connecting direction vector is registered for each edge. The diffusion weight along the structural direction and the lateral diffusion weight are calculated based on the angle between the main structural direction vector and the connecting direction vector, the structural similarity, and the spatial distance. The diffusion weight parameters are registered as anisotropic weights based on the Finsler–Laplace–Beltrami metric. A weighted adjacency matrix and a node degree matrix are constructed based on the graph node set, edge set, and anisotropic weights. The weighted adjacency matrix and the node degree matrix are combined into a discrete Finsler–Laplace–Beltrami spectral operator matrix under a unified reference coordinate system. The graph node number is used as the row index and column index of the Finsler–Laplace–Beltrami spectral operator matrix to obtain the Finsler–Laplace–Beltrami spectral operator.
5. The unregistered image fusion method based on pseudo-supervised structural consistency according to claim 1, characterized in that, The generation of the initial structure embedding feature map specifically includes: Read the Finsler–Laplace–Beltrami spectral operator matrix constructed in a unified reference coordinate system, set the target feature dimension and the number of eigenvalues to be retained according to the number of rows and columns of the spectral operator matrix, and establish the eigenvalue decomposition task of Finsler–Laplace–Beltrami spectral operator; Perform eigenvalue decomposition on the Finsler–Laplace–Beltrami spectral operator matrix to obtain a sequence of eigenvalues and a sequence of eigenvectors corresponding to each eigenvalue. Sort the eigenvalue sequence according to the numerical value and register the corresponding eigenvalue index and the corresponding graph node number for each eigenvector. Select a preset number of non-zero feature values from the sorted feature value sequence, extract the feature vectors corresponding to the selected feature values from the feature vector sequence, and combine the components of each graph node on the selected feature vector in a fixed order to form the graph node-level initial structure embedding feature set. Under a unified reference coordinate system, based on the correspondence between graph node numbers and pixel blocks in the structural feature map, the initial structural embedding features at the graph node level are assigned to the corresponding pixel blocks. The initial structural embedding features are then filled in the spatial locations covered by the structural feature map. Normalization is performed on each feature dimension to generate the initial structural embedding feature map.
6. The unregistered image fusion method based on pseudo-supervised structural consistency according to claim 1, characterized in that, The generation of the structure-consistent embedding feature map specifically includes: Under a unified reference coordinate system, read the pseudo-supervised structural consistency map and the initial structural embedding feature map, register the structural consistency confidence and the initial structural embedding feature for each spatial location, and establish a spatial location index; Based on the structural consistency confidence level, a first threshold and a second threshold are set. If the first threshold is greater than the second threshold, the spatial locations with a structural consistency confidence level not less than the first threshold are divided into a high consistency location set, the spatial locations with a structural consistency confidence level not greater than the second threshold are divided into a low consistency location set, and the remaining spatial locations are divided into a medium consistency location set. The first adjustment coefficient is registered in the set of highly consistent positions, the second adjustment coefficient is registered in the set of medium consistent positions, and the third adjustment coefficient is registered in the set of low consistent positions. The first adjustment coefficient is greater than the second adjustment coefficient and the third adjustment coefficient. Under a unified reference coordinate system, a neighborhood window is determined for each spatial location. Based on the initial structural embedding features and corresponding adjustment coefficients within the neighborhood window, the initial structural embedding features are updated using a weighted average method in the high consistency location set, and using a weighted difference method in the low consistency location set. The initial structural embedding features are kept unchanged in the medium consistency location set. The updated structural embedding features are then registered as structurally consistent embedding feature maps according to their spatial locations.
7. The unregistered image fusion method based on pseudo-supervised structural consistency according to claim 1, characterized in that, The generation of the fusion weight graph specifically includes: Under a unified reference coordinate system, read the structure-consistent embedding feature map and the preprocessed source image, register the coordinate index and structure-consistent embedding feature vector for each spatial location, register the image number and the pixel value at each spatial location for each preprocessed source image, and establish the index relationship between spatial location and image number. Set the fusion weight generation parameters, which include the fusion weight dimension, non-negativity constraint parameters, and normalization constraint parameters. Assign fusion weight component indices to each image number according to the fusion weight dimension. Under a unified reference coordinate system, for each spatial location, the initial fusion weight component of each image number is calculated based on the corresponding structurally consistent embedded feature vector and fusion weight generation parameters. All initial fusion weight components are combined into an initial fusion weight vector according to the image number order. The initial fusion weight vector is normalized according to the non-negativity constraint parameter and the normalization constraint parameter, so that all components in the initial fusion weight vector are non-negative and the sum of all image number dimensions is one. The normalization result is registered as the fusion weight vector. Under a unified reference coordinate system, the fusion weight vector is written into the fusion weight storage structure, and the fusion weight value is registered for each spatial location and each image number to generate a fusion weight map.
8. The unregistered image fusion method based on pseudo-supervised structural consistency according to claim 1, characterized in that, The generation of the structurally consistent fused image specifically includes: Read the preprocessed source image and the fusion weight map in a unified reference coordinate system, register the image number and the pixel value of each spatial location for each preprocessed source image, register the fusion weight value for each spatial location and each image number, and establish the index relationship between pixel value and fusion weight value. For each spatial location, the pixel value and fusion weight value of the corresponding preprocessed source image are read in the order of image number. The pixel value corresponding to each image number is multiplied by the fusion weight value to obtain the weighted pixel component. All weighted pixel components are summed in the image number dimension to obtain the fusion pixel value of the corresponding spatial location. The fused pixel values are written to the fused image storage area based on the spatial location index. The fused image storage area is then filled with the fused pixel values of all spatial locations to generate the fused image.