Image processing method and system based on hierarchical fusion of multi-source sensor data

By performing geometric registration, spectral normalization, and feature extraction on multi-source sensor data, calculating cross-modal attention weights, and optimizing consistency constraints, the problem of poor cross-modal feature association modeling in multi-modal image fusion is solved, achieving high-quality multi-source data fusion and interpretable output.

CN121883249APending Publication Date: 2026-04-17STATE POWER INVESTMENT GRP DAMAOQI NEW ENERGY POWER GENERATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing multimodal image fusion techniques suffer from poor cross-modal feature correlation modeling, low fusion weight adaptability, and insufficient optimization of spectral and semantic consistency, making it difficult to achieve hierarchical fusion and interpretable output of multi-source data under a unified network framework.

Method used

By performing geometric registration and spectral normalization on multi-source sensor data, hierarchical features are extracted and cross-modal attention weights are calculated. Consistency constraint optimization and feature reconstruction are performed on the fused features to generate a fused image and output interpretable saliency results.

Benefits of technology

It achieves consistent processing of multi-source data in geometric and spectral dimensions, ensuring the complementary use of modalities and generating fused images with simultaneous improvements in spatial structure, spectral restoration and semantic consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883249A_ABST
    Figure CN121883249A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and system based on multi-source sensor data hierarchical fusion, and relates to the technical field of multi-source sensor data fusion and image processing, and the method comprises the steps: carrying out the geometric registration and spectrum normalization processing of the data collected by a multi-source sensor; extracting hierarchical features from the standardized multi-modal data, and calculating a cross-modal attention weight; and performing consistency constraint optimization and feature reconstruction processing on the fusion features, generating a fusion image and outputting an interpretable significance result. According to the method, through geometric registration and spectrum normalization processing, unified standardization of multi-source sensor data in space and spectrum levels is realized; constructing a multi-scale semantic association and adaptive information fusion mechanism through hierarchical feature extraction and cross-modal attention weight calculation; and the fusion features are coordinated and balanced in structure, spectrum and semantic levels, so that a high-fidelity, semantic-consistent and interpretable fusion image is obtained, and the precision and stability of multi-source data fusion are integrally improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-source sensor data fusion and image processing technology, specifically to an image processing method and system based on hierarchical fusion of multi-source sensor data. Background Technology

[0002] With the rapid development of multi-source sensor technology and intelligent imaging systems, data from various sensors such as optics, infrared, radar, and lidar are widely used in image processing and scene understanding. Multi-source data fusion technology has become an important means for high-precision environmental perception, remote sensing identification, and intelligent monitoring. Existing multimodal fusion methods mostly employ convolutional neural networks, attention mechanisms, or feature encoding frameworks to achieve information complementarity, thereby improving imaging quality under complex lighting, occlusion, and noise conditions. However, due to the differences in spatial resolution, spectral response, and signal characteristics among different sensors, how to achieve hierarchical correlation and effective fusion of cross-modal features remains a research hotspot and technical challenge in this field.

[0003] Existing multi-source sensor image fusion techniques mainly perform feature stitching or weighted fusion in single-layer or shallow feature spaces, lacking unified modeling of semantic relationships between multi-scale and hierarchical features. This results in significant deficiencies in maintaining spatial structure and spectral consistency of the fusion results. Traditional geometric registration methods have limited accuracy under conditions of multi-sensor nonlinear distortion or complex viewpoint differences, and spectral normalization processing is also difficult to achieve adaptive matching of response functions in dynamic scenes. Existing deep learning fusion models often employ fixed weights or static attention mechanisms in their fusion strategies, failing to adaptively adjust according to differences in multimodal features, leading to redundancy or loss of modal information. Current fusion optimization stages often rely on a single loss function, lacking a collaborative optimization mechanism for structural, spectral, and semantic constraints, resulting in deviations in detail restoration and semantic consistency of the fused image. In summary, existing technologies cannot simultaneously achieve accurate registration, hierarchical feature fusion, and consistency optimization of multi-source data within a unified framework, making it difficult to generate fusion results with synergy at the spatial, spectral, and semantic levels. Therefore, it is necessary to propose an image processing method based on hierarchical fusion of multi-source sensor data to achieve adaptive fusion and interpretable output of multimodal features. Summary of the Invention

[0004] In view of the above-mentioned problems, the present invention is proposed.

[0005] Therefore, the technical problem solved by this invention is that existing multimodal image fusion technologies suffer from poor cross-modal feature association modeling, low adaptability of fusion weights, insufficient optimization of spectral and semantic consistency, and the problem of how to achieve hierarchical fusion and interpretable output of multi-source data under a unified network framework.

[0006] To address the aforementioned technical problems, this invention provides the following technical solution: an image processing method based on hierarchical fusion of multi-source sensor data, comprising geometric registration and spectral normalization processing of multi-source sensor acquired data; extracting hierarchical features from the normalized multimodal data and calculating cross-modal attention weights; performing consistency constraint optimization and feature reconstruction processing on the fused features to generate a fused image and output interpretability saliency results.

[0007] As a preferred embodiment of the image processing method based on hierarchical fusion of multi-source sensor data described in this invention, the geometric registration includes: determining corresponding point pairs by calculating key feature points between different sensor images or point clouds, and performing spatial alignment between multiple modalities by establishing a spatial mapping relationship; performing feature extraction and descriptor calculation on each sensor image, and establishing matching pairs by judging feature similarity; if the matching error exceeds the preset error threshold range, performing affine transformation iterative correction on the local region until all modal data are aligned in a unified spatial reference coordinate system; during the registration process, dynamically adjusting the preset error threshold according to the residual value after each iteration, and outputting the final spatial correction result when the residual converges to the allowable error range.

[0008] As a preferred embodiment of the image processing method based on hierarchical fusion of multi-source sensor data described in this invention, the spectral normalization processing includes: calculating spectral response characteristics and establishing a unified brightness and band mapping model for the registered multimodal data; analyzing the image brightness distribution and color histogram of each modality, and performing linear or nonlinear mapping on the data according to a preset standard response curve to ensure that the grayscale distribution of each modality is within a similar range; when the spectral brightness difference between different modalities is detected to exceed a set spectral brightness difference threshold, a dynamic correction unit is automatically triggered to recalculate the spectral mapping parameters and update the conversion coefficients; based on comparative analysis of the average brightness difference of each modality, the output is continuously adjusted through an adaptive correction function to ensure that the final spectral characteristics remain consistent in the multimodal fusion space.

[0009] As a preferred embodiment of the image processing method based on hierarchical fusion of multi-source sensor data described in this invention, the extraction of hierarchical features from standardized multimodal data includes: extracting local texture features, structural features, and high-level semantic information of images at different spatial resolutions through a feature extraction network composed of multiple hierarchical structures; extracting feature representations at different scales step by step through hierarchical encoding, with the bottom layer capturing edge and detail information, the middle layer extracting shape and region contours, and the high layer extracting global context information through a deep attention mechanism; during feature extraction, the network automatically adjusts the convolutional kernel size and the number of channels to encode input features at different scales step by step; when gradient instability or feature extraction overfitting is detected, a residual connection mechanism is activated to jointly input low-level and high-level information; the encoding network shares some parameters among different modalities to maintain consistent feature distribution and performs local adaptive adjustments based on modal differences.

[0010] As a preferred embodiment of the image processing method based on hierarchical fusion of multi-source sensor data described in this invention, the calculation of cross-modal attention weights includes: automatically assigning contribution weights of each modality in the fusion process based on the similarity of features between different modalities; calculating the correlation between different modalities by mapping the feature representations of each modality to a unified feature space, and weighting the features according to the correlation level; during the weight allocation process, if a large modal difference is detected, performing local feature enhancement or suppression operations, and dynamically adjusting the weight of the current modality through an attention mechanism.

[0011] As a preferred embodiment of the image processing method based on hierarchical fusion of multi-source sensor data described in this invention, the following steps are included: The consistency constraint optimization and feature reconstruction processing of the fused features comprises establishing three optimization mechanisms simultaneously in the fused feature space: structural consistency constraint, spectral consistency constraint, and semantic consistency constraint. Structural consistency constraint corrects the geometric structure by calculating the difference between the fused result and the reference modality in edge and texture gradients, ensuring that the fused image retains the contour information of the reference modality. Spectral consistency constraint confirms the physical authenticity of the brightness and color of the fused output by comparing the weighted average difference between the fused features and the features of each modality. Semantic consistency constraint confirms high-level semantic consistency by comparing the difference between the predicted semantic distribution of the fused features and the true semantic label distribution. During the optimization process, the weight coefficients of the three constraints are dynamically adjusted. When an increase in structural error is detected, the structural consistency constraint is strengthened first. When the spectral deviation exceeds the spectral consistency deviation threshold, the weight of the spectral consistency constraint is automatically increased. When the semantic distribution change of the fused result exceeds the semantic distribution offset threshold, the weight of the semantic consistency constraint is strengthened. The model parameters are updated through continuous iteration and backpropagation to optimize the structural consistency, spectral consistency, and semantic consistency constraint terms, thereby performing joint optimization processing on the fused features.

[0012] As a preferred embodiment of the image processing method based on multi-source sensor data hierarchical fusion described in this invention, the step of generating a fused image and outputting interpretability saliency results includes: performing layer-by-layer upsampling on multi-layer fusion features to generate a fused image, generating a saliency distribution map to indicate the contribution area of ​​each modality feature; in the image reconstruction stage, restoring the original spatial resolution through deconvolution operation, and retaining early feature information in each layer using skip connections; when the output image pixel value is detected to exceed the preset output pixel amplitude threshold range, automatically triggering a normalization unit to adjust the pixel data amplitude; in the interpretability analysis stage, calculating the saliency distribution map according to the degree of influence of different modality features on the output category, and visually displaying the contribution area of ​​each modality feature during the fusion process.

[0013] Another objective of this invention is to provide an image processing system based on hierarchical fusion of multi-source sensor data. This system can generate fused images and output interpretable saliency results by performing consistency constraint optimization and feature reconstruction processing on fused features. This solves the problem that current multi-source sensor fusion technologies lack structural, spectral, and semantic co-constraints in their fused features.

[0014] As a preferred embodiment of the image processing system based on hierarchical fusion of multi-source sensor data described in this invention, it includes: a sensor calibration module, a feature fusion module, and a reconstruction optimization module; the sensor calibration module is used to perform geometric registration and spectral normalization processing on the data acquired by the multi-source sensors; the feature fusion module is used to extract hierarchical features from the normalized multimodal data and calculate cross-modal attention weights; the reconstruction optimization module is used to perform consistency constraint optimization and feature reconstruction processing on the fused features, generate a fused image, and output interpretability and significance results.

[0015] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement an image processing method based on hierarchical fusion of multi-source sensor data.

[0016] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of an image processing method based on hierarchical fusion of multi-source sensor data.

[0017] The beneficial effects of this invention are as follows: The image processing method based on hierarchical fusion of multi-source sensor data provided by this invention performs geometric registration and spectral normalization on multi-source sensor data, achieving consistent processing of multi-source data in both geometric and spectral dimensions, providing high-quality, low-bias input data for subsequent feature extraction; hierarchical features are extracted from the standardized multimodal data, and cross-modal attention weights are calculated, ensuring that the complementarity between modalities is fully utilized and effectively avoiding excessive dominance or redundant interference of a certain modality feature; consistency constraint optimization and feature reconstruction processing are performed on the fused features to generate a fused image and output interpretable saliency results, realizing the collaborative optimization and reconstruction of fused features at multiple scales and multiple semantic levels, so that the output image is simultaneously improved in terms of spatial structure, spectral restoration, and semantic consistency. This invention achieves better results in geometric and spectral consistency processing of multi-source data, hierarchical fusion modeling of cross-modal features, and multidimensional consistency optimization and high-fidelity reconstruction of fusion results. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is an overall flowchart of an image processing method based on hierarchical fusion of multi-source sensor data provided in Embodiment 1 of the present invention. Detailed Implementation

[0020] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0021] Example 1, referring to Figure 1 As an embodiment of the present invention, an image processing method based on hierarchical fusion of multi-source sensor data is provided, comprising:

[0022] S1: Perform geometric registration and spectral normalization on data acquired from multiple sensors.

[0023] Furthermore, geometric registration includes: determining corresponding point pairs by calculating key feature points between different sensor images or point clouds, and performing spatial alignment between multiple modalities by establishing spatial mapping relationships; extracting features and calculating descriptors for each sensor image, and establishing matching pairs by judging feature similarity; if the matching error exceeds the preset error threshold range, performing affine transformation iterative correction of the local region until all modal data are aligned in a unified spatial reference coordinate system; during the registration process, dynamically adjusting the preset error threshold based on the residual value after each iteration, and outputting the final spatial correction result when the residual converges to the allowable error range.

[0024] It should be noted that a preferred scheme for the preset error threshold range specifically includes the following: when the matching error value is below 1.0 pixel, it indicates that the image has reached an acceptable level in terms of geometric alignment; when the matching error value is between 1.0 and 3.0 pixels, it indicates that there is a local offset in the registration, and local affine adjustment is performed according to the error gradient direction; when the matching error exceeds 3.0 pixels, it is considered a serious mismatch, and the spatial mapping matrix is ​​re-established; after each round of correction, the residual descent rate is calculated to determine whether it has entered the stable stage, and when the global residual converges to below 0.5 pixels, the final registration result is output.

[0025] It should be noted that the spectral normalization process includes: calculating the spectral response characteristics and establishing a unified brightness and band mapping model for the registered multimodal data; analyzing the image brightness distribution and color histogram of each mode, and performing linear or nonlinear mapping on the data according to the preset standard response curve to make the grayscale distribution of each mode within a similar range; when the spectral brightness difference between different modes exceeds the set spectral brightness difference threshold, the dynamic correction unit is automatically triggered to recalculate the spectral mapping parameters and update the conversion coefficients; based on the comparative analysis of the average brightness difference of each mode, the output is continuously adjusted through an adaptive correction function to ensure that the final spectral characteristics remain consistent in the multimodal fusion space.

[0026] It should also be noted that a preferred scheme for spectral normalization processing specifically includes: calculating the difference in the mean brightness of different modal images in a unified grayscale space to obtain the spectral brightness difference, and setting a threshold for the spectral brightness difference. When the spectral brightness difference value is less than 0.03, it indicates that the lighting conditions are consistent with the photosensitivity response; when the spectral brightness difference value is between 0.03 and 0.05, it indicates that there is a slight difference in local reflectivity, which is compensated by proportional adjustment; when the spectral brightness difference value exceeds 0.05, it indicates that there is a significant inconsistency in the sensor response, and the spectral mapping model is refitted; when the difference still exceeds 0.10, it is considered that the physical property differences between modal responses are large, and global band remapping is required.

[0027] It should also be noted that by performing geometric registration and spectral normalization on data acquired from multiple sensors, spatial and spectral consistency of cross-modal data is achieved. Geometric registration extracts key feature points and establishes matching pairs, using residual iteration and threshold self-adjustment mechanisms to complete spatial alignment between multiple modes, effectively reducing geometric mismatch caused by differences in sensor viewing angles. Spectral normalization dynamically corrects the grayscale and color distribution of each mode through a brightness and band mapping model. When a spectral difference is detected to exceed a threshold, adaptive correction is automatically triggered, and mapping parameters are updated in real time to eliminate illumination and response deviations. This achieves unified standardization of input data in geometric position and spectral response, ensuring that subsequent feature extraction is performed under the same reference system, thereby reducing the impact of data offset on fusion accuracy and improving the stability and reliability of the overall fusion system.

[0028] S2: Extract hierarchical features from the standardized multimodal data and calculate cross-modal attention weights.

[0029] Furthermore, the extraction of hierarchical features from standardized multimodal data includes: extracting local texture features, structural features, and high-level semantic information of images at different spatial resolutions through a feature extraction network composed of multiple hierarchical structures; extracting feature representations at different scales step by step through hierarchical encoding, with the bottom layer capturing edge and detail information, the middle layer extracting shape and region contours, and the top layer extracting global context information through a deep attention mechanism; during feature extraction, the network automatically adjusts the convolutional kernel size and the number of channels to encode input features at different scales step by step; when gradient instability or feature extraction overfitting is detected, a residual connection mechanism is activated to jointly input low-level and high-level information; the encoding network shares some parameters among different modalities to maintain consistent feature distribution and performs local adaptive adjustments based on modal differences.

[0030] It should be noted that calculating cross-modal attention weights includes: automatically assigning contribution weights of each modality to the fusion process based on the similarity of features between different modalities; calculating the correlation between different modalities by mapping the feature representations of each modality to a unified feature space, and weighting the features according to the correlation level; and if a large modal difference is detected during the weight allocation process, performing local feature enhancement or suppression operations and dynamically adjusting the weight of the current modality through the attention mechanism.

[0031] It should also be noted that a preferred scheme for calculating cross-modal attention weights specifically includes: standardizing the feature representations from different sensors to a unified dimension; converting the original features of each modality into an alignable feature form through a shared feature mapping structure, eliminating differences in scale and distribution between modalities; calculating the similarity relationship between any two modal features sequentially within a unified feature space, the calculation of the similarity relationship being based on element-wise feature response comparison, and quantifying the degree of correlation between different modal features in local regions by normalizing the differences in corresponding feature values ​​at spatial locations; weighting and aggregating the feature similarity relationships of each modality within a local range to form a cross-modal correlation distribution, and assigning corresponding attention weights to each modal feature based on this distribution, the attention weights reflecting the modality's importance in fusion. The dynamic contribution ratio during the fusion process; during the weight allocation process, when a significant difference is detected between the feature distribution of a certain modality and other modalities, an adaptive weight adjustment step is triggered. In this step, the feature response of the current modality is locally enhanced or suppressed, so that it maintains the feature strength in the high-relevance region and weakens the response in the low-relevance region, thereby maintaining the weight balance of different modal features; throughout the feature fusion stage, the distribution of attention weights is continuously updated through iterative calculation. The update process is constrained by the local consistency of the fused features. When the similarity distribution of the fused features tends to stabilize, the final weight matrix is ​​output; finally, the weight matrix is ​​used to weight and integrate the features of each modality to generate cross-modal fusion features, so that the effective information of different modalities forms a continuous and distinguishable fusion representation in the spatial hierarchy.

[0032] It should also be noted that by performing hierarchical feature extraction on standardized data and calculating cross-modal attention weights, multi-scale semantic information extraction and dynamic modeling of modal relevance are achieved. A multi-layer coding structure is used to extract local texture, shape, and global semantic features to form a hierarchical feature representation. Subsequently, the local similarity between each modality is calculated through a unified feature space, and fusion weights are automatically assigned according to the similarity level. When a modality is detected to be significantly different from other modalities, a local enhancement or suppression mechanism is triggered to adjust its weight. This dynamic weight allocation process makes the modal contribution adaptively change, retaining high-confidence modal features while reducing redundant information. Through this mechanism, balanced fusion and relevance enhancement of cross-modal features are achieved, effectively improving the integrity and discriminativeness of the fused feature expression and avoiding information bias caused by the dominance of a single modality.

[0033] S3: Perform consistency constraint optimization and feature reconstruction on the fused features to generate a fused image and output interpretable saliency results.

[0034] Furthermore, the consistency constraint optimization and feature reconstruction processing of the fused features includes simultaneously establishing three optimization mechanisms—structural consistency constraint, spectral consistency constraint, and semantic consistency constraint—within the fused feature space. Structural consistency constraint corrects the geometric structure by calculating the differences in edge and texture gradients between the fused result and the reference modality, ensuring the fused image retains the contour information of the reference modality. Spectral consistency constraint confirms the physical realism of the brightness and color of the fused output by comparing the weighted average difference between the fused features and the features of each modality. Semantic consistency constraint confirms high-level semantic consistency by comparing the difference between the predicted semantic distribution of the fused features and the true semantic label distribution. During the optimization process, the weight coefficients of the three constraints are dynamically adjusted. When an increase in structural error is detected, the structural consistency constraint is strengthened first. When the spectral deviation exceeds the spectral consistency deviation threshold, the weight of the spectral consistency constraint is automatically increased. When the semantic distribution change of the fused result exceeds the semantic distribution offset threshold, the weight of the semantic consistency constraint is strengthened. The model parameters are updated through continuous iteration and backpropagation to optimize the structural consistency, spectral consistency, and semantic consistency constraint terms, performing joint optimization processing on the fused features.

[0035] It should be noted that a preferred scheme for consistency constraint optimization and feature reconstruction of fused features specifically includes: calculating structural consistency constraints by iteratively correcting the geometric shape and structural contour through calculating the differences between the fused result and the reference mode in edge intensity, gradient direction, and local texture distribution; and calculating structural consistency constraints. , represented as:

[0036]

[0037] in, This represents the fused hierarchical feature map. This represents the selected reference modal feature used to maintain structural edge consistency.

[0038] when When the fusion feature is highly matched with the reference modal structure, the edges are clear and there is no obvious deformation, and the existing structural constraint weights are maintained.

[0039] when When this occurs, it indicates that there is a slight misalignment or ambiguity in some structural regions, and the structural constraint weights are adjusted accordingly. Increase by 10%-15%.

[0040] when If the fusion result shows significant mismatch or edge breakage in the geometric edge region, a local structure correction mechanism is activated. This mechanism recalculates the gradient direction in the edge region and performs affine compensation, while simultaneously adjusting the structural consistency constraint weights. Increase by 30%.

[0041] Spectral consistency constraints are achieved by constructing a spectral mapping loss function, comparing the differences between the fused features and each mode in band response and brightness distribution, and updating the spectral weight matrix using a weighted average. When local brightness or color deviations from the standard response curve are detected in the fused result, the spectral mapping parameters are automatically re-estimated, compensation corrections are performed on the deviation areas, and the spectral consistency constraints are calculated. , represented as:

[0042]

[0043] in, Represented as the first Overall characteristics of a modality Represented as the first Spectral weighting coefficients for each mode, This represents the total number of modes.

[0044] when This indicates that the brightness and color of the fused image are consistent with the weighted average of each modality, the physical spectral characteristics are well maintained, and the original spectral weight configuration is preserved.

[0045] when This indicates that the fusion result has slight brightness deviations or color shifts in some areas, and the spectral consistency constraint weights are adjusted accordingly. Improvement by 20%, while re-estimating the spectral weighting coefficients of each mode. .

[0046] when When this occurs, it indicates that the fused image has significant brightness or color distortion, triggering the spectral recalibration mechanism to retrain the spectral mapping network and forcibly limit the output brightness range to [0,1].

[0047] Semantic consistency constraints are implemented by establishing a distance function between the predicted semantic distribution and the true semantic labels to monitor the shift in high-level semantics. When the difference in semantic distribution exceeds a threshold, the semantic layer features are adaptively adjusted through backpropagation to calculate the semantic consistency constraint. , represented as:

[0048]

[0049] in, This is represented as the Kullback-Leibler divergence, used to measure the similarity of distributions. This is represented as the semantic probability distribution of the fused features. It is represented as the target semantic distribution from the annotation or teacher model.

[0050] when When the semantic distribution of the fused features is highly consistent with the target label, the current semantic constraint strength is maintained.

[0051] when This indicates a slight shift in semantic distribution, with some category predictions being unstable; therefore, the semantic consistency constraint weights are enhanced. Approximately 10%-15%.

[0052] when When the semantic distribution of the fusion result deviates significantly from the reference model, a semantic recalibration mechanism is initiated. This mechanism involves introducing auxiliary semantic supervision or a teacher network to retrain the semantic distribution in order to restore the semantic consistency of the fusion features.

[0053] Through joint optimization, the fused features maintain consistency in the three-dimensional space (structure-spectrum-semantics), and the total loss is calculated. , represented as:

[0054]

[0055] in, The structural consistency constraint weight parameters, For spectral consistency constraint weight parameters, The semantic consistency constraint weight parameter is used to control the relative importance of each loss term.

[0056] It should be noted that generating the fused image and outputting interpretable saliency results includes: upsampling the multi-layer fusion features layer by layer to generate the fused image; generating a saliency distribution map to indicate the contribution area of ​​each modality feature; in the image reconstruction stage, restoring the original spatial resolution through deconvolution operation and retaining early feature information in each layer using skip connections; when the output image pixel value is detected to exceed the preset output pixel amplitude threshold range, a normalization unit is automatically triggered to adjust the pixel data amplitude; in the interpretability analysis stage, a saliency distribution map is calculated based on the degree of influence of different modality features on the output category, and the contribution area of ​​each modality feature during the fusion process is displayed in a visual manner.

[0057] It should also be noted that when the output image pixel value is detected to exceed the preset output pixel amplitude threshold range, a preferred scheme for automatically triggering the normalization unit to adjust the pixel data amplitude specifically includes: setting the preset output pixel amplitude threshold range; determining whether the pixel value of the fused and reconstructed image is within the effective dynamic range; the value is set according to the image type: for normalized floating-point images, it is set to [0.00, 1.00], and for 8-bit integer images, it is set to [0, 255]. When the pixel amplitude range is [0.00, 1.00], it indicates that the output image brightness and color are within the normal range, the image has not experienced numerical overflow, the current output result is maintained, and no adjustment is required; when the pixel amplitude slightly exceeds the upper limit of the threshold (1.00 < pixel value ≤ 1), the threshold value is adjusted accordingly. When the pixel value is 10 or 255 < pixel value ≤ 280, it indicates a slight brightness drift during the fusion and reconstruction stage. The normalization unit is automatically triggered to perform linear compression mapping, scaling the pixel value back to the range of [0.00, 1.00] or [0, 255]. When the pixel value exceeds a large range (pixel value > 1.10 or > 280), it indicates that the image has obvious brightness saturation or numerical overflow. The forced normalization correction mechanism is triggered, and saturation value clipping and local recalibration operations are performed on the pixels that exceed the limit. When the pixel value is detected to be below the lower limit (pixel value < 0.00 or < 0), it indicates that the image has numerical drift or reverse overflow. Offset compensation processing is performed, and the overall grayscale of the image is returned to the normal range through global brightness translation correction.

[0058] It should also be noted that by performing consistency constraint optimization and feature reconstruction processing on the fused features, fusion feature optimization under the consistency of structure, spectrum, and semantics is achieved. Three constraint mechanisms are established in the fusion space: structural constraints correct edge and texture errors, spectral constraints maintain the physical consistency of brightness and color, and semantic constraints ensure the coherence of high-level semantic distribution. The weights of the three types of constraints are dynamically adjusted according to the error changes, so that the model maintains feature balance in iterative optimization. The features after joint optimization are reconstructed at high resolution through layer-by-layer upsampling and deconvolution, and a saliency distribution map is generated to identify the contribution areas of each modality. Through this step, the fused features achieve geometric, spectral, and semantic consistency in multidimensional space, and the output image has high reliability in both visual quality and semantic interpretability.

[0059] Example 2, an embodiment of the present invention, provides an image processing system based on hierarchical fusion of multi-source sensor data, including a sensor calibration module, a feature fusion module, and a reconstruction optimization module.

[0060] The sensor calibration module is used to perform geometric registration and spectral normalization on data collected by multi-source sensors; the feature fusion module is used to extract hierarchical features from the standardized multimodal data and calculate cross-modal attention weights; the reconstruction optimization module is used to perform consistency constraint optimization and feature reconstruction on the fused features, generate a fused image and output interpretability and significance results.

Claims

1. An image processing method based on hierarchical fusion of multi-source sensor data, characterized in that, include: Geometric registration and spectral normalization are performed on data acquired from multiple sensors. Hierarchical features are extracted from the standardized multimodal data, and cross-modal attention weights are calculated. Consistency constraint optimization and feature reconstruction are performed on the fused features to generate a fused image and output interpretability saliency results.

2. The image processing method based on hierarchical fusion of multi-source sensor data as described in claim 1, characterized in that: The geometric registration includes, By calculating key feature points between images or point clouds from different sensors, corresponding point pairs are determined, and spatial alignment between multiple modalities is performed by establishing spatial mapping relationships. Feature extraction and descriptor calculation are performed on images from each sensor, and matching pairs are established based on feature similarity. If the matching error exceeds the preset error threshold, then perform local affine transformation iterative correction until all modal data are aligned in a unified spatial reference coordinate system. During the registration process, the preset error threshold is dynamically adjusted based on the residual value after each iteration. When the residual converges to the allowable error range, the final spatial correction result is output.

3. The image processing method based on hierarchical fusion of multi-source sensor data as described in claim 2, characterized in that: The spectral normalization process includes, For the registered multimodal data, spectral response characteristics are calculated and a unified brightness and band mapping model is established; The image brightness distribution and color histogram of each modality are analyzed, and the data are linearly or nonlinearly mapped according to the preset standard response curve so that the grayscale distribution of each modality is in a similar range. When the difference in spectral brightness between different modes exceeds the set threshold, the dynamic correction unit is automatically triggered to recalculate the spectral mapping parameters and update the conversion coefficients. Based on comparative analysis of the average brightness differences of each mode, the output is continuously adjusted through an adaptive correction function to ensure that the final spectral characteristics remain consistent in the multimodal fusion space.

4. The image processing method based on hierarchical fusion of multi-source sensor data as described in claim 3, characterized in that: The extraction of hierarchical features from the standardized multimodal data includes, By using a feature extraction network composed of multiple hierarchical structures, local texture features, structural features, and high-level semantic information of images at different spatial resolutions are extracted. Feature representations at different scales are extracted step by step through hierarchical encoding. The bottom layer captures edge and detail information, the middle layer extracts shape and region contours, and the top layer extracts global context information through a deep attention mechanism. During feature extraction, the network automatically adjusts the convolution kernel size and the number of channels to encode input features at different scales step by step. When gradient instability or feature extraction overfitting is detected, the residual connection mechanism is enabled to jointly input low-level and high-level information. The encoding network shares some parameters across modalities to maintain a consistent feature distribution and performs local adaptive adjustments based on modal differences.

5. The image processing method based on hierarchical fusion of multi-source sensor data as described in claim 4, characterized in that: The calculation of cross-modal attention weights includes, Based on the similarity of features between different modalities, the contribution weight of each modality in the fusion process is automatically assigned. By mapping the feature representations of each modality to a unified feature space, the correlation between different modalities is calculated, and the features are weighted according to the degree of correlation. During the weight allocation process, if a large modal difference is detected, local feature enhancement or suppression operations are performed, and the weights of the current modality are dynamically adjusted through the attention mechanism.

6. The image processing method based on hierarchical fusion of multi-source sensor data as described in claim 5, characterized in that: The process of performing consistency constraint optimization and feature reconstruction on the fused features includes, Three optimization mechanisms—structural consistency constraint, spectral consistency constraint, and semantic consistency constraint—are established simultaneously in the fused feature space. The structural consistency constraint corrects the geometric structure by calculating the difference between the fused result and the reference mode in the edge and texture gradients, so that the fused image retains the contour information of the reference mode. Spectral consistency constraints confirm that the brightness and color of the fused output maintain physical authenticity by comparing the weighted average difference between the fused features and the features of each modality. Semantic consistency constraints confirm high-level semantic consistency by comparing the degree of difference between the predicted semantic distribution of fused features and the true semantic label distribution. During the optimization process, the weight coefficients of the three constraints are dynamically adjusted. When an increase in structural error is detected, the structural consistency constraint is strengthened first. When the spectral deviation exceeds the spectral consistency deviation threshold, the weight of the spectral consistency constraint is automatically increased; When the semantic distribution change of the fusion result exceeds the semantic distribution offset threshold, the semantic consistency constraint weight is enhanced. The model parameters are updated through continuous iteration and backpropagation, and the structural consistency, spectral consistency, and semantic consistency constraints are optimized to jointly optimize the fused features.

7. The image processing method based on hierarchical fusion of multi-source sensor data as described in claim 6, characterized in that: The generation of the fused image and the output of interpretable saliency results include, The multi-layer fusion features are upsampled layer by layer to generate a fusion image, and a saliency distribution map is generated to indicate the contribution area of ​​each modality feature; In the image reconstruction stage, the original spatial resolution is restored through deconvolution operation, and skip connections are used in each layer to preserve early feature information; When the output image pixel value is detected to exceed the preset output pixel amplitude threshold range, the normalization unit is automatically triggered to adjust the pixel data amplitude. During the interpretability analysis phase, a significance distribution map is calculated based on the degree of influence of different modal features on the output category, and the contribution area of ​​each modal feature during the fusion process is displayed in a visual way.

8. An image processing system based on hierarchical fusion of multi-source sensor data, employing the image processing method based on hierarchical fusion of multi-source sensor data as described in any one of claims 1 to 7, characterized in that: Includes a sensor calibration module, a feature fusion module, and a reconstruction optimization module; The sensor calibration module is used to perform geometric registration and spectral normalization processing on data collected by multi-source sensors. The feature fusion module is used to extract hierarchical features from the standardized multimodal data and calculate cross-modal attention weights. The reconstruction optimization module is used to perform consistency constraint optimization and feature reconstruction processing on the fused features, generate a fused image, and output interpretability saliency results.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the image processing method based on multi-source sensor data hierarchical fusion as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the image processing method based on hierarchical fusion of multi-source sensor data as described in any one of claims 1 to 7.