Multi-modal image fusion method and device based on multi-level alignment and task perception
By employing a multi-level alignment and task-aware approach, the problem of multimodal image fusion in heterogeneous images and complex environments was solved, achieving accurate fusion of object-level image content and improving image detail fidelity and recognition accuracy.
Patent Information
- Application Number
- CN202511593368.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-02-10
AI Technical Summary
Existing multimodal image fusion methods suffer from several problems when dealing with images of different sizes and complex environments. These problems include: scale inconsistency leading to geometric distortion of the target; global scale alignment ignoring local differences; loss of target region information and lack of task perception; instability of fusion in irregular environments; and error accumulation caused by separation of alignment and fusion.
Employing a multi-level alignment and task-aware approach, this method combines radiometric correction, scale normalization, global affine and local geometric extension, minimum stretching constraints, and task saliency-driven fusion to achieve accurate fusion of object-level image content.
It improves image detail fidelity, enhances geometric consistency and robustness, ensures the stability and precise alignment of target contours and details in complex environments, and improves object-level recognition accuracy.
Smart Images

Figure CN121504740A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and in particular to a multimodal image fusion method and apparatus based on multi-level alignment and task awareness. Background Technology
[0002] With the rapid development of multi-sensor imaging and intelligent sensing technologies, multimodal image fusion has gradually become an important means to enhance perception capabilities in fields such as power line inspection, rail transit, security monitoring, autonomous driving, and industrial inspection. Different modalities of images each have their advantages: visible light images possess high spatial resolution and rich texture details; infrared images can reflect temperature and thermal radiation characteristics; depth images provide spatial geometry and structural information; and radar or SAR images maintain stable imaging capabilities under adverse weather conditions. Therefore, by fusing multi-source heterogeneous data, more robust target detection, recognition, and semantic understanding can be achieved in complex environments.
[0003] While existing multimodal fusion methods have effectively improved image fusion performance to some extent, they still have the following limitations when dealing with images of different sizes and irregular complex environments:
[0004] 1) Inconsistent scale leads to geometric distortion of the target.
[0005] Existing methods typically unify the image size acquired by different sensors through simple interpolation or cropping. These methods do not consider the geometric consistency of the target at different resolutions, and are prone to introducing edge blurring or shape stretching, especially for slender or small-scale targets (such as power lines, traffic signs, and the edges of people and vehicles), leading to a significant decrease in the accuracy of subsequent detection and segmentation.
[0006] 2) Global scale alignment ignores local differences
[0007] Images of varying sizes often exhibit non-uniform scaling, such as differences in scale between foreground and background regions. A single global affine or perspective transformation cannot simultaneously meet the scale requirements of different regions, resulting in over-compression or over-expansion of certain targets. Geometric distortion and ghosting occur in local target regions, directly affecting the consistency of object-level features.
[0008] 3) Loss of target area information and lack of task awareness
[0009] Most fusion methods only focus on the overall pixel consistency of the image, without introducing task-level target saliency constraints. In scenes with varying sizes, interpolation and registration often lead to the cropping or distortion of target edge regions, causing misalignment between the detection box / segmentation region and the actual target. As a result, although the fused image is visually coherent, it performs poorly in target recognition tasks.
[0010] 4) Unrobust fusion of heterogeneous targets in irregular environments
[0011] In complex environments (strong light, weak light, haze, uneven thermal radiation), the cross-modal feature differences of targets further amplify the difficulty of aligning different sizes. Traditional methods often rely on brightness or texture consistency, which is prone to failure in the target edge region, resulting in target contour mismatch, double images, or partial missing parts, making the object-level fusion results lack robustness.
[0012] 5) Alignment and fusion separation leads to error accumulation.
[0013] Existing technologies typically perform heterogeneous size registration first, followed by modal fusion, but the two lack a unified optimization objective. This pipeline-like operation accumulates errors within the target area: for example, insufficient handling of target deformation during the registration stage, and inability to compensate for geometric misalignment during the fusion stage, ultimately leading to inconsistencies between the target contour and semantic content in the fusion result. Summary of the Invention
[0014] The main technical problems solved by this invention include: how to overcome the many defects in the existing technology when facing complex environments and object-level applications, and achieve accurate fusion of object-level image content.
[0015] To address the aforementioned issues, this invention proposes a multi-level alignment and task-aware multimodal image fusion method. Through in-depth analysis of the shortcomings of existing technologies, six technical challenges are identified: heterogeneous size processing, multimodal fusion, adaptation to irregular environments, global-local alignment, geometric extension constraints, and affine compensation. Based on this, the inventors propose a novel methodological framework combining scale normalization, radiometric correction, global affine and local geometric extension, minimum stretching constraints, and task saliency-driven fusion, achieving accurate fusion of object-level image content. This approach not only improves upon the shortcomings of existing methods but also proposes a new overall framework for alignment and fusion mechanisms.
[0016] In a first aspect, embodiments of the present invention provide a multimodal image fusion method based on multi-level alignment and task awareness, the method comprising:
[0017] Acquire source modal images of different sizes of the target region from different sensors, perform radial symmetry correction on each source modal image, and perform scale normalization on each source modal image after radial symmetry correction.
[0018] After global affine initialization of each source modality image after scale normalization, local geometric expansion is performed. Based on task saliency and geometric constraints, region fusion is performed on each source modality image after local geometric expansion to obtain a fused source modality image after region fusion.
[0019] The fused source modal image is subjected to occlusion / hole detection and completion. The completed fused source modal image is corrected in the reflectivity domain. The corrected fused source modal image is input into the task network for processing, and the object-level task result is output.
[0020] Using a loss function, the fused source modal image after region fusion, the object-level task result, each source modal image after local geometric expansion, the fused source modal image after completion, and each source modal image after radial symmetry are optimized to obtain the final fused source modal image and the object and task result.
[0021] In a specific embodiment of the present invention, the step of performing radial symmetry correction on each of the source modal images includes:
[0022] Transform each of the source modal images to the logarithmic domain. ,in, For source modal images, For reference modal image, It is a modal set. It is a modal type. It is a tiny positive number, used to prevent the value from becoming unstable when the logarithm of zero is taken.
[0023] The learned mapping function is applied to the source modality image to obtain the radiation-symmetric source modality image. The optimization objective is to minimize the difference between the transformed sample source modal image and the reference sample modal image in the logarithmic domain, thereby optimizing the mapping function to obtain the optimized mapping function.
[0024] In a specific embodiment of the present invention, the scale normalization of the source modal images after radiation symmetry includes:
[0025] Set normalization optimization target :
[0026]
[0027] in Indicates the first Layered Laplacian energy operators are used to construct multi-scale pyramid representations of images, with each pyramid layer capturing structural information at different scales. It is the number of layers in the Laplacian capability operator, and Down is a differentiable scaling operation used to scalate the source modality image after radiation symmetry. Scale by a ratio s;
[0028] Calculate the scaled source modal image Compared with reference image Multiscale differences:
[0029] In the Energy characteristics on the layered Laplace pyramid , forming a difference
[0030] and its square norm 2 As intra-layer differences; accumulated across all layers. As an optimization target;
[0031] Using optimization algorithms The search space is used to find the optimal scaling factor that minimizes the total L2 norm difference.
[0032] After using the optimal scaling ratio, the source modal images are subjected to the corresponding scaling transformation, and the scale-normalized source modal images are output.
[0033] In a specific embodiment of the present invention, the step of performing global affine initialization on each of the scale-normalized source modal images, followed by local geometric expansion, includes:
[0034] A global affine model is used to establish a globally aligned overall structure. Global parameters , These are the pixel coordinates in each of the source modal images after scale normalization;
[0035] By utilizing a robust optimization objective function, the reference image is made possible. Gradient and transformed source image The minimum value of the gradient difference is used to obtain the optimal global affine transformation parameters. and ;
[0036] The robust optimization objective function is:
[0037]
[0038] in, It is gradient feature matching, used to compare the gradient information of the reference image and the transformed source image. It is a robust loss function, referring to the image domain. All pixels within, The reference image ;
[0039] Utilizing the optimized global affine transformation parameters and The global affine result is obtained. ;
[0040] Reference image domain Divided into K grids / superpixels A local linear extension is superimposed on the global affine result:
[0041]
[0042] in, It is a global affine result;
[0043] It is a local expansion. It is a partition indicator, a local parameter. It is the first A local region in mode Local linear transformation matrix on, local parameters It is the local translation vector of the k-th local region in mode m;
[0044] Constructing heterogeneous size consistent energy targets The heterogeneous size uniform energy target This includes gradient consistency terms and structure tensor consistency terms, minimum stretching constraint terms, and smoothing regularization terms;
[0045]
[0046] Among them, the gradient consistency term and the structure tensor consistency term are:
[0047]
[0048] Minimum tensile constraint term:
[0049] Smoothing regularization terms:
[0050] in, Represents the set of adjacent partition pairs. and It is the partition index number ( ∈ K is the number of partitions. It is a set of neighborhood relationships, containing all adjacent partition pairs; It is an unordered pair, representing a partition. and partitions Adjacent; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first Each partition in the model Local translation vectors on each, It is a weighted hyperparameter used to balance the importance of the gradient consistency term in the overall alignment loss function;
[0051] A hierarchical alternating optimization framework is adopted, with fixed global affine transformation parameters. and Parallel optimization of local parameters and Fix local parameters and iteratively update global parameters. and When the preset convergence condition is met: the change in the energy function The global affine transformation parameters are obtained. and and the local parameters and ;
[0052] This results in the structure after local geometric expansion:
[0053] .
[0054] In a specific embodiment of the present invention, the step of performing region fusion on each of the source modal images after local geometric expansion based on task saliency and geometric constraints to obtain a fused source modal image includes:
[0055] By minimizing the comprehensive energy function, the source modal images are fused to obtain the optimal fusion domain:
[0056]
[0057] in, It is the geometric deformation cost, used to measure the cumulative geometric alignment error within the region Ω. Reference image coordinate system A region or subset thereof;
[0058] This refers to task coverage, which ensures that the fusion area covers key content related to the task. It is the average of the task saliency maps within region Ω; the task saliency maps are generated from pre-trained... generate;
[0059] It refers to boundary discontinuity, used to penalize the unnaturalness of the merged boundary;
[0060] It is a context-expansion reward, used to incentivize the inclusion of task-related contextual information. It is the radius of the expansion operation;
[0061] It is a weight that controls the tolerance for geometric deformation. It is a weight that controls the balance between task performance and the compactness of the fusion region. Weights that control the naturalness of the fusion domain boundaries. It is a weight that adjusts the degree of inclusion of contextual information.
[0062] In a specific embodiment of the present invention, the step of performing region fusion on each of the source modal images after local geometric expansion based on task saliency and geometric constraints to obtain a fused source modal image includes:
[0063] Set detection conditions to obtain a set of void regions. ;
[0064] For each pixel position x, a hole is determined if any of the following conditions are met:
[0065] Condition 1: ;
[0066] Condition 2: No effective support for cross-modal communication;
[0067] Among them, in condition 1 These are the pixel coordinates in the source modal image. express ; It is the position deviation threshold, which indicates that if the transformation is irreversible or there is a large error, the registration at that position is unreliable.
[0068] Condition 2, cross-modal validity, means checking whether there are valid corresponding pixels in the source modality, excluding invalid values, boundary overflow, sensor blind spots, etc., to ensure that each pixel has reliable cross-modal support;
[0069] By optimizing the objective function and constraints, the set of hole regions detected during multimodal image registration is optimized. Perform minimum stretching completion;
[0070] The objective function and constraints are as follows:
[0071]
[0072]
[0073] in, It is a position correction vector, which defines the mapping relationship from the initial registration position to the optimized position;
[0074] Smoothing terms The requirement is to make the entire surface as smooth as possible within the void areas. It is the gradient of the entire field;
[0075] It ensures a smooth transition between the completed area and the surrounding known area; It is the boundary of the hollow area; These are the known pixel values on the boundary. These are boundary matching weight coefficients that control the strength of the connection.
[0076] It is an area conservation constraint that avoids excessive stretching or compression during the completion process; The soft constraint threshold is equivalent to area conservation completion guided by ARAP / Poisson. This constraint keeps the local area unchanged, thereby avoiding excessive expansion or contraction during completion and ensuring the target shape.
[0077] In a specific embodiment of the present invention, the step of performing region fusion on each of the source modal images after local geometric expansion based on task saliency and geometric constraints to obtain a fused source modal image includes:
[0078] Reflectivity decomposition is used to decompose the source modality image of each modality into two physical components: ;
[0079] in, It is the source modal image, Indicates the lighting components, Represents the reflectivity component;
[0080] The illumination components are solved using the Tikhonov regularization method. :
[0081]
[0082] Among them, data fidelity item To ensure that the reconstruction error of the decomposed image is minimized. This is the initial estimate of the reflectivity component;
[0083] Smoothing regularization terms The lighting components are required to vary smoothly in space (in accordance with the characteristics of light). It is a regularization parameter that controls the smoothness; it utilizes the obtained illumination components. Calculate reflectivity Using the reflectivity Noise reliability weight and cross-modal consistency value The fusion weights are obtained as follows:
[0084]
[0085] in, Control the contribution level of each factor separately;
[0086] Thus, a reflectivity domain fused image is obtained. :
[0087]
[0088] Obtain the final reconstructed image :
[0089]
[0090] in, It is the fused reflectance image, representing the weighted fusion result of all modal reflectance information at pixel position x; It uses the illumination component of the reference mode. The final multimodal fused image represents the optimized fusion result generated after the complete processing at pixel position x; The data is directly fed into the task network, and the object-level task results are output. .
[0091] Secondly, embodiments of the present invention provide a multimodal image fusion apparatus based on multi-level alignment and task awareness, the apparatus comprising:
[0092] The first processing module is used to acquire source modal images of different sizes of the target area from different sensors, perform radial symmetry correction on each of the source modal images, and perform scale normalization on each of the radially symmetric source modal images.
[0093] The second processing module is used to perform global affine initialization on each of the source modal images after scale normalization, perform local geometric expansion, and perform region fusion on each of the source modal images after local geometric expansion based on task saliency and geometric constraints to obtain a fused source modal image after region fusion.
[0094] The third processing module is used to perform occlusion / hole detection and completion on the fused source modal image, correct the completed fused source modal image in the reflectivity domain, input the corrected fused source modal image into the task network for processing, and output the object-level task result.
[0095] The fourth processing module is used to optimize the fused source modal image after region fusion, the object-level task result, each source modal image after local geometric expansion, the fused source modal image after completion, and each source modal image after radial symmetry using a loss function, so as to obtain the final fused source modal image and the object and task result.
[0096] In a specific embodiment of the present invention, the first processing module includes:
[0097] Transform each of the source modal images to the logarithmic domain. ,in, For source modal images, For reference modal image, It is a modal set. It is a modal type. It is a tiny positive number, used to prevent the value from becoming unstable when the logarithm of zero is taken.
[0098] The learned mapping function is applied to the source modality image to obtain the radiation-symmetric source modality image. The optimization objective is to minimize the difference between the transformed sample source modal image and the reference sample modal image in the logarithmic domain, thereby optimizing the mapping function to obtain the optimized mapping function.
[0099] In a specific embodiment of the present invention, the first processing module includes:
[0100] Set normalization optimization target :
[0101]
[0102] in Indicates the first Layered Laplacian energy operators are used to construct multi-scale pyramid representations of images, with each pyramid layer capturing structural information at different scales. It is the number of layers in the Laplacian capability operator, and Down is a differentiable scaling operation used to scalate the source modality image after radiation symmetry. proportionally Scaling;
[0103] Calculate the scaled source modal image Compared with reference image Differences in the L2 norm of energy characteristics at each pyramid level;
[0104] Using optimization algorithms The search space is used to find the optimal scaling factor that minimizes the total L2 norm difference.
[0105] After using the optimal scaling ratio, the source modal images are subjected to the corresponding scaling transformation, and the scale-normalized source modal images are output.
[0106] In a specific embodiment of the present invention, the second processing module includes:
[0107] A global affine model is used to establish a globally aligned overall structure. Global parameters , These are the pixel coordinates in each of the source modal images after scale normalization;
[0108] By utilizing a robust optimization objective function, the reference image is made possible. Gradient and transformed source image The minimum value of the gradient difference is used to obtain the optimal global affine transformation parameters. and ;
[0109] The robust optimization objective function is:
[0110]
[0111] in, It is gradient feature matching, used to compare the gradient information of the reference image and the transformed source image. It is a robust loss function, referring to the image domain. All pixels within, The reference image ;
[0112] Utilizing the optimized global affine transformation parameters and The global affine result is obtained. ;
[0113] Reference image domain Divided into K grids / superpixels A local linear extension is superimposed on the global affine result:
[0114]
[0115] in, It is a global affine result;
[0116] It is a local expansion. It is a partition indicator, a local parameter. It is the local linear transformation matrix of the k-th local region on mode m, and the local parameters are... It is the local translation vector of the k-th local region in mode m;
[0117] Constructing heterogeneous size consistent energy targets The heterogeneous size uniform energy target This includes gradient consistency terms and structure tensor consistency terms, minimum stretching constraint terms, and smoothing regularization terms;
[0118]
[0119] Among them, the gradient consistency term and the structure tensor consistency term are:
[0120]
[0121] Minimum tensile constraint term:
[0122] Smoothing regularization terms:
[0123] in, Represents the set of adjacent partition pairs. and It is the partition index number ( (∈ {1,2, ..., K}) where K is the number of partitions. It is a set of neighborhood relationships, containing all adjacent partition pairs; It is an unordered pair, representing a partition. and partitions Adjacent; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first Each partition in the model Local translation vectors on each, It is a weighted hyperparameter used to balance the importance of the gradient consistency term in the overall alignment loss function;
[0124] A hierarchical alternating optimization framework is adopted, with fixed global affine transformation parameters. and Parallel optimization of local parameters and Fix local parameters and iteratively update global parameters. and When the preset convergence condition is met: the change in the energy function The global affine transformation parameters are obtained. and and the local parameters and ;
[0125] This results in the structure after local geometric expansion:
[0126] .
[0127] In a specific embodiment of the present invention, the second processing module includes:
[0128] By minimizing the comprehensive energy function, the source modal images are fused to obtain the optimal fusion domain:
[0129]
[0130] in, It is the cost of geometric deformation, used to measure the region. Cumulative geometric alignment error within, Reference image coordinate system A region or subset thereof;
[0131] This refers to task coverage, which ensures that the fusion area covers key content related to the task. It is the average of the task saliency map within region Ω; the task saliency map is generated by the pre-trained task network. generate; It refers to boundary discontinuity, used to penalize the unnaturalness of the merged boundary; It is a context-expansion reward, used to incentivize the inclusion of task-related contextual information. It is the radius of the expansion operation; It is a weight that controls the tolerance for geometric deformation. It is a weight that controls the balance between task performance and the compactness of the fusion region. Weights that control the naturalness of the fusion domain boundaries. It is a weight that adjusts the degree of inclusion of contextual information.
[0132] In a specific embodiment of the present invention, the third processing module includes:
[0133] Set detection conditions to obtain a set of void regions. ;
[0134] For each pixel position x, a hole is determined if any of the following conditions are met:
[0135] Condition 1: ;
[0136] Condition 2: No effective support for cross-modal communication;
[0137] Among them, in condition 1 These are the pixel coordinates in the source modal image. express ; It is the position deviation threshold, which indicates that if the transformation is irreversible or there is a large error, the registration at that position is unreliable.
[0138] Condition 2, cross-modal validity, means checking whether there are valid corresponding pixels in the source modality, excluding invalid values, boundary overflow, sensor blind spots, etc., to ensure that each pixel has reliable cross-modal support;
[0139] By optimizing the objective function and constraints, the set of hole regions detected during multimodal image registration is optimized. Perform minimum stretching completion;
[0140] The objective function and constraints are as follows:
[0141]
[0142]
[0143] in, It is the position correction vector, which defines the mapping relationship from the initial registration position to the optimized position; the smoothing term The requirement is to make the entire surface as smooth as possible within the void areas. It is the gradient of the entire field; It ensures a smooth transition between the completed area and the surrounding known area; It is the boundary of the hollow area; These are the known pixel values on the boundary. These are boundary matching weight coefficients that control the strength of the connection.
[0144] It is an area conservation constraint that avoids excessive stretching or compression during the completion process; The soft constraint threshold is equivalent to area conservation completion guided by ARAP / Poisson. This constraint keeps the local area unchanged, thereby avoiding excessive expansion or contraction during completion and ensuring the target shape.
[0145] In a specific embodiment of the present invention, the third processing module includes:
[0146] Reflectivity decomposition is used to decompose the source modality image of each modality into two physical components:
[0147] ;
[0148] in, It is the source modal image, Indicates the lighting components, Represents the reflectivity component;
[0149] The illumination components are solved using the Tikhonov regularization method. :
[0150]
[0151] Among them, data fidelity item To ensure that the reconstruction error of the decomposed image is minimized. This is the initial estimate of the reflectivity component;
[0152] Smoothing regularization terms The lighting components are required to vary smoothly in space (in accordance with the characteristics of light).
[0153] It is a regularization parameter that controls the degree of smoothness;
[0154] Using the obtained illumination components Calculate reflectivity
[0155] Using the reflectivity Noise reliability weight and cross-modal consistency value The fusion weights are obtained as follows:
[0156]
[0157] in, Control the contribution level of each factor separately;
[0158] Thus, a reflectivity domain fused image is obtained. :
[0159]
[0160] Obtain the final reconstructed image :
[0161]
[0162] in, It is the fused reflectance image, representing the weighted fusion result of all modal reflectance information at pixel position x; It uses the illumination component of the reference mode. The final multimodal fused image represents the optimized fusion result generated after the complete processing at pixel position x; The data is directly fed into the task network, and the object-level task results are output. .
[0163] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multimodal image fusion method based on multi-level alignment and task awareness.
[0164] Fourthly, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the multimodal image fusion method based on multi-level alignment and task awareness as described above.
[0165] Compared with existing technologies, it has the following outstanding advantages:
[0166] Compared with existing technologies, this invention has significant advantages in object-level processing of multimodal heterogeneous image fusion:
[0167] Firstly, in the processing of images of different sizes, this invention avoids detail blurring and target distortion caused by traditional interpolation or cropping methods by introducing scale normalization and geometry preservation mechanisms, achieving object-level geometric consistency under different resolution inputs. This improvement effectively reduces the loss of target region information and improves the detail fidelity of the fused image.
[0168] Secondly, regarding multimodal fusion, this invention effectively bridges the differences between multi-source imaging mechanisms such as visible light, infrared, and depth through radiometric correction and feature space mapping, establishing a unified representation space that is comparable across modalities. This design ensures accurate alignment of target boundaries under cross-modal conditions, significantly improving the clarity and semantic consistency of the fused target region.
[0169] Furthermore, regarding adaptability to irregular lighting and complex environments, this invention employs a radiation symmetry and environment-adaptive weighting mechanism to effectively suppress the interference of strong light, backlight, haze, and uneven thermal radiation on image fusion. Compared with existing methods, this invention significantly improves the stability and robustness of target region alignment, maintaining consistent target contours and details even in complex scenes.
[0170] Regarding the alignment mechanism, this invention proposes a multi-layered framework combining global affine mapping and local geometric expansion. This achieves rigid preservation of the global structure and flexible optimization of local targets, overcoming the shortcomings of traditional single models that cannot simultaneously consider both global and local aspects. This enables the fused image to remain stable not only in the overall space but also to achieve higher-precision matching in the target region.
[0171] In terms of geometric extension modeling, this invention introduces minimum stretching and area conservation constraints to ensure the morphological stability of the target during local deformation, avoiding distortion in slender targets and geometrically regular targets. This method makes the fusion results more consistent with the true structural characteristics of the target, improving the accuracy of object-level recognition.
[0172] Regarding the applicability of affine models, this invention overcomes the limitations of single affine models in complex viewpoints, focal lengths, and distortion scenarios by combining global affine constraints with local compensation mechanisms, ensuring a balance between global consistency and local flexibility. This improvement significantly enhances target alignment accuracy and fusion stability in complex scenarios.
[0173] In summary, this invention outperforms existing technologies in terms of image detail preservation, geometric consistency, fusion robustness, and task applicability. Its beneficial effects are manifested in: improved information fidelity: reducing target detail loss and retaining more object-level key information; enhanced geometric consistency: global and local coordination ensures stable target morphology; improved environmental adaptability: achieving stable fusion even under complex lighting and environmental conditions; and optimized task performance: the fusion results better serve downstream detection and segmentation tasks, improving recognition accuracy.
[0174] Enhanced deployment flexibility: The method design takes into account the computing power constraints of edge devices and cloud platforms, enabling highly efficient and usable intelligent applications. Therefore, this invention not only breaks through the limitations of existing technologies in terms of theoretical methods, but also significantly improves the accuracy and robustness of target-level image fusion in engineering applications, possessing broad application value. Attached Figure Description
[0175] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0176] Figure 1 This is a schematic diagram of the multimodal image fusion method based on multi-level alignment and task awareness of the present invention;
[0177] Figure 2 This is a schematic flowchart illustrating the method architecture of an embodiment of the present invention;
[0178] Figure 3 This is a schematic diagram of the system framework of an embodiment of the present invention;
[0179] Figure 4 This is a schematic diagram of a multimodal image fusion device module based on multi-level alignment and task awareness according to an embodiment of the present invention;
[0180] Figure 5 This is a schematic diagram of the computer hardware of the present invention. Detailed Implementation
[0181] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0182] It should also be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0183] It should also be understood that, in various embodiments of the present invention, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0184] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0185] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0186] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0187] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0188] To make the above-mentioned features and effects of the present invention clearer and easier to understand, specific embodiments are described below in conjunction with the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are merely illustrative. The scope of protection of the present invention is not limited to the disclosed embodiments, but is defined by the appended claims.
[0189] The following are system embodiments corresponding to the above method embodiments. This embodiment can be implemented in conjunction with the above embodiments. The relevant technical details mentioned in the above embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.
[0190] With the rapid development of multi-sensor imaging and intelligent sensing technologies, multimodal image fusion has gradually become an important means to enhance perception capabilities in fields such as power line inspection, rail transit, security monitoring, autonomous driving, and industrial inspection. Different modalities of images each have their advantages: visible light images possess high spatial resolution and rich texture details; infrared images can reflect temperature and thermal radiation characteristics; depth images provide spatial geometry and structural information; and radar or SAR images maintain stable imaging capabilities under adverse weather conditions. Therefore, by fusing multi-source heterogeneous data, more robust target detection, recognition, and semantic understanding can be achieved in complex environments.
[0191] Existing technologies mainly include the following types of methods.
[0192] The first category is registration and fusion methods based on traditional image processing. Typical methods use feature points (such as SIFT, SURF), edges, or mutual information for registration, and fuse them through wavelet transform, Laplacian pyramid, or multi-scale decomposition. These methods have the advantages of simple implementation and low computational cost, but in irregular lighting or low-contrast scenes, feature points are difficult to extract stably, leading to a decrease in registration accuracy. In addition, traditional methods usually assume that the input images have the same resolution, and images of different sizes can only be processed by simple interpolation or cropping, which often causes loss of detail and geometric distortion.
[0193] The second category is registration methods based on geometric and affine transformations. Common geometric registration methods include affine, perspective, and projection models. These methods can handle multi-viewpoint or resolution differences to some extent, but their core is still based on global transformations, making them difficult to adapt to local scale inconsistencies or non-uniform deformations. For scenes with perspective distortion, local stretching, or occlusion, global affine models often produce problems such as blurred blending boundaries, ghosting, and geometric distortion.
[0194] The third category is deep learning-based fusion methods. In recent years, with the development of Convolutional Neural Networks (CNNs), Generative Adversarial Networks (GANs), and Transformers, numerous end-to-end multimodal fusion methods have emerged, such as infrared and visible light image fusion methods. These methods, through deep feature extraction and learned cross-modal mappings, can mitigate the effects of uneven illumination to some extent and improve the performance of the fused target detection and recognition. However, deep learning methods typically require pre-scaling or cropping the input image to ensure dimensionality consistency, which still results in structural information loss for images of varying sizes. Furthermore, deep learning models have large parameters and are highly dependent on computational and storage resources, making them difficult to meet real-time requirements on edge devices.
[0195] Existing patented technologies also involve multimodal image fusion. For example, US11830222B2 proposes an infrared and visible light fusion method based on two-level optimization. This method uses binocular infrared and visible light cameras to acquire images and then uses mathematical modeling to divide the fused image into an image domain and a gradient domain for fusion. However, this method does not effectively address the alignment problem of images of different sizes. Furthermore, this patent does not address task-driven dynamic sensor selection mechanisms or consider how to automatically adjust sensor configurations according to different scenarios to improve the flexibility of image acquisition and processing. Although some fusion strategies have been proposed in the literature, they mainly focus on sensor combinations and processing flows in fixed scenarios, lacking specific methods for effectively fusing images of different sizes, and also failing to address how to improve target detection accuracy through adaptive sensor scheduling.
[0196] While existing multimodal fusion methods have effectively improved image fusion performance to some extent, they still have the following limitations when dealing with images of different sizes and irregular complex environments:
[0197] 1) Limited ability to process images of different sizes
[0198] Existing methods typically unify the image size acquired by different sensors through simple interpolation or cropping. These methods do not consider the geometric consistency of the target at different resolutions, and are prone to introducing edge blurring or shape stretching, especially for slender or small-scale targets (such as power lines, traffic signs, and the edges of people and vehicles), leading to a significant decrease in the accuracy of subsequent detection and segmentation.
[0199] (2) Multimodal fusion is not adaptable to cross-sensor characteristics.
[0200] Most existing methods rely on photometric consistency or simple feature matching to achieve fusion between different modalities. However, different sensors (visible light, infrared, depth, etc.) have different physical imaging mechanisms: visible light reflects reflection features, infrared reflects radiation features, and depth sensors provide geometry. These differences cause matching methods based directly on brightness or texture to fail in cross-modal scenes, resulting in target boundary mismatch or insufficient complementarity of modal information.
[0201] (3) Insufficient robustness under irregular lighting and complex environments
[0202] Existing registration and fusion methods lack the ability to model irregular environments. When there is strong light, backlight, shadows, fog, or uneven thermal radiation in the scene, traditional feature point or mutual information methods are prone to failure, resulting in misalignment, ghosting, or missing target areas. This is because existing methods assume that the imaging conditions of each modality are stable, while ignoring the dynamic changes in radiation characteristics and signal-to-noise levels under complex environments.
[0203] (4) Lack of a multi-level alignment mechanism that combines global and local aspects
[0204] Most current methods rely on global affine or projective transformation models to align multimodal images. While these models can handle overall scale differences well, they struggle to address situations where the target and background regions have local scale inconsistencies or non-uniform deformations. For example, in rail transit or autonomous driving scenarios, there are often significant scale differences between near-field targets and distant backgrounds. A single global model cannot simultaneously guarantee the geometric consistency of both, leading to distortion or misalignment of local targets.
[0205] (5) Lack of geometric extension modeling makes it difficult to maintain the target shape.
[0206] While some non-rigid registration methods introduce local deformation models, they lack sufficient constraints on local geometric expansion (expansion) and are deficient in area conservation or minimum stretching constraints. When dealing with slender targets (such as power lines, tower edges, and vehicle outlines), excessive stretching or compression of the target can easily occur, leading to geometric distortion. The root cause of this deficiency is that existing methods do not explicitly introduce local scale constraints and lack optimization mechanisms oriented towards preserving the target's shape.
[0207] (6) A single affine model is difficult to meet the alignment requirements of complex scenes.
[0208] While affine transformation is a common image registration technique, it assumes a linear mapping between modalities. In multimodal fusion scenarios, especially with non-uniform distortion, different focal lengths, or different viewpoints, a single affine model often cannot guarantee simultaneous alignment of the target region and the background, leading to ghosting or distortion in local areas. Current technologies lack comprehensive modeling methods that combine affine transformation with local geometric extension, thus limiting their adaptability to complex scenes.
[0209] To address the above issues, this invention proposes a scale normalization method for multimodal images of varying sizes. This method not only unifies resolution but also introduces target edge preservation and detail protection mechanisms. This avoids the problem of target detail loss caused by simple interpolation or cropping in existing technologies. By adding geometric consistency constraints during the normalization process, images of different resolutions can maintain object-level proportions and structural integrity during alignment. Furthermore, this mechanism effectively reduces the risk of blurring or distortion of small and elongated targets during the fusion process, ensuring the consistency of spatial structure and detail fidelity of the fused image, thereby improving the accuracy of subsequent detection and recognition.
[0210] This invention addresses the significant differences in imaging mechanisms across different modalities, such as visible light, infrared, and depth imaging. It employs a cross-modal feature alignment method, transforming images from different modalities into a unified representation space through radiometric correction and feature space mapping. Feature alignment and information fusion are then performed within this unified representation space. This approach fully considers the differences in spectral density, radiometric intensity, and signal-to-noise ratio across modalities, ensuring that fusion relies not only on pixel-level matching but also incorporates cross-modal feature consistency constraints. Furthermore, information from different modalities can achieve higher-precision alignment and complementarity within the target region. The fusion result, at the object level, exhibits clear target boundaries and semantic consistency, avoiding the modal mismatch and information loss problems common in traditional methods.
[0211] This invention employs a robustness enhancement method combining radiation symmetry correction and environment-adaptive weighting. By uniformly correcting the differences in illumination and radiation characteristics across different modalities and dynamically adjusting weights for overexposure, underexposure, shadows, haze, and uneven thermal radiation, it can suppress the influence of abnormal regions at the pixel level, highlighting reliable information in the target area. Furthermore, it significantly improves fusion stability in complex environments, ensuring that the target area maintains clear boundaries and consistent semantics even in strong light, backlight, or low-visibility scenes, effectively avoiding ghosting, information loss, or morphological distortion in the fusion results.
[0212] This invention employs a multi-layered alignment framework combining global affine transformation and local geometric extension. Global affine transformation maintains the consistency of the overall spatial structure, while local geometric extension performs fine-grained morphological adjustments within the target region. This hierarchical alignment strategy combines global stability with local flexibility, thus avoiding the problem of insufficient adaptation of a single affine model to local targets. This invention simultaneously ensures the stability of the overall structure and accurate local alignment of the target region, preventing boundary mismatches or geometric deformations of object-level targets due to global model limitations, and significantly improving the geometric consistency of the fused image.
[0213] This invention introduces a geometric expansion mechanism into local geometric modeling, combined with minimum stretching and area conservation constraints, to limit scale changes in the target region. This method ensures that the target is not excessively stretched or compressed during local alignment, while simultaneously constraining the target boundary and structural morphology during model optimization to guarantee geometric stability. The target region retains its original shape and proportions in the fusion result, effectively avoiding morphological distortion, especially for slender targets or targets with strong geometric regularity, thus improving the structural integrity and recognition accuracy of object-level targets.
[0214] This invention, based on traditional affine models, combines global affine constraints with local compensation mechanisms to maintain overall affine consistency. Simultaneously, it introduces nonlinear compensation and fine-tuning mechanisms in local regions, ensuring that different modal images maintain synchronous alignment between the target and background even in complex scenes. This method overcomes the shortcomings of single affine models under multi-viewpoint, different focal length, and distortion conditions. It guarantees the stability of the global structure while providing flexible adaptation to local targets, enabling the fusion result to achieve high-precision object-level alignment under various shooting conditions, thus improving the applicability of multimodal image fusion in complex practical applications.
[0215] The following key technical points emerged during the development of this invention:
[0216] Firstly, in the processing of images of different sizes, this invention avoids detail blurring and target distortion caused by traditional interpolation or cropping methods by introducing scale normalization and geometry preservation mechanisms, achieving object-level geometric consistency under different resolution inputs. This improvement effectively reduces the loss of target region information and improves the detail fidelity of the fused image.
[0217] Secondly, regarding multimodal fusion, this invention effectively bridges the differences between multi-source imaging mechanisms such as visible light, infrared, and depth through radiometric correction and feature space mapping, establishing a unified representation space that is comparable across modalities. This design ensures accurate alignment of target boundaries under cross-modal conditions, significantly improving the clarity and semantic consistency of the fused target region.
[0218] Furthermore, regarding adaptability to irregular lighting and complex environments, this invention employs a radiation symmetry and environment-adaptive weighting mechanism to effectively suppress the interference of strong light, backlight, haze, and uneven thermal radiation on image fusion. Compared with existing methods, this invention significantly improves the stability and robustness of target region alignment, maintaining consistent target contours and details even in complex scenes.
[0219] Regarding the alignment mechanism, this invention proposes a multi-layered framework combining global affine mapping and local geometric expansion. This achieves rigid preservation of the global structure and flexible optimization of local targets, overcoming the shortcomings of traditional single models that cannot simultaneously consider both global and local aspects. This enables the fused image to remain stable not only in the overall space but also to achieve higher-precision matching in the target region.
[0220] In terms of geometric extension modeling, this invention introduces minimum stretching and area conservation constraints to ensure the morphological stability of the target during local deformation, avoiding distortion in slender targets and geometrically regular targets. This method makes the fusion results more consistent with the true structural characteristics of the target, improving the accuracy of object-level recognition.
[0221] Regarding the applicability of affine models, this invention overcomes the limitations of single affine models in complex viewpoints, focal lengths, and distortion scenarios by combining global affine constraints with local compensation mechanisms, ensuring a balance between global consistency and local flexibility. This improvement significantly enhances target alignment accuracy and fusion stability in complex scenarios.
[0222] In summary, this invention outperforms existing technologies in terms of image detail preservation, geometric consistency, fusion robustness, and task applicability. Its beneficial effects are manifested in: improved information fidelity: reducing target detail loss and retaining more object-level key information; enhanced geometric consistency: global and local coordination ensures stable target morphology; improved environmental adaptability: achieving stable fusion even under complex lighting and environmental conditions; and optimized task performance: the fusion results better serve downstream detection and segmentation tasks, improving recognition accuracy.
[0223] Enhanced deployment flexibility: The methodology is designed to balance the computing power constraints of edge devices and cloud platforms, enabling highly efficient and usable intelligent applications.
[0224] Therefore, this invention not only breaks through the limitations of existing technologies in terms of theoretical methods, but also significantly improves the accuracy and robustness of target-level image fusion in engineering applications, and has broad application value.
[0225] like Figure 1 As shown, this embodiment of the invention provides a multimodal image fusion method based on multi-level alignment and task awareness, including:
[0226] 110. Acquire source modal images of different sizes of the target region from different sensors, perform radial symmetry correction on each of the source modal images, and perform scale normalization on each of the radially symmetric source modal images;
[0227] 120. After performing global affine initialization on each of the source modal images after scale normalization, local geometric expansion is performed. Based on task saliency and geometric constraints, region fusion is performed on each of the source modal images after local geometric expansion to obtain the fused source modal image after region fusion.
[0228] 130. Perform occlusion / hole detection and completion on the fused source modal image, correct the completed fused source modal image in the reflectivity domain, input the corrected fused source modal image into the task network for processing, and output the object-level task result;
[0229] 140. Using a loss function, optimize the fused source modal image after region fusion, the object-level task result, each source modal image after local geometric expansion, the fused source modal image after completion, and each source modal image after radial symmetry to obtain the final fused source modal image and the object and task result.
[0230] In a specific embodiment of the present invention, step 110 includes:
[0231] Transform each of the source modal images to the logarithmic domain. ,in, For source modal images, For reference modal image, It is a modal set. It is a modal type. It is a tiny positive number, used to prevent the value from becoming unstable when the logarithm of zero is taken.
[0232] The learned mapping function is applied to the source modality image to obtain the radiation-symmetric source modality image. The optimization objective is to minimize the difference between the transformed sample source modal image and the reference sample modal image in the logarithmic domain, thereby optimizing the mapping function to obtain the optimized mapping function.
[0233] In a specific embodiment of the present invention, step 110 includes:
[0234] Set normalization optimization target :
[0235]
[0236] in Indicates the first Layered Laplacian energy operators are used to construct multi-scale pyramid representations of images, with each pyramid layer capturing structural information at different scales. It is the number of layers in the Laplacian capability operator, and Down is a differentiable scaling operation used to scalate the source modality image after radiation symmetry. Scale by a ratio s;
[0237] Calculate the scaled source modal image Compared with reference image Multiscale differences:
[0238] In the Energy characteristics on the layered Laplace pyramid This creates a difference.
[0239] and its square norm 2 As intra-layer differences; accumulated across all layers. As an optimization target;
[0240] Using optimization algorithms The search space is used to find the optimal scaling factor that minimizes the total L2 norm difference.
[0241] After using the optimal scaling ratio, the source modal images are subjected to the corresponding scaling transformation, and the scale-normalized source modal images are output.
[0242] In a specific embodiment of the present invention, step 120 includes:
[0243] A global affine model is used to establish a globally aligned overall structure. Global parameters , These are the pixel coordinates in each of the source modal images after scale normalization;
[0244] By utilizing a robust optimization objective function, the reference image is made possible. Gradient and transformed source image The minimum value of the gradient difference is used to obtain the optimal global affine transformation parameters. and ;
[0245] The robust optimization objective function is:
[0246]
[0247] in, It is gradient feature matching, used to compare the gradient information of the reference image and the transformed source image. It is a robust loss function, referring to the image domain. All pixels within, The reference image ;
[0248] Utilizing the optimized global affine transformation parameters and The global affine result is obtained. ;
[0249] Reference image domain Divided into individual grids / superpixels A local linear extension is superimposed on the global affine result:
[0250]
[0251] in, It is a global affine result;
[0252] It is a local expansion. It is a partition indicator, a local parameter. It is the local linear transformation matrix of the k-th local region on mode m, and the local parameters are... It is the local translation vector of the k-th local region in mode m;
[0253] Constructing heterogeneous size consistent energy targets The heterogeneous size uniform energy target This includes gradient consistency terms and structure tensor consistency terms, minimum stretching constraint terms, and smoothing regularization terms;
[0254]
[0255] Among them, the gradient consistency term and the structure tensor consistency term are:
[0256]
[0257] Minimum tensile constraint term:
[0258] Smoothing regularization terms:
[0259] in, Represents the set of adjacent partition pairs. and It is the partition index number ( ∈ K is the number of partitions. It is a set of neighborhood relationships, containing all adjacent partition pairs; It is an unordered pair, representing a partition. and partitions Adjacent; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first Each partition in the model Local translation vectors on each, It is a weighted hyperparameter used to balance the importance of the gradient consistency term in the overall alignment loss function;
[0260] A hierarchical alternating optimization framework is adopted, with fixed global affine transformation parameters. and Parallel optimization of local parameters and Fix local parameters and iteratively update global parameters. and When the preset convergence condition is met: the change in energy function ΔL < threshold ε, the global affine transformation parameters are obtained. and and the local parameters and ;
[0261] This results in the structure after local geometric expansion:
[0262] .
[0263] In a specific embodiment of the present invention, step 120 includes:
[0264] By minimizing the comprehensive energy function, the source modal images are fused to obtain the optimal fusion domain:
[0265]
[0266] in, It is the geometric deformation cost, used to measure the cumulative geometric alignment error within the region Ω. Reference image coordinate system A region or subset thereof;
[0267] This refers to task coverage, which ensures that the fusion area covers key content related to the task. It is the average of the task saliency map within region Ω; the task saliency map is generated by the pre-trained task network. generate;
[0268] It refers to boundary discontinuity, used to penalize the unnaturalness of the merged boundary;
[0269] It is a context-expansion reward, used to incentivize the inclusion of task-related contextual information. It is the radius of the expansion operation;
[0270] It is a weight that controls the tolerance for geometric deformation. It is a weight that controls the balance between task performance and the compactness of the fusion region. Weights that control the naturalness of the fusion domain boundaries. It is a weight that adjusts the degree of inclusion of contextual information.
[0271] In a specific embodiment of the present invention, step 130 includes:
[0272] Set detection conditions to obtain a set of void regions. ;
[0273] For each pixel position x, a hole is determined if any of the following conditions are met:
[0274] Condition 1: ;
[0275] Condition 2: No effective support for cross-modal communication;
[0276] Among them, in condition 1 These are the pixel coordinates in the source modal image. express ; It is the position deviation threshold, which indicates that if the transformation is irreversible or there is a large error, the registration at that position is unreliable.
[0277] Condition 2, cross-modal validity, means checking whether there are valid corresponding pixels in the source modality, excluding invalid values, boundary overflow, sensor blind spots, etc., to ensure that each pixel has reliable cross-modal support;
[0278] By optimizing the objective function and constraints, the set of hole regions detected during multimodal image registration is optimized. Perform minimum stretching completion;
[0279] The objective function and constraints are as follows:
[0280]
[0281]
[0282] in, It is the position correction vector, which defines the mapping relationship from the initial registration position to the optimized position; the smoothing term The requirement is to make the entire surface as smooth as possible within the void areas. It is the gradient of the entire field; It ensures a smooth transition between the completed area and the surrounding known area; It is the boundary of the hollow area; These are the known pixel values on the boundary. These are boundary matching weight coefficients that control the strength of the connection. It is an area conservation constraint that avoids excessive stretching or compression during the completion process; The soft constraint threshold is equivalent to area conservation completion guided by ARAP / Poisson. This constraint keeps the local area unchanged, thereby avoiding excessive expansion or contraction during completion and ensuring the target shape.
[0283] In a specific embodiment of the present invention, step 130 includes:
[0284] Reflectivity decomposition is used to decompose the source modality image of each modality into two physical components: ;
[0285] in, It is the source modal image, Indicates the lighting components, Represents the reflectivity component;
[0286] The illumination components are solved using the Tikhonov regularization method. :
[0287]
[0288] Among them, data fidelity item To ensure that the reconstruction error of the decomposed image is minimized. This is the initial estimate of the reflectivity component;
[0289] Smoothing regularization terms The lighting components are required to vary smoothly in space (in accordance with the characteristics of light).
[0290] It is a regularization parameter that controls the degree of smoothness;
[0291] Using the obtained illumination components Calculate reflectivity
[0292] Using the reflectivity Noise reliability weight and cross-modal consistency value The fusion weights are obtained as follows:
[0293]
[0294] in, , and Control the contribution level of each factor separately;
[0295] Thus, a reflectivity domain fused image is obtained. :
[0296]
[0297] Obtain the final reconstructed image :
[0298]
[0299] in, It is the fused reflectance image, representing the weighted fusion result of all modal reflectance information at pixel position x;
[0300] It uses the illumination component of the reference mode. The final multimodal fused image represents the optimized fusion result generated after the complete processing flow at pixel position x;
[0301] Will The data is directly fed into the task network, and the object-level task results are output. .
[0302] In step 140, alignment, fusion, and task objectives are optimized under a single loss function:
[0303]
[0304] Among them, geometric alignment loss To ensure geometric consistency, merge consistency loss Ensuring modal consistency results in a performance penalty for the task. Ensure downstream target recognition performance.
[0305] Radiative symmetry loss Ensure the accuracy of the radiation correction process.
[0306] It is used to compensate for quality loss and to evaluate and optimize the effectiveness of void filling.
[0307] This unified framework avoids the error accumulation of traditional pipelined methods, achieving an end-to-end optimal solution. The cross-modal consistency term is as follows:
[0308]
[0309] Representative loss of TRL task:
[0310]
[0311] Finally, the results are output and deployed collaboratively between the edge and cloud. The final output includes: 1) the fused high-precision image. 2) Object-level task results In terms of deployment, this invention supports edge-cloud collaboration: edge devices are responsible for front-end ERS correction, scale normalization, and preliminary registration, while the cloud platform completes task perception fusion and deep optimization, thus balancing real-time performance and accuracy.
[0312] As shown in Table 1, the embodiments of this application solve the following technical problems:
[0313] Key Point 1: Scale Normalization and Target Preservation Mechanism for Images of Different Sizes
[0314] This application proposes a scale normalization method for processing multimodal images of varying sizes. This method not only unifies the resolution but also introduces target edge preservation and detail protection mechanisms. It avoids the problem of target detail loss caused by relying on simple interpolation or cropping in existing technologies. By adding geometric consistency constraints during the normalization process, it ensures that images of different resolutions maintain object-level proportions and structural integrity during alignment.
[0315] The technical effect is that this mechanism can effectively reduce the risk of small and slender targets being blurred or distorted during the fusion process, ensuring the consistency of the spatial structure and the fidelity of details in the fused image, thereby improving the accuracy of subsequent detection and recognition.
[0316] Key Point 2: Multimodal Cross-Sensor Feature Alignment and Fusion
[0317] This application addresses the significant differences in imaging mechanisms across different modalities, such as visible light, infrared, and depth imaging, by proposing a cross-modal feature alignment method. Through radiometric correction and feature space mapping, images from different modalities are transformed into a unified representation space, where feature alignment and information fusion are performed. This method fully considers the differences in spectral density, radiometric intensity, and signal-to-noise ratio among different modalities, ensuring that fusion relies not only on pixel-level matching but also incorporates cross-modal feature consistency constraints.
[0318] The technical effect is that information from different modalities can achieve higher precision alignment and complementarity in the target region. The fusion result is characterized at the object level as having clear target boundaries and consistent semantics, avoiding the modality mismatch and information loss problems common in traditional methods.
[0319] Key Point 3: Robustness Enhancement Mechanisms in Irregular Illumination and Complex Environments
[0320] This application proposes a robust enhancement method that combines radiation symmetry correction with environment-adaptive weights. By uniformly correcting the differences in illumination and radiation characteristics of different modes, and dynamically adjusting the weights for overexposure, underexposure, shadows, haze, and uneven thermal radiation, the method can suppress the influence of abnormal regions at the pixel level and highlight reliable information in the target area.
[0321] The technical effect is that the mechanism significantly improves the fusion stability in complex environments, enabling the target area to maintain clear boundaries and consistent semantics even in strong light, backlight, or low visibility scenarios, effectively avoiding ghosting, information loss, or shape distortion in the fusion results.
[0322] Key Point 4: A multi-layered alignment framework combining global affine mapping and local geometric expansion
[0323] This application proposes a multi-level alignment framework that combines global affine transformation with local geometric extension. Global affine transformation is used to maintain the consistency of the overall spatial structure, while local geometric extension is used for fine-grained morphological adjustments within the target region. This method combines global stability with local flexibility through a hierarchical alignment strategy, thereby avoiding the problem of insufficient adaptation of a single affine model to local targets.
[0324] The technical effect is that the framework can simultaneously ensure the stability of the overall structure and the local accurate alignment of the target area, avoiding boundary mismatch or geometric deformation of object-level targets due to global model limitations, and significantly improving the geometric consistency of the fused image.
[0325] Key Point 5: Geometric Expansion Modeling and Minimum Tension Constraints of the Target Region
[0326] This application introduces a geometric expansion mechanism in local geometric modeling and combines minimum stretching and area conservation constraints to limit the scale changes of the target region. This method ensures that the target is not overstretched or compressed during local alignment, and at the same time, it constrains the target boundary and structural morphology during model optimization to guarantee the stability of the geometric shape.
[0327] The technical effect is that the target region can maintain its original shape and proportion in the fusion result. Especially for slender targets or targets with strong geometric regularity, it can effectively avoid shape distortion and improve the structural integrity and recognition accuracy of object-level targets.
[0328] Key Point 6: Combining Affine Global Constraints with Local Compensation Mechanisms
[0329] This application proposes a scheme combining global affine constraints and local compensation mechanisms based on traditional affine models. By maintaining overall affine consistency while introducing nonlinear compensation and fine-tuning mechanisms in local regions, the target and background can still be synchronized even in complex scenes with different modalities. This method overcomes the shortcomings of single affine models under conditions of multiple viewpoints, different focal lengths, and distortions.
[0330] The technical benefits are as follows: This mechanism ensures the stability of the global structure while providing the ability to flexibly adapt to local targets, enabling the fusion results to achieve high-precision alignment at the object level under different shooting conditions, thus improving the applicability of multimodal image fusion in complex practical applications. (See Table 1.)
[0331]
[0332] Table 1. Comparison between the corresponding solution of the present invention and the prior art
[0333] Symbol conventions:
[0334] Modal set The reference mode is denoted as .
[0335] Source Image (Grayscale or multi-channel), reference image .
[0336] Deformation mapping Reference domain coordinates Mapping to mode The source coordinates.
[0337] Sampling uses a differentiable bilinear operator .
[0338] Conventional operators such as structure tensor and gradient are defined according to the standard.
[0339] like Figure 2 As shown, this invention addresses multimodal scenarios of varying sizes (visible light VIS / infrared IR / depth DEP, etc.) and aligns and fuses content at the "object level (target level)". The core features include: radial symmetry (ERS) and scale normalization to establish cross-modal comparability at the pixel intensity and resolution levels; multi-level registration using global affine + local geometric extension (NUAE) to explicitly constrain minimum stretching / area conservation, ensuring target geometry; and task-aware fusion domain decision-making and minimum stretching completion to preserve object-level contours.
[0340] Reflectivity domain reliability-gated fusion suppresses the impact of irregular lighting / environment; unified optimization objectives (including alignment / fusion / task-related losses) and edge-cloud collaborative deployment.
[0341] The methods of the present invention will be described in detail below with reference to specific embodiments:
[0342] In step S1, multimodal data acquisition and preliminary correction are performed.
[0343] First, acquire a collection of images of varying sizes from different sensors. ,in This indicates the specific modal type, such as visible light (VIS), infrared (IR), depth (DEP), etc. The source modality image to be corrected. This refers to the total mode type, with each mode corresponding to different resolutions, spectral characteristics, and signal-to-noise levels. If the system has extrinsic parameters or a synchronization mechanism, preliminary temporal and geometric corrections are performed on the multimodal data to reduce the burden of subsequent alignment.
[0344] In step S2, radiation symmetry (ERS) correction is performed.
[0345] In multimodal image processing, due to differences in imaging principles, different sensors exhibit significant inconsistencies in the radiometric characteristics of the same scene across different modal images. This inconsistency is mainly reflected in: differences in brightness levels (the global brightness distribution differs across modal images); differences in contrast range (significant differences exist in the dynamic response range of each modality); and complex local response characteristics (the pixel intensity relationships of the same object are complex across different modalities).
[0346] By using logarithmic domain transformation and learning of monotonic mapping functions, the statistical characteristics of each modality image are aligned with the reference modality, eliminating radiometric inconsistencies while preserving the unique information content of each modality.
[0347] First, convert each modality image to the logarithmic domain: , ,in, The source modality image to be corrected. For reference modal image, It is a tiny positive number, used to prevent the value from becoming unstable when the logarithm of zero is taken.
[0348] Learn monotonically non-decreasing mapping functions through optimization methods. The optimization objective is... ,expression:
[0349]
[0350] in, It is a data fidelity term that ensures that the transformed source modal image is as close as possible to the reference modal image in the logarithmic domain. It is a smoothing constraint term: to prevent the mapping function from fluctuating excessively and to maintain the natural smoothness of the transformation. It is the regularization coefficient, which balances the weights of the two terms. To prevent unstable logarithmic values.
[0351] Mapping function From the family of monotonically non-decreasing functions The learning process employs the equal-order spline technique to ensure the monotonicity of the function while maintaining its differentiability. It supports gradient propagation and end-to-end training, ensuring the smoothness and stability of the mapping relationship.
[0352] The learned mapping function is applied to the source modality image, and the output image is radially symmetric. :
[0353]
[0354] This operation ensures the comparability of cross-modal images in terms of brightness, contrast, and local dynamic range, thereby improving the robustness of subsequent registration.
[0355] In step S3, scale normalization is performed.
[0356] In multimodal image processing, images acquired by different sensors often have different resolutions. This scale difference can lead to: structural misalignment: the same object appears at different sizes in different modal images; registration difficulties: scale mismatch increases the complexity of geometric registration; and fusion distortion: directly fusing images of different scales can cause target deformation.
[0357] This application embodiment analyzes the multi-scale structural information of the image and automatically finds the optimal scaling ratio, so that the scaled source modal image is consistent with the reference modal in terms of global structural characteristics.
[0358] Set core optimization goals:
[0359]
[0360] in Indicates the first A layered Laplacian energy operator is used to construct a multi-scale pyramid representation of the image, with each layer capturing structural information at different scales. The energy operator quantifies the strength of structural features at each scale.
[0361] Down is a differentiable scaling mode that scales the source modal image by a ratio s. It is implemented using differentiability, supports gradient propagation, and facilitates automatic parameter tuning during the optimization process.
[0362] By comparing the energy differences between the scaled image and the reference image at each pyramid level, the optimal scaling ratio is found by minimizing the total energy difference, ensuring the consistency of multi-scale structural characteristics.
[0363] Unlike traditional interpolation, this invention introduces edge preservation and geometric consistency constraints during the normalization process to prevent the target proportions from being destroyed.
[0364] In step S4, global affine initialization is performed.
[0365] Even after scale normalization, there are still overall geometric differences between images of different modalities, including: translational offset: image positions do not coincide; rotational differences: inconsistent orientation angles; scaling residue: subtle size differences after scale normalization; affine deformation: slight stretching or shearing deformation.
[0366] This application optimizes the affine transformation parameters to achieve the best matching of gradient features between the transformed source modal image and the reference modal image, while using a robust loss function to enhance the adaptability to illumination changes.
[0367] First, a global affine model is used to align the overall structure and narrow the search domain.
[0368]
[0369] in, It is a 2×2 linear transformation matrix that controls rotation, scaling, and shearing. It is a 2D translation vector that controls the position offset. These are the pixel coordinates in the source modal image.
[0370] in Robust optimization of the objective function:
[0371]
[0372] in, It is gradient feature matching, used to compare the gradient information of the reference image and the transformed source image. It uses the L1 norm to enhance robustness to outliers, and gradient features are relatively insensitive to changes in illumination.
[0373] It is a robust loss function, either Huber or Cauchy, to enhance robustness under varying illumination. It achieves initial geometric consistency. This step ensures the rigid constraints of the global framework, providing a reference benchmark for subsequent local optimization. The optimization scope is within the reference image domain. Summation is performed on all pixels within the range.
[0374] In step S5, the local block non-uniform geometry is expanded.
[0375] Even with global affine transformation, local geometric differences still exist. These differences stem from: non-rigid deformation: local deformation caused by differences in the viewpoints of different modal sensors; target specificity: differences in the performance of different objects in different modalities; and resolution inconsistency: despite scale normalization, there are still matching deviations in local details.
[0376] While maintaining the overall geometric framework, this application allows for moderate linear adjustments to each local region, achieving multi-level geometric alignment from coarse to fine.
[0377] First, the reference image domain Divided into K grids / superpixels Based on the global affine transformation, a local linear extension is superimposed:
[0378]
[0379] in, It is the global foundation, namely the global affine transformation in the previous step.
[0380] It is a local expansion. It is a partition indicator. It is the first Local linear transformation matrix of a local region on mode m It is the local translation vector of the k-th local region on mode m.
[0381] Constructing heterogeneous uniform energy (HSCE):
[0382]
[0383] Its optimization objective function consists of three parts:
[0384] Alignment term: Ensures consistency between gradient and structure tensor; Minimum stretch constraint: To avoid excessive deformation of the target; smoothing regularization: neighborhood Maintain continuity between blocks to avoid tearing. Denotes the set of adjacent partition pairs, where: and It is the partition index number N is the neighborhood relation set, which contains all adjacent partition pairs. It is an unordered pair, representing a partition. and partitions Adjacent. It is the first In the mode, the th Local linear transformation matrix of a local region. It is the first In the mode, the th Local linear transformation matrix of a local region. It is the first A local region in mode Local translation vector on. It is a weighted hyperparameter used to balance the importance of the gradient consistency term in the overall alignment loss function. It controls the contribution of the gradient matching term: Gradient term: Gradient information is crucial for the structural alignment of images, measuring the consistency of edge and contour features between the transformed image and the reference image. It is a weighted hyperparameter used to balance the importance of the structure tensor consistency term in the overall alignment loss function.
[0385] The NUAE mechanism enables multi-level alignment of global and local collaboration, ensuring that object-level targets are not distorted while flexibly adapting to local differences.
[0386] For structural tensors or directional consistency measures;
[0387] Reliability weight
[0388] The optimization strategy employs alternating minimization under a multi-scale pyramid:
[0389] fixed optimization (Each quadratic subproblem can be parallelized);
[0390] Fixed local terms iterative update (Gauss-Newton);
[0391] Refinement layer by layer, convergence criterion is .
[0392] In step S6, the optimal fusion domain decision for task awareness is made.
[0393] After completing the geometric alignment of multimodal images, it is necessary to determine the optimal fusion region. Traditional methods only consider geometric consistency, while this application selects the fusion region to simultaneously satisfy geometric alignment quality and downstream task requirements, achieving task-driven intelligent fusion domain decision-making.
[0394] The optimal fusion domain is determined by minimizing the following energy function:
[0395]
[0396] in, It is the geometric deformation cost, used to measure the cumulative geometric alignment error within the region Ω. It refers to task coverage, which is used to ensure that the merged area covers the key content related to the task. This is the average of the task saliency maps within region Ω. The task saliency maps are generated by a pre-trained task network. generate. This refers to boundary discontinuity, used to penalize the unnaturalness of the blended boundary. It measures the length and curvature of the blended domain boundary; considers the alignment of the boundary with the natural edges of the image; and avoids producing jagged or abrupt boundaries. It is a context-extended reward, used to incentivize the inclusion of task-related contextual information. It controls the tolerance for geometric deformation. It controls the balance between task performance and the compactness of the fusion region. Controlling the naturalness of the fusion domain boundaries, It adjusts the degree to which contextual information is included.
[0397] Fusion not only needs to satisfy geometric consistency, but should also be task-oriented (such as detection and segmentation). Therefore, this invention introduces task-saliency-guided pruning-union optimization: utilizing task networks (Detection / Segmentation) Generate a saliency map on the alignment results. The fusion domain is determined by minimizing the clipping-union function:
[0398]
[0399] in radius The morphological dilation (which can be approximated as differentiable using soft morphology) is considered. This energy function simultaneously takes into account geometric distortion, task coverage, boundary continuity, and context extension, ensuring the final fusion domain... It is optimal for target detection tasks.
[0400] In step S7, occlusion / hole detection and minimum stretching completion are performed.
[0401] In the process of multimodal image registration, due to differences in sensor viewpoints, occlusion phenomena, or registration errors, the following problems often occur: some objects are visible in one mode but are occluded in another mode, resulting in areas without corresponding pixels after registration; interpolation errors or boundary effects during the transformation process cause some pixels to fail to find effective corresponding points in the inverse transformation.
[0402] Testing conditions:
[0403] For each pixel position x, a hole is determined if any of the following conditions are met:
[0404] Condition 1:
[0405] Condition 2: No effective support for cross-modal communication
[0406] Among them, in condition 1 These are the pixel coordinates in the source modal image. express . This is the positional deviation threshold. It indicates that if the transformation is irreversible or there is a large error, the registration at that position is unreliable.
[0407] Condition 2, cross-modal validity, means checking whether there are valid corresponding pixels in the source modality, excluding invalid values, boundary overflow, sensor blind spots, etc., to ensure that each pixel has reliable cross-modal support.
[0408] The set of hole regions detected during multimodal image registration Perform minimum stretching completion. Set the optimization objective function and constraints:
[0409]
[0410] in, It is a position correction vector, which defines the mapping relationship from the initial registration position to the optimized position.
[0411] Smoothing terms The requirement is to make the entire surface as smooth as possible within the void areas. It is the gradient of the entire field;
[0412] It ensures that the completed area is smoothly connected to the surrounding known area. It is the boundary of the hollow area; These are known pixel values on the boundary (from the surrounding reliable region). It is the boundary matching weight coefficient, which controls the connection strength.
[0413] It is an area conservation constraint that avoids excessive stretching or compression during the completion process.
[0414] This is a soft constraint threshold, equivalent to area conservation completion guided by ARAP / Poisson. This constraint keeps the local area unchanged, thus avoiding excessive expansion or contraction during completion and ensuring the target shape.
[0415] In step S8, the reliability of the reflectivity domain is gated and fused.
[0416] In multimodal image fusion, complex lighting conditions severely affect the fusion quality. Traditional methods perform fusion directly in the original image domain, which is easily affected by changes in lighting.
[0417] This application achieves intelligent weighting by decomposing reflectance and illumination, fusing data in the reflectance domain where illumination remains constant, and combining this with reliability assessment to fundamentally eliminate the influence of illumination and improve the fusion quality.
[0418] Each modality's image is decomposed into two physical components using reflectance decomposition: ,in, It is the source modal image after the previous steps. Indicates the lighting components, This represents the reflectivity component.
[0419] The illumination components were estimated using the Tikhonov regularization method:
[0420]
[0421] Among them, data fidelity item This ensures that the reconstruction error of the decomposed image is minimized; This is the initial estimate of the reflectance component. Smoothing regularization term. The lighting components are required to vary smoothly in space (in accordance with the characteristics of light). This is a regularization parameter that controls the smoothness. (Reflectance calculation)
[0422] Fusion weights (reliability gating):
[0423]
[0424] in, It is the noise reliability weight. It is the reflectivity gradient. It is a consistency indicator. , , All of these are weights.
[0425] The reliability of different modes is quantified and then fused.
[0426]
[0427]
[0428] in, It is the fused reflectance image, representing the weighted fusion result of all modal reflectance information at pixel position x;
[0429] It uses the illumination component of the reference mode. The final multimodal fused image represents the optimized fusion result generated after the complete processing flow at pixel position x.
[0430] Can Directly input the task network to output object-level task results This mechanism suppresses the interference of low-confidence modes in the fusion process, achieving robust weighting of the target region.
[0431] In step S9, a unified optimization objective and training are performed.
[0432] This invention optimizes alignment, fusion, and task objectives under a single loss function:
[0433]
[0434] Among them, geometric alignment loss To ensure geometric consistency, merge consistency loss Ensuring modal consistency results in a performance penalty for the task. Ensure downstream target recognition performance.
[0435] Radiative symmetry loss Ensure the accuracy of the radiation correction process. It is used to compensate for quality loss and to evaluate and optimize the effectiveness of void filling. Number of source modes (e.g., visible light, infrared, SAR, etc.); Each loss weight (hyperparameter) can be fixed or use segmented / adaptive weights (such as GradNorm or uncertainty weighting).
[0436] This unified framework avoids the error accumulation of traditional pipelined methods, achieving an end-to-end optimal solution. The cross-modal consistency term is as follows:
[0437]
[0438] in, For spatial gradient operators, For the first The modal response after radiation symmetry and geometric registration For the fusion result, For pixel-wise modal weights and satisfying , Provides robustness against noise; This is a task representativeness regularization used to enhance the structural and discriminative features related to downstream targets. It can be implemented as an edge-preserving term focusing on candidate target regions or a consistency term in the task feature space. Its weight.
[0439] Representative loss of TRL task:
[0440]
[0441] In step S10, the results are output and deployed collaboratively between the edge and cloud.
[0442] The final output includes: 1) the fused high-precision image 2) Object-level task results In terms of deployment, this invention supports edge-cloud collaboration: edge devices are responsible for front-end ERS correction, scale normalization, and preliminary registration, while the cloud platform completes task perception fusion and deep optimization, thus balancing real-time performance and accuracy.
[0443] like Figure 3 As shown in the embodiments of this application, a multimodal image fusion system based on multi-level alignment and task awareness is also provided. It includes: an input and synchronization module; an ERS radiation symmetry module; a scale normalization module; a global affine estimation module; a NUAE local extended registration module; a task saliency and fusion domain decision module; an occlusion detection and minimum stretching completion module; a reflectivity domain gated fusion module; and a task inference and result output module.
[0444] The system can be implemented on CPU+GPU / FPGA / ASIC; morphological "expansion", differentiable sampling, graph Laplacian regularization and other operators are provided for implementation.
[0445] Readable storage medium: The program instructions for implementing steps S1–S10 above are stored in a non-transitory computer-readable medium (such as flash memory / solid-state drive / EEPROM) and loaded and executed by the processor to implement the method; if firmware is involved, it is located in the microcode area of the vision accelerator and is used to accelerate S, gradient, structural tensor and morphological operators.
[0446] Edge side: S1–S5 (including ERS / scale / affine + NUAE alignment) and initial screening for task saliency, outputting sparse deformation parameters. With low bit reflectivity characteristics;
[0447] Cloud-based: S6–S9 (Domain decision refinement, completion, fusion, and task reasoning), returning object-level results and sending back adaptive parameters;
[0448] Communication load: Transmit parameters instead of the original multimodal video stream, reducing bandwidth and latency.
[0449] Furthermore, in this embodiment of the application, block-based linear expansion is performed. Replace with B-spline / TPS free deformation parameters However, the area conservation / minimum stretching term is retained: Adjacent smoothing is replaced with graph Laplacian regularization of surface control points. — Equivalent coverage of the non-rigid body case.
[0450] The clipping-union objective can be replaced with graph cut / energy minimization: for pixel / superpixel binary labels (in-domain / out-domain), data terms are... Compared to geometric distortion, the smoothness term, using cross-modal boundary inconsistency, still retains... Context extension.
[0451] In this embodiment, in addition to the reflectivity domain, gated fusion can be performed in the gradient / Laplace domain, and then Poisson reconstruction can be used to obtain the final product. Reliability gating weights Keep it unchanged (or add directional consistency) .
[0452] In the embodiments of this application, if there is depth / geometric prior, the occlusion mask can be inferred from depth consistency and Z-buffer; if the depth is missing, the joint threshold of forward and reverse consistency and photometric residual is used.
[0453] In this embodiment of the application, the monotonic mapping of ERS It can be replaced by piecewise linear histogram matching or monotonic neural mapping (Isotonic Net); as long as it satisfies monotonicity and smoothness regularity, it is within the scope of this invention.
[0454] In this embodiment of the application, if there is no task label, Replace it with cross-modal consistency + temporal consistency; or promote it with pseudo-label self-training, while still maintaining a unified loss framework.
[0455] In this embodiment, each layer and block of NUAE is a small-scale quadratic subproblem that can be parallelized (GPU / FPGA block-wise).
[0456] The morphological "expansion" uses a soft-dilation approximation (smooth substitution of max pooling) to ensure differentiability and hardware friendliness;
[0457] The area conservation and minimum stretching terms only introduce a quadratic penalty for the diagonal Jacobian, which facilitates Gaussian-Newton acceleration.
[0458] The various metrics of NRW weights (saturation / blurring / SNR / edge consistency) are calculated quickly on the edge side using lightweight convolution and integral images;
[0459] Typical end-to-end configuration: 720p×2 modal, 30fps target, <5ms / layer on A100; <33ms / frame on Jetson Orin NX (NUAE 2 layers, K≈400).
[0460] Next, the method of the present invention will be described in detail with reference to another specific embodiment. For example... Figure 4 As shown, a multimodal image fusion device based on multi-level alignment and task awareness is disclosed, the device comprising:
[0461] The first processing module is used to acquire source modal images of different sizes of the target area from different sensors, perform radial symmetry correction on each of the source modal images, and perform scale normalization on each of the radially symmetric source modal images.
[0462] The second processing module is used to perform global affine initialization on each of the source modal images after scale normalization, perform local geometric expansion, and perform region fusion on each of the source modal images after local geometric expansion based on task saliency and geometric constraints to obtain a fused source modal image after region fusion.
[0463] The third processing module is used to perform occlusion / hole detection and completion on the fused source modal image, correct the completed fused source modal image in the reflectivity domain, input the corrected fused source modal image into the task network for processing, and output the object-level task result.
[0464] The fourth processing module is used to optimize the fused source modal image after region fusion, the object-level task result, each source modal image after local geometric expansion, the fused source modal image after completion, and each source modal image after radial symmetry using a loss function, so as to obtain the final fused source modal image and the object and task result.
[0465] In a specific embodiment of the present invention, the first processing module includes:
[0466] Transform each of the source modal images to the logarithmic domain. ,in, For source modal images, For reference modal image, It is a modal set. It is a modal type. It is a tiny positive number, used to prevent the value from becoming unstable when the logarithm of zero is taken.
[0467] The learned mapping function is applied to the source modality image to obtain the radiation-symmetric source modality image. The optimization objective is to minimize the difference between the transformed sample source modal image and the reference sample modal image in the logarithmic domain, thereby optimizing the mapping function to obtain the optimized mapping function.
[0468] In a specific embodiment of the present invention, the first processing module includes:
[0469] Set normalization optimization target :
[0470]
[0471] in Indicates the first Layered Laplacian energy operators are used to construct multi-scale pyramid representations of images, with each pyramid layer capturing structural information at different scales. It is the number of layers in the Laplacian capability operator, and Down is a differentiable scaling operation used to scalate the source modality image after radiation symmetry. Scale by a ratio s;
[0472] Calculate the scaled source modal image Compared with reference image Multiscale differences: in the first Energy characteristics on the layered Laplace pyramid This creates a difference. and its square norm 2 As intra-layer differences; accumulated across all layers. As an optimization target;
[0473] Using optimization algorithms The search space is used to find the optimal scaling factor that minimizes the total L2 norm difference.
[0474] After using the optimal scaling ratio, the source modal images are subjected to the corresponding scaling transformation, and the scale-normalized source modal images are output.
[0475] In a specific embodiment of the present invention, the second processing module includes:
[0476] A global affine model is used to establish a globally aligned overall structure. Global parameters , These are the pixel coordinates in each of the source modal images after scale normalization;
[0477] By utilizing a robust optimization objective function, the reference image is made possible. Gradient and transformed source image The minimum value of the gradient difference is used to obtain the optimal global affine transformation parameters. and ;
[0478] The robust optimization objective function is:
[0479]
[0480] in, It is gradient feature matching, used to compare the gradient information of the reference image and the transformed source image. It is a robust loss function, referring to the image domain. All pixels within, The reference image ;
[0481] Utilizing the optimized global affine transformation parameters and The global affine result is obtained. ;
[0482] Reference image domain Divided into K grids / superpixels A local linear extension is superimposed on the global affine result:
[0483]
[0484] in, It is a global affine result;
[0485] It is a local expansion. It is a partition indicator, a local parameter. It is the local linear transformation matrix of the k-th local region on mode m, and the local parameters are... It is the local translation vector of the k-th local region in mode m;
[0486] Constructing heterogeneous size consistent energy targets The heterogeneous size uniform energy target This includes gradient consistency terms and structure tensor consistency terms, minimum stretching constraint terms, and smoothing regularization terms;
[0487]
[0488] Among them, the gradient consistency term and the structure tensor consistency term are:
[0489]
[0490] Minimum tensile constraint term:
[0491] Smoothing regularization terms:
[0492] in, Represents the set of adjacent partition pairs. and It is the partition index number ( ∈ K is the number of partitions. It is a set of neighborhood relationships, containing all adjacent partition pairs; It is an unordered pair, representing a partition. and partitions Adjacent; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first Each partition in the model Local translation vectors on each, It is a weighted hyperparameter used to balance the importance of the gradient consistency term in the overall alignment loss function;
[0493] A hierarchical alternating optimization framework is adopted, with fixed global affine transformation parameters. and Parallel optimization of local parameters and Fix local parameters and iteratively update global parameters. and When the preset convergence condition is met: the change in the energy function The global affine transformation parameters are obtained. and and the local parameters and ;
[0494] This results in the structure after local geometric expansion:
[0495] .
[0496] In a specific embodiment of the present invention, the second processing module includes:
[0497] By minimizing the comprehensive energy function, the source modal images are fused to obtain the optimal fusion domain:
[0498]
[0499] in, It is the geometric deformation cost, used to measure the cumulative geometric alignment error within the region Ω. Reference image coordinate system A region or subset thereof; This refers to task coverage, which ensures that the fusion area covers key content related to the task. It is the average of the task saliency map within region Ω; the task saliency map is generated by the pre-trained task network. generate; It refers to boundary discontinuity, used to penalize the unnaturalness of the merged boundary;
[0500] It is a context-expansion reward, used to incentivize the inclusion of task-related contextual information. It is the radius of the expansion operation;
[0501] It is a weight that controls the tolerance for geometric deformation. It is a weight that controls the balance between task performance and the compactness of the fusion region. Weights that control the naturalness of the fusion domain boundaries. It is a weight that adjusts the degree of inclusion of contextual information.
[0502] In a specific embodiment of the present invention, the third processing module includes:
[0503] Set detection conditions to obtain a set of void regions. ;
[0504] For each pixel position x, a hole is determined if any of the following conditions are met:
[0505] Condition 1: ;
[0506] Condition 2: No effective support for cross-modal communication;
[0507] Among them, in condition 1 These are the pixel coordinates in the source modal image. express ; It is the position deviation threshold, which indicates that if the transformation is irreversible or there is a large error, the registration at that position is unreliable.
[0508] Condition 2, cross-modal validity, means checking whether there are valid corresponding pixels in the source modality, excluding invalid values, boundary overflow, sensor blind spots, etc., to ensure that each pixel has reliable cross-modal support;
[0509] By optimizing the objective function and constraints, the set of hole regions detected during multimodal image registration is optimized. Perform minimum stretching completion;
[0510] The objective function and constraints are as follows:
[0511]
[0512]
[0513] in, It is a position correction vector, which defines the mapping relationship from the initial registration position to the optimized position;
[0514] Smoothing terms The requirement is to make the entire surface as smooth as possible within the void areas. It is the gradient of the entire field;
[0515] It ensures a smooth transition between the completed area and the surrounding known area; It is the boundary of the hollow area; These are the known pixel values on the boundary. These are boundary matching weight coefficients that control the strength of the connection.
[0516] It is an area conservation constraint that avoids excessive stretching or compression during the completion process; The soft constraint threshold is equivalent to area conservation completion guided by ARAP / Poisson. This constraint keeps the local area unchanged, thereby avoiding excessive expansion or contraction during completion and ensuring the target shape.
[0517] In a specific embodiment of the present invention, the third processing module includes:
[0518] Reflectivity decomposition is used to decompose the source modality image of each modality into two physical components: ;
[0519] in, It is the source modal image, Indicates the lighting components, Represents the reflectivity component;
[0520] The illumination components are solved using the Tikhonov regularization method. :
[0521]
[0522] Among them, data fidelity item To ensure that the reconstruction error of the decomposed image is minimized. This is the initial estimate of the reflectivity component;
[0523] Smoothing regularization terms The lighting components are required to vary smoothly in space (in accordance with the characteristics of light).
[0524] It is a regularization parameter that controls the degree of smoothness;
[0525] Using the obtained illumination components Calculate reflectivity
[0526] Using the reflectivity Noise reliability weight and cross-modal consistency value The fusion weights are obtained as follows:
[0527]
[0528] in, , and Control the contribution level of each factor separately;
[0529] Thus, a reflectivity domain fused image is obtained. :
[0530]
[0531] Obtain the final reconstructed image :
[0532]
[0533] in, It is the fused reflectance image, representing the weighted fusion result of all modal reflectance information at pixel position x;
[0534] It uses the illumination component of the reference mode. The final multimodal fused image represents the optimized fusion result generated after the complete processing flow at pixel position x;
[0535] Will The data is directly fed into the task network, and the object-level task results are output. .
[0536] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a multimodal image fusion method based on multi-level alignment and task awareness.
[0537] This invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of a multimodal image fusion method based on multi-level alignment and task awareness as described above.
[0538] In addition, combined Figure 1 The multimodal image fusion method based on multi-level alignment and task awareness described in this embodiment of the invention can be implemented by an electronic device, such as a computer device. Figure 5 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention.
[0539] In some embodiments, the computer device may further include a communication interface 83 and a bus 80. For example, Figure 5 As shown, the processor 81, memory 82, and communication interface 83 are connected through bus 80 and complete communication with each other.
[0540] Specifically, the processor 81 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of the present invention.
[0541] The memory 82 can be used to store or cache various data files that need to be processed and / or communicated, as well as possible computer program instructions executed by the processor 81.
[0542] The processor 81 reads and executes computer program instructions stored in the memory 82 to implement any of the multimodal image fusion methods based on multi-level alignment and task awareness in the above embodiments.
[0543] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0544] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A multimodal image fusion method based on multi-level alignment and task awareness, characterized in that, The method includes: Acquire source modal images of different sizes of the target region from different sensors, perform radial symmetry correction on each source modal image, and perform scale normalization on each source modal image after radial symmetry correction. After global affine initialization of each source modality image after scale normalization, local geometric expansion is performed. Based on task saliency and geometric constraints, region fusion is performed on each source modality image after local geometric expansion to obtain a fused source modality image after region fusion. The fused source modal image is subjected to occlusion / hole detection and completion. The completed fused source modal image is corrected in the reflectivity domain. The corrected fused source modal image is input into the task network for processing, and the object-level task result is output. Using a loss function, the fused source modal image after region fusion, the object-level task result, each source modal image after local geometric expansion, the fused source modal image after completion, and each source modal image after radial symmetry are optimized to obtain the final fused source modal image and the object and task result.
2. The method according to claim 1, characterized in that, The radial symmetry correction performed on each of the source modal images includes: Transform each of the source modal images to the logarithmic domain. ,in, For source modal images, For reference modal image, It is a modal set. It is a modal type. It is a tiny positive number, used to prevent the value from becoming unstable when the logarithm of zero is taken. The learned mapping function is applied to the source modality image to obtain the radiation-symmetric source modality image. The optimization objective is to minimize the difference between the transformed sample source modal image and the reference sample modal image in the logarithmic domain, thereby optimizing the mapping function to obtain the optimized mapping function.
3. The method according to claim 2, characterized in that, The scaling normalization of each of the source modal images after radiation symmetry includes: Set normalization optimization target : in Indicates the first Layered Laplacian energy operators are used to construct multi-scale pyramid representations of images, with each pyramid layer capturing structural information at different scales. It is the number of layers in the Laplacian capability operator, and Down is a differentiable scaling operation used to scalate the source modality image after radiation symmetry. Scaling by a ratio s This is a reference image; Calculate the scaled source modal image Compared with reference image Multiscale differences: In the Energy characteristics on the layered Laplace pyramid , forming a difference and its square norm 2 As intra-layer differences; accumulated across all layers. As an optimization target; Using optimization algorithms The search space is used to find the optimal scaling factor that minimizes the total L2 norm difference. Using the optimal scaling ratio, the source modal images are subjected to corresponding scaling transformations, and scale-normalized source modal images are output.
4. The method according to claim 3, characterized in that, After performing global affine initialization on each of the scale-normalized source modal images, local geometric expansion is then performed, including: A global affine model is used to establish a globally aligned overall structure. Global parameters , These are the pixel coordinates in each of the source modal images after scale normalization; By utilizing a robust optimization objective function, the reference image is made possible. Gradient and transformed source image The minimum value of the gradient difference is used to obtain the optimal global affine transformation parameters. and ; The robust optimization objective function is: in, It is gradient feature matching, used to compare the gradient information of the reference image and the transformed source image. It is a robust loss function, referring to the image domain. All pixels within, The reference image The image domain; Utilizing the optimized global affine transformation parameters and The global affine result is obtained. ; Reference image domain Divided into individual grids / superpixels A local linear extension is superimposed on the global affine result: in, It is a global affine result; It is a local expansion. It is a partition indicator, a local parameter. It is the local linear transformation matrix of the k-th local region on mode m, and the local parameters are... It is the local translation vector of the k-th local region in mode m; Constructing heterogeneous size consistent energy targets The heterogeneous size uniform energy target This includes gradient consistency terms and structure tensor consistency terms, minimum stretching constraint terms, and smoothing regularization terms; Among them, the gradient consistency term and the structure tensor consistency term are: Minimum tensile constraint term: Smoothing regularization terms: in, Represents the set of adjacent partition pairs. and It is the partition index number ( ∈ ) It refers to the number of partitions. It is a set of neighborhood relationships, containing all adjacent partition pairs; It is an unordered pair, representing a partition. and partitions Adjacent; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first In the mode, the th Local linear transformation matrices for each partition; It is the first Each partition in the model Local translation vectors on each, , , and It is a weight hyperparameter; A hierarchical alternating optimization framework is adopted, with fixed global affine transformation parameters. and Parallel optimization of local parameters and Fix local parameters and iteratively update global parameters. and When the preset convergence condition is met: the change in the energy function The global affine transformation parameters are obtained. and and the local parameters and ; This results in the structure after local geometric expansion: 。 5. The method according to claim 4, characterized in that, Based on task saliency and geometric constraints, the source modal images after local geometric expansion are fused into a fused source modal image, including: By minimizing the comprehensive energy function, the source modal images are fused to obtain the optimal fusion domain: in, It is the cost of geometric deformation, used to measure the region. Cumulative geometric alignment error within, Reference image coordinate system A region or subset thereof; This refers to task coverage, which ensures that the fusion area covers key content related to the task. It is the average of the task saliency maps within region Ω; the task saliency maps are generated from pre-trained... generate; It refers to boundary discontinuity, used to penalize the unnaturalness of the merged boundary; It is a context-expansion reward, used to incentivize the inclusion of task-related contextual information. It is the radius of the expansion operation; It is a weight that controls the tolerance for geometric deformation. It is a weight that controls the balance between task performance and the compactness of the fusion region. Weights that control the naturalness of the fusion domain boundaries. It is a weight that adjusts the degree of inclusion of contextual information.
6. The method according to claim 5, characterized in that, The step of detecting and completing occlusions / holes in the fused source modality image includes: Set detection conditions to obtain a set of void regions. ; For each pixel position x, a hole is determined if any of the following conditions are met: Condition 1: ; Condition 2: No effective support for cross-modal communication; Among them, in condition 1 These are the pixel coordinates in the source modal image. express ; It is the position deviation threshold, which indicates that if the transformation is irreversible or there is a large error, the registration at that position is unreliable. Condition 2, cross-modal validity, means checking whether there are valid corresponding pixels in the source modality, excluding invalid values, boundary overflow, sensor blind spots, etc., to ensure that each pixel has reliable cross-modal support; By optimizing the objective function and constraints, the set of hole regions detected during multimodal image registration is optimized. Perform minimum stretching completion; The objective function and constraints are as follows: in, It is a position correction vector, which defines the mapping relationship from the initial registration position to the optimized position; Smoothing Term The requirement is to make the entire surface as smooth as possible within the void areas. It is the gradient of the entire field; It ensures a smooth transition between the completed area and the surrounding known area; It is the boundary of the hollow area; These are the known pixel values on the boundary. These are boundary matching weight coefficients that control the strength of the connection. It is an area conservation constraint that avoids excessive stretching or compression during the completion process; The soft constraint threshold is equivalent to area conservation completion guided by ARAP / Poisson. This constraint keeps the local area unchanged, thereby avoiding excessive expansion or contraction during completion and ensuring the target shape.
7. The method according to claim 6, characterized in that, The process involves correcting the fused source modality image in the reflectivity domain, inputting the corrected fused source modality image into the task network for processing, and outputting object-level task results, including: Reflectivity decomposition is used to decompose the source modality image of each modality into two physical components: ; in, It is the source modal image, Indicates the lighting components, Represents the reflectivity component; The illumination components are solved using the Tikhonov regularization method. : Among them, data fidelity item To ensure that the reconstruction error of the decomposed image is minimized. This is the initial estimate of the reflectivity component; Smoothing regularization terms The lighting components are required to vary smoothly in space (in accordance with the characteristics of light). It is a regularization parameter that controls the degree of smoothness; Using the obtained illumination components Calculate reflectivity Using the reflectivity Noise reliability weight and cross-modal consistency value The fusion weights are obtained as follows: in, Control the contribution level of each factor separately; Thus, a reflectivity domain fused image is obtained. : Obtain the final reconstructed image : in, This is the fused reflectance image, representing the reflectance at pixel locations. The weighted fusion result of all modal reflectivity information; It uses the illumination component of the reference mode. The final multimodal fused image represents the pixel location. The optimized fusion result generated after a complete processing flow; Will The data is directly fed into the task network, and the object-level task results are output. .
8. A multimodal image fusion device based on multi-level alignment and task awareness, characterized in that, The device includes: The first processing module is used to acquire source modal images of different sizes of the target area from different sensors, perform radial symmetry correction on each of the source modal images, and perform scale normalization on each of the radially symmetric source modal images. The second processing module is used to perform global affine initialization on each of the source modal images after scale normalization, perform local geometric expansion, and perform region fusion on each of the source modal images after local geometric expansion based on task saliency and geometric constraints to obtain a fused source modal image after region fusion. The third processing module is used to perform occlusion / hole detection and completion on the fused source modal image, correct the completed fused source modal image in the reflectivity domain, input the corrected fused source modal image into the task network for processing, and output the object-level task result. The fourth processing module is used to optimize the fused source modal image after region fusion, the object-level task result, each source modal image after local geometric expansion, the fused source modal image after completion, and each source modal image after radial symmetry using a loss function, so as to obtain the final fused source modal image and the object and task result.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the multimodal image fusion method based on multi-level alignment and task awareness as described in any one of claims 1-7.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the multimodal image fusion method based on multi-level alignment and task awareness as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Bi-level optimization-based infrared and visible light fusion method
US11830222B2