An image multimodal fusion method and system for complex dynamic environments
By introducing a bidirectional adapter module and a multi-level feature-level fusion layer in the multimodal perception system, the semantic alignment and computing efficiency problems between modes are solved, and efficient fusion of infrared and visible images and more accurate perception effects are achieved.
Patent Information
- Application Number
- CN202411885090.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-20
AI Technical Summary
When existing multimodal perception technologies deal with multimodal data fusion in complex dynamic systems, there are problems with semantic alignment between modals, balance of computing efficiency and precision, and robustness.
A bidirectional adapter module and a multi-level feature-level fusion layer are designed. Multi-scale features are extracted through a multi-layer encoder, and the bidirectional adapter performs dynamic adaptation and semantic alignment to achieve efficient fusion of infrared images and visible light images.
It effectively solves the problems of semantic alignment and representation differences between modals, realizes more accurate perception of objects in complex environments, and provides more accurate and robust perception capabilities.
Smart Images

Figure CN119339201B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multimodal image processing, and in particular relates to an image multimodal fusion method and system for complex dynamic environments. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] In modern complex dynamic systems, such as smart cities, Industry 4.0, and intelligent transportation, real-time and accurate environmental perception is crucial for the efficient operation of the system. However, with the increasing complexity of the system and the rapid changes in the environment, traditional single-modal perception methods can no longer meet the needs of these complex scenarios. Multimodal perception is considered an important means to solve these problems because it can fuse data from different sources and provide more comprehensive information. However, multimodal perception also brings challenges such as data heterogeneity, information redundancy, and increased processing complexity.
[0004] At present, existing multimodal perception research mainly focuses on three directions: data-level, feature-level, and decision-level fusion. Data-level fusion directly processes the original data and is difficult to deal with the scale and representation differences between different modalities. Feature-level fusion fuses data in the feature space. Although it can partially solve the heterogeneity problem, it is difficult to capture the complex relationship between different modalities. Decision-level fusion is achieved by integrating the independent decision results of each modality, but this method tends to ignore the complementary information between modalities, thus affecting the perception effect. Existing perception methods have the following main problems when dealing with multimodal data fusion in complex dynamic systems:
[0005] (1) Semantic alignment between modalities: Data from different modalities may describe the same semantic information, but the representation forms are different. It is still difficult to effectively align this information during fusion.
[0006] (2) Balance between computational efficiency and accuracy: In resource-constrained scenarios, how to improve computational efficiency while ensuring perception accuracy is an important issue that multimodal perception systems need to address.
[0007] (3) Robustness issue: When certain modal data are unavailable or unreliable, how to ensure high performance output of the perception system is one of the key challenges of the system. Summary of the invention
[0008] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides an image multimodal fusion method and system for complex dynamic environments, innovatively designs a bidirectional adapter module and a multi-level feature-level fusion layer, effectively solves the problems of semantic alignment and representation differences between modalities, and can achieve efficient fusion of infrared images and visible light images. By combining the advantages of infrared images in low-light conditions with the high-resolution characteristics of visible light images, more accurate perception of objects in complex environments can be achieved.
[0009] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0010] The first aspect of the present invention provides an image multimodal fusion method for complex dynamic environments;
[0011] An image multimodal fusion method for complex dynamic environments, comprising:
[0012] Acquire a visible light image and an infrared image, and preprocess the visible light image and the infrared image;
[0013] The preprocessed visible light image and infrared image are respectively input into a multi-layer encoder to obtain a comprehensive feature map of the visible light image and a comprehensive feature map of the infrared image containing multi-scale information;
[0014] The comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image are input into the bidirectional adapter, and the multimodal features of the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image are respectively extracted, and the multimodal features are converted into two-dimensional features suitable for linear operations, and dynamic alignment of the two-dimensional features is realized through dynamic adaptive adjustment, and then the two-dimensional features adjusted by dynamic alignment are restored to the original dimension; the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image restored to the original dimension are semantically aligned;
[0015] The image gradient information of the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map is calculated, the fusion weight is determined based on the image gradient information, and the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map are fused using the fusion weight to obtain a fused image.
[0016] As a further technical solution, the preprocessing process includes: cutting the visible light image and the infrared image into image blocks of a specific size; and expanding the number of the image blocks by using data expansion technology.
[0017] As a further technical solution, the multi-layer encoder includes multiple encoders connected in sequence. Except for the first encoder, each subsequent encoder is respectively provided with a dense connection block, and all the previous encoders are connected through the dense connection block, so that the outputs of all the previous encoders are multiplexed.
[0018] As a further technical solution, the process of converting the multimodal features into two-dimensional features suitable for linear operations is:
[0019] Reshape the multimodal features of the comprehensive feature map of the visible light image and the multimodal features of the comprehensive feature map of the infrared image to flatten them into a two-dimensional matrix:
[0020] ;
[0021] In the formula, is the multimodal feature of the input, , where B is the batch size, C is the number of channels, H and W are the height and width of the comprehensive feature map respectively;
[0022] Map the flattened 2D matrix to a lower dimension to obtain 2D features suitable for linear operations:
[0023] ;
[0024] In the formula, is the reduced-dimensional weight matrix, is the bias term, and D is the number of channels after dimensionality reduction.
[0025] As a further technical solution, the process of achieving dynamic alignment of two-dimensional features through dynamic adaptive adjustment is:
[0026] Calculate the global mean of two-dimensional features and extract global information:
[0027] The global information is further processed based on the conditional network to generate a conditional vector for adjusting the features:
[0028] The conditional vector is multiplied bit by bit with the two-dimensional feature to obtain the adaptively adjusted feature, ensuring the dynamic alignment of different modal features in the shared space.
[0029] As a further technical solution, a reconstruction module is also used to process the fused image and generate two reconstructed images of different modalities.
[0030] As a further technical solution, the reconstruction module is composed of two independent convolution sub-modules, each module processes the fusion features respectively through two layers of convolution kernels and generates two reconstructed images of different modalities.
[0031] A second aspect of the present invention provides an image multimodal fusion system for complex dynamic environments.
[0032] An image multimodal fusion system for complex dynamic environments, comprising:
[0033] The image preprocessing module is configured to: acquire a visible light image and an infrared image, and preprocess the visible light image and the infrared image;
[0034] The encoder module is configured to: input the preprocessed visible light image and infrared image into the multi-layer encoder respectively, and obtain a comprehensive feature map of the visible light image and a comprehensive feature map of the infrared image containing multi-scale information;
[0035] The bidirectional adapter module is configured to: input the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image into the bidirectional adapter, respectively extract the multimodal features of the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image, convert the multimodal features into two-dimensional features suitable for linear operations, realize dynamic alignment of the two-dimensional features through dynamic adaptive adjustment, and then restore the two-dimensional features adjusted by dynamic alignment to the original dimension; and semantically align the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image restored to the original dimension;
[0036] The image fusion module is configured to: calculate the image gradient information of the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map, determine the fusion weight based on the image gradient information, and fuse the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map using the fusion weight to obtain a fused image. The third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the image multimodal fusion method for complex dynamic environments as described in the first aspect of the present invention.
[0037] The fourth aspect of the present invention provides an electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the method for multimodal image fusion for complex dynamic environments as described in the first aspect of the present invention are implemented.
[0038] One or more of the above technical solutions have the following beneficial effects:
[0039] The present invention realizes the multimodal image fusion of infrared and visible light, and achieves more accurate perception of objects in complex environments by combining the advantages of infrared images in low-light conditions with the high-resolution characteristics of visible light images; performs adaptive two-way adjustment of dimensionality reduction and dimensionality increase, and can adaptively process different modal data and scene changes by dynamically generating conditional vectors, thereby achieving efficient semantic alignment of multimodal features; provides more accurate and robust perception capabilities, and realizes efficient and comprehensive utilization of multimodal information.
[0040] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0042] Figure 1 It is a structural diagram of the SDNet network model of Example 1 of the present invention.
[0043] Figure 2 It is a structural diagram of a bidirectional adapter module of the SDNet network model of Example 1 of the present invention.
[0044] Figure 3 This is a flowchart of an image multimodal fusion method for complex dynamic environments according to Example 1 of the present invention.
[0045] Figure 4 This is a simulation effect diagram of the first embodiment.
[0046] Figure 5 It is a system structure diagram of the second embodiment. DETAILED DESCRIPTION
[0047] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0048] It should be noted that the terms used herein are for describing specific embodiments only and are not intended to be limiting of exemplary embodiments according to the present invention.
[0049] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0050] Embodiment 1
[0051] This embodiment discloses a method for multimodal image fusion in a complex dynamic environment;
[0052] like Figures 1 to 3 As shown, a multimodal image fusion method for complex dynamic environments includes:
[0053] Step S1, acquiring a visible light image and an infrared image, and preprocessing the visible light image and the infrared image;
[0054] Step S2, inputting the preprocessed visible light image and infrared image into a multi-layer encoder respectively to obtain a comprehensive feature map of the visible light image and a comprehensive feature map of the infrared image containing multi-scale information;
[0055] Step S3, inputting the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image into the bidirectional adapter, extracting multimodal features of the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image respectively, converting the multimodal features into two-dimensional features suitable for linear operations, realizing dynamic alignment of the two-dimensional features through dynamic adaptive adjustment, and then restoring the two-dimensional features after dynamic alignment adjustment to the original dimensions; semantically aligning the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image restored to the original dimensions;
[0056] Step S4, calculate the image gradient information of the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map, determine the fusion weight based on the image gradient information, and use the fusion weight to fuse the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map to obtain a fused image. In this embodiment, multimodal image fusion of infrared and visible light is realized, and more accurate perception of objects in complex environments is achieved by combining the advantages of infrared images under low light conditions with the high resolution characteristics of visible light images; adaptive dimensionality reduction and dimensionality increase bidirectional adjustment is performed, and different modal data and scene changes can be adaptively processed by dynamically generating conditional vectors, realizing efficient semantic alignment of multimodal features; providing more accurate and robust perception capabilities, and realizing efficient comprehensive utilization of multimodal information.
[0057] The above steps are implemented by building a semantic-driven fusion network (SDNet), which can be referred to as the SDNet network model. The SDNet network model is not only innovative in feature alignment and fusion, but also improves the computational efficiency and robustness of the system through efficient network architecture design. It can adapt to changing environments and unreliable single-modal inputs, ensuring the stability and reliability of the system in practical applications.
[0058] In some embodiments, the SDNet network model includes a multi-layer encoder, a bidirectional adapter module, a fusion layer, and a reconstruction module, which can effectively process the feature differences between different modalities and achieve high-quality image fusion; the details are described below.
[0059] Step S1, acquiring a visible light image and an infrared image, and preprocessing the visible light image and the infrared image;
[0060] In this embodiment, the public and highly registered MFNet dataset is selected as the main data source. The dataset contains 1083 pairs of infrared and visible light images of different scenes, covering targets in a variety of complex environments, such as pedestrians, vehicles, etc. All image pairs are strictly registered to ensure the consistency of information between modalities, providing a reliable foundation for multimodal image fusion. The image size in the dataset is 640x480 pixels. In order to enhance the generalization ability of the model, the paired infrared images and visible light images are preprocessed. The preprocessing process includes cropping the images into image blocks of 128×128 pixels. In addition, the dataset is further enriched through data expansion technology, including rotation, occlusion, illumination change and other operations. After expansion processing, a total of 18,000 pairs of images and a total of 36,000 images were generated, providing sufficient samples for subsequent model training and evaluation.
[0061] Step S2, inputting the preprocessed visible light image and infrared image into a multi-layer encoder respectively to obtain a comprehensive feature map of the visible light image and a comprehensive feature map of the infrared image containing multi-scale information;
[0062] In step S2, each multi-layer encoder includes multiple encoders connected in sequence. Except for the first encoder, each subsequent encoder is provided with a dense connection block (Dense Connection Block), which connects all the previous encoders through the dense connection block (Dense Connection Block) to reuse the outputs of all the previous encoders; wherein each encoder includes a convolution layer, a batch normalization layer, and an activation layer connected in sequence. The dense connection block includes a splicing module, a convolution layer, a batch normalization layer, and an activation layer, which perform splicing operations, convolution operations, batch normalization operations, and activation function activation operations in sequence; the encoder can capture image features at different scales through a multi-layer convolution structure, from low-level detail information to high-level semantic information, and finally output a comprehensive feature map containing multi-scale information, providing a comprehensive feature basis for the fusion of multimodal images.
[0063] Furthermore, these multi-layer encoders extract features from images of different modalities through multi-scale convolutional layers, ensuring that local details can be obtained while capturing global information. This multi-scale feature extraction strategy enhances the model's perception ability when dealing with complex dynamic environments, especially when the information between modalities is highly complementary.
[0064] Specifically, the input image ,in, is the image height, is the width, is the number of channels, for infrared images =1, grayscale processing is performed on visible light images, so the number of channels is also 1;
[0065] In the first layer of the encoder, in this embodiment, a convolution operation with a convolution kernel of size 3×3 is used to perform preliminary feature extraction on the input image. This process can be expressed as:
[0066] ;
[0067] in, Represents the convolution operation; Represents batch normalization operation; is the LeakyReLU activation function.
[0068] The encoder module further processes the image features through densely connected blocks. Each densely connected block contains multiple layers of convolution operations, and the input of each layer is composed of the concatenation of the output of the current layer and all previous layers, so as to reuse the features of different layers. For the i-th densely connected block, its output feature is expressed as:
[0069] ;
[0070] in: Indicates that the The output features of the layer are concatenated along the channel dimension. Represents convolution operation, batch normalization (BN) and LeakyReLU activation function Used to improve the nonlinear expression ability of the model.
[0071] This densely connected structure can effectively capture feature information at different levels and fuse feature representations from different layers in the channel dimension, thereby capturing global features while retaining details.
[0072] Specifically, in each densely connected block, the input feature map The number of channels gradually increases, taking the setting of three-layer encoder as an example:
[0073] For encoder 1: the input is The output is obtained through convolution and activation function ;
[0074] For encoder 2: the input is , the output is ;
[0075] For encoder 3: the input is , the output is ;
[0076] This design of multi-layer encoders ensures that features are continuously accumulated in the channel dimension, so that the network can reuse feature information from previous layers at each layer, enhancing the feature transfer capability. After being processed by multiple convolutional layers, the final output of the encoder module is a comprehensive feature map that combines low-level and high-level features, expressed as:
[0077] ;
[0078] The final feature representation The dimensions are:
[0079] ;
[0080] Through the structure of the multi-layer encoder of this embodiment, it is possible to capture the features of the input image at different scales, retain low-level detail information through dense feature splicing, and extract higher-level semantic information through multi-layer convolution.
[0081] Step S3, input the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image into the bidirectional adapter, extract the multimodal features of the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image respectively, convert the multimodal features into two-dimensional features suitable for linear operations, realize dynamic alignment of the two-dimensional features through dynamic adaptive adjustment, and then restore the two-dimensional features after dynamic alignment adjustment to the original dimension; semantically align the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image restored to the original dimension; traditional multimodal fusion methods often have difficulty in achieving effective semantic alignment when dealing with representation differences between modalities. Figure 1 ,The SDNet network model of this embodiment innovatively designs a bidirectional adapter module,which is used to perform adaptive bidirectional adjustment in the feature space of different,modalities.
[0082] Specifically, in step S3, the bidirectional adapter module adjusts the features between the modalities by reducing the dimension and upsampling to ensure that the features achieve higher compatibility and consistency before fusion. Through this bidirectional adaptation mechanism, the differences in the representation forms of infrared and visible light images can be effectively processed to achieve semantic alignment between the modalities.
[0083] The bidirectional adapter module of this embodiment is the core innovative part of the system to realize multimodal feature alignment and fusion. This module can dynamically adjust the feature representation of different modes through an adaptive feature adjustment mechanism combined with the Reshape operation, ensuring the effective processing of features in the process of feature dimension reduction, adjustment and dimension increase. Especially when processing multimodal data such as infrared and visible light, a more efficient modal alignment is achieved through a bidirectional transformation strategy of dimension reduction and dimension increase, combined with a conditional network for dynamic feature adjustment. In this embodiment, the bidirectional adapter module further improves the stability and efficiency of model training through appropriate parameter initialization methods. The strategy of Xavier initialization and bias initialization to 0 is adopted to ensure that the network can maintain stable gradient propagation in the early stages of training, thereby avoiding the phenomenon of gradient disappearance or explosion.
[0084] Step S31, extracting multimodal features from the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image obtained in step S2, first performing dimensionality reduction processing on the extracted multimodal features, and then performing a flattening (Reshape) operation to convert the four-dimensional multimodal features into two-dimensional features suitable for linear operations;
[0085] Assume that the multimodal features of its input are , where B is the batch size, C is the number of channels, H and W are the height and width of the feature map respectively;
[0086] To perform linear transformation, first reshape the features and flatten them into a two-dimensional matrix. The formula is as follows:
[0087] ;
[0088] On this basis, a dimensionality reduction operation is performed to further map the flattened high-dimensional features to lower dimensions:
[0089] ;
[0090] in, is the reduced-dimensional weight matrix, is the bias term, and D is the number of channels after dimensionality reduction.
[0091] The above process will input the feature map From high-dimensional space Projection into low-dimensional space , effectively compresses high-dimensional features and reduces computational complexity;
[0092] Furthermore, Xavier uniform initialization is used to initialize the weights of the linear layer Initialization is performed and all bias items are initialized to zero; this ensures that the gradient can be propagated normally during dimensionality reduction, adaptive adjustment, and dimensionality increase, avoiding the problem of gradient disappearance or explosion.
[0093] Among them, Xavier initialization (also known as Glorot initialization) is a weight initialization method;
[0094] In this embodiment, Xavier initialization is used to maintain the stability of feature mapping in each stage of dimensionality reduction, adjustment, and dimensionality increase, and all bias items are initialized to zero, which can prevent unnecessary output offsets in the early stage of training, thereby improving the stability and efficiency of model training.
[0095] Step S32, introduce a conditional network to dynamically and adaptively adjust the two-dimensional features to achieve dynamic alignment of features of different modalities. After dimensionality reduction, the dynamic adaptive adjustment of features is further enhanced by the introduced conditional network. A conditional vector is generated based on the global information of the input multimodal features. , used to guide the adjustment of features. Specifically,
[0096] Step S321, calculate the global mean of the two-dimensional features and extract the global information of the two-dimensional features:
[0097] ;
[0098] Step S322: further process the global information through a conditional network to generate a conditional vector for adjusting features:
[0099] condition ;
[0100] in, , are the weight matrices of the conditional network, , is the bias term, the activation function is a nonlinear function, such as ReLU. The generated conditional vector will dynamically adjust the feature representation after dimensionality reduction.
[0101] Step S323, multiplying the conditional vector and the two-dimensional feature bit by bit to obtain the adaptively adjusted feature, so as to ensure the dynamic alignment of different modal features in the shared space.
[0102] ;
[0103] The above adaptive adjustment mechanism ensures that features of different modalities can be dynamically aligned in a shared space, ensuring the consistency of feature performance in different modalities and scenarios.
[0104] The adjusted feature representation is further optimized through nonlinear mapping, and the Reshape operation is performed to restore the feature to the dimension C before the flattening operation in step 31 to obtain the deformed low-dimensional feature. The formula is as follows:
[0105] ;
[0106] in, Indicates activation operation, such as Figure 2 As shown, the LeakyRelu function can be used for activation operation;
[0107] The adjusted and optimized two-dimensional features are restored to the original four-dimensional form through the Reshape operation, preparing for the next dimensionality increase operation:
[0108] ;
[0109] Through the Reshape operation, the feature is restored from the flattened two-dimensional form to the original four-dimensional form so that the dimensionality increase process can continue.
[0110] Step S33, after using upsampling to restore the two-dimensional features after dynamic adaptive adjustment to the original dimension, semantically align the multimodal features of the infrared image comprehensive feature map and the visible light image comprehensive feature map to obtain the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map. Specifically, based on the features after adaptive adjustment and Reshape restoration, the module restores the features to the original high-dimensional space through an upsampling (dimensionality increase) operation to ensure that the feature dimension is consistent with the input. First, the low-dimensional features are flattened through the Reshape operation:
[0111] ;
[0112] Then perform a dimensionality increase operation to restore the low-dimensional features to the same dimension as the input features:
[0113] ;
[0114] in, is the dimension-raising weight matrix, is the bias term.
[0115] Finally, the original four-dimensional features are restored through the Reshape operation:
[0116] ;
[0117] This step ensures that the adjusted features are consistent with the input features in terms of dimension. In the bidirectional adapter, in order to achieve semantic alignment of the two modal features, the infrared modal features and the visible light modal features are first concatenated in the channel dimension to obtain a temporary joint feature representation; then the joint features are adaptively adjusted to align the features of the two modalities in the semantic space; finally, the adjusted features are re-separated to obtain semantically aligned infrared modal features and visible light modal features. This processing method achieves effective alignment of different modal features in the semantic space through joint analysis and adjustment of features, laying the foundation for subsequent adaptive fusion.
[0118] Assume that the infrared modal characteristics are , the visible light modal characteristics are , then the fused features are expressed as:
[0119] ;
[0120] Through this feature processing method of concatenating first and then separating, the bidirectional adapter achieves semantic alignment of features of different modalities, making the features of the two modalities consistent in the semantic space. The aligned features retain the characteristics of each modality while ensuring their comparability at the semantic level, laying the foundation for subsequent adaptive fusion based on gradient information.
[0121] This embodiment solves the problem of processing the representation differences and information alignment between different modal data such as infrared and visible light images by setting a bidirectional adapter module, and realizes accurate semantic alignment and information fusion;
[0122] In the above scheme, the bidirectional adapter module in step 2 realizes flexible adjustment and adaptive mapping of features in the process of dimensionality reduction, adjustment and dimensionality increase by introducing the Reshape operation and conditional network. The Reshape operation ensures that features can be seamlessly converted between different dimensions, enhancing the flexibility of features in multimodal tasks. The conditional network can adaptively process different modal data and scene changes by dynamically generating conditional vectors, ensuring accurate alignment of features. This design not only solves the heterogeneity and feature alignment problems of multimodal data, but also greatly improves the robustness of the system. Even in the case of incomplete or low-quality data, it can still maintain high perceptual performance through adaptive adjustment to ensure stable output of the system. In addition, the combination of Xavier initialization and bias initialization further enhances the stability of the network during training, avoids the phenomenon of gradient disappearance or gradient explosion, and ensures efficient training and stable convergence of the model.
[0123] Step S4, calculate the image gradient information of the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map, determine the fusion weight based on the image gradient information, and use the fusion weight to fuse the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map to obtain a fused image. , the comprehensive feature map of visible light after semantic alignment is , the feature map of the infrared mode and the feature map of the visible light mode to be spliced are spliced in the channel dimension, and the feature maps of the two modes are spliced as follows:
[0124] ;
[0125] in, This operation retains the feature information of the two modalities and combines them into a higher-dimensional feature representation;
[0126] The concatenated features are convolved (conv), batch normalized (BN), and activated to obtain the fused features;
[0127] ;
[0128] in, and The convolution kernels are 3*3 and 1*1 respectively. and is the bias term, is an activation function used to limit the output range to (-1, 1); is the LeakyReLU activation function.
[0129] In step 4 of this embodiment, the fusion weight is adaptively adjusted by dynamically calculating the image gradient information. Specifically, the dynamic weight factor S1 is generated by calculating the gradient difference between the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map. The weight factor automatically adjusts the fusion ratio of the two modalities according to the gradient features of different regions. When the gradient information of a certain modality (infrared or visible light) is more significant, the modality obtains a higher weight in the fusion process, thereby ensuring that the fusion result can better retain significant features.
[0130] Specifically, the method of dynamically calculating image gradient information to adaptively adjust fusion weights includes the following:
[0131] Using the Sobel operator and Compute the gradient of the input image:
[0132] ;
[0133] For the input image , the gradient amplitude is defined as:
[0134] ;
[0135] Perform low-pass filtering on the image to be fused: Smooth the image. By preprocessing the original image in this way, the subsequent adaptive weight calculation can be more accurate and reliable.
[0136] The filter matrix is defined as:
[0137] ;
[0138] Significant gradient region selection: Based on the low-pass filtered image, the gradient difference between the infrared image and the visible light image is calculated to generate a dynamic weight factor S1, and S1 is used to extract regions with significant gradient changes;
[0139] The dynamic weight factor S1 is calculated as follows:
[0140] ;
[0141] in; Indicates gradient calculation; low_pass and Respectively represent the infrared image and visible light images Perform low-pass filtering, Represents the absolute value operation, and sign is the sign function, which is used to extract the sign (positive or negative) of the result.
[0142] This formula is used to calculate the mask of the significant gradient area, and the significant area is determined by comparing the difference in gradient size between infrared and visible light images.
[0143] based on Calculate the fusion loss :
[0144] ;
[0145] in, Represents the gradient of the fused image. Represents the gradient of the infrared image Represents the gradient of a visible light image. is a mask that indicates which regions of gradient differences should be given more attention in infrared images and which regions should be given more attention in visible light images. The value of is 0 or 1, depending on the significance of the current region. Express The inversion of is used to focus on the opposite area. Mean is used to calculate the average value of the gradient differences in different regions, with the aim of smoothing these gradient differences and obtaining an overall loss value.
[0146] When the gradient features of the infrared image are more significant, S1=1, and more emphasis is placed on maintaining the gradient information of the infrared image. When the gradient features of the visible light image are more significant, S1=-1, and more emphasis is placed on maintaining the gradient information of the visible light image. The weight distribution is automatically determined based on the local features of each position.
[0147] In addition, in this embodiment, in order to verify the quality of the fused features and enhance the robustness of the system, a reconstruction module is set as a feature decoder. The reconstruction module contains two independent decoding submodules, corresponding to infrared reconstruction and visible light reconstruction, and plays multiple roles in the entire system: first, as a decoder, the fused high-dimensional features are mapped back to visible image data through a two-layer convolution structure to achieve the conversion from feature space to image space; secondly, the network is guided to optimize feature extraction and fusion strategies through the back propagation of reconstruction loss, helping the network to establish associations between different modalities during the training process and enhance the complementarity of features; in addition, the reconstruction module also provides a supervision verification mechanism to evaluate whether the fused features effectively retain the original image information through the reconstruction quality. More importantly, the reconstruction module significantly improves the robustness of the system by introducing a redundancy mechanism. Even if the quality of a certain modality data is reduced or missing, the system can still use the information of another modality for effective reconstruction, thereby ensuring the stability of perception. This multifunctional design makes the reconstruction module not just a simple decoder, but a key component to ensure system performance and reliability.
[0148] Specifically, each reconstruction module includes two convolutional layers, which perform reverse processing on the fused features through convolution operations to restore the image details of the original modality;
[0149] In this embodiment, the multimodal image includes two, and the reconstruction module receives the fused features , and generates two independent outputs through two different reconstruction modules, namely and , used to reconstruct images of two different modalities;
[0150] Specifically, the reconstruction process of each reconstruction module includes the following:
[0151] The fused feature map is convolved to expand the number of channels, batch normalized, and activated to obtain the intermediate features;
[0152] Each reconstruction consists of two convolutional layers. The first convolution layer performs convolution operations on the fused features, expands the number of input channels from 1 to 64, and uses batch normalization (BN) and LeakyReLU activation function ( ) to perform nonlinear transformation.
[0153] The infrared reconstruction module is used to reconstruct the feature map of the infrared modality after the fusion layer. The formula of the first convolution layer is as follows:
[0154] ;
[0155] in, is the convolution kernel, is the bias term; output feature is an intermediate feature;
[0156] For the second layer of convolution, the intermediate features are convolved to reduce the number of channels and activated to obtain the reconstructed image. Specifically, the intermediate features are processed again, the number of channels is reduced from 64 to 1, and the Tanh activation function is used to limit the output value range to (-1,1) to generate the final reconstructed image:
[0157] ;
[0158] in, is the convolution kernel, is the bias term. Output image is the first modality image after reconstruction.
[0159] The visible light reconstruction module is the same as the infrared reconstruction module. The visible light reconstruction module is used to reconstruct visible light images.
[0160] Specifically, the output of the infrared reconstruction module is:
[0161] ;
[0162] The output of the visible light reconstruction module is:
[0163] ;
[0164] The reconstruction module of this embodiment consists of two independent convolution submodules. Each module processes the fused features through two layers of convolution kernels (both 3×3) and Tanh activation functions, and generates two reconstructed images of different modalities. This design ensures that the fused features can restore the image details of the original infrared and visible light modalities, and improves the accuracy and robustness of the reconstruction.
[0165] In this embodiment, the reconstruction module not only undertakes the task of modal data reconstruction, but more importantly, provides a verification mechanism through the reconstruction task. In addition, the function of the decoder is also realized to a certain extent. For the reconstruction module in this embodiment, on the one hand, it can verify whether the fused features effectively retain the key information of the original image by calculating the reconstruction loss; on the other hand, by using the reconstruction loss as a supervisory signal for the training process, it can guide the network to learn better feature representation. This module reconstructs the original images of the infrared and visible light modalities from the fused features, thereby enhancing the robustness of the system in the case of unreliable or missing modal data. In addition, the reconstruction module provides a redundancy mechanism, which can still recover useful information from the features of another modality even if the data quality of a certain modality is degraded or missing, ensuring the stability of perception. In actual complex scenes, this design significantly improves the reliability and adaptability of the system. Therefore, the reconstruction module is not only used to map the high-dimensional fused features back to the image space, but also retains the details and resolution of the generated image through multi-layer convolution operations, thereby improving the clarity and accuracy of perception. This module effectively combines the reconstruction and decoding functions to achieve fine reconstruction after multi-modal fusion.
[0166] A further technical solution is to use an end-to-end learning architecture during the training of the SDNet network model, input the preprocessed image into the SDNet network model for training, and obtain the reconstructed image; the end-to-end learning architecture jointly optimizes the feature extraction, alignment and fusion processes, making the system more efficient in training and application.
[0167] In addition, the loss value of the reconstructed image and the fused image is calculated based on the loss function, the model parameters are adjusted according to the loss value, and the training is iterated until the loss value reaches the set accuracy requirement to obtain the trained SDNet network model;
[0168] In this embodiment, the loss function includes reconstruction loss, structural similarity (SSIM) loss, and mean square error (MSE) loss;
[0169] The two reconstruction output reconstruction losses can be used to measure the difference between the fused image and the reconstruction results of each modality;
[0170] (1) Reconstruction loss, expressed as mean square error (MSE):
[0171] ;
[0172] is the fused image, and Corresponding to the results reconstructed by infrared mode and visible light mode respectively;
[0173] (2) Structural similarity (SSIM) loss is used to measure the structural similarity between two images. SSIM loss is calculated as 1 minus the SSIM index:
[0174] ;
[0175] in, and They are the infrared modality and visible light modality input images respectively.
[0176] (3) MSE loss:
[0177] The mean square error (MSE) loss calculates the pixel-level difference between two images and is expressed as:
[0178] ;
[0179] in, is the total number of pixels in the image. It can be an infrared or visible light modality image, or a reconstructed target image.
[0180] The final loss function is expressed as:
[0181] + + ;
[0182] in, , and It is the weight hyperparameter of the loss, which is used to balance the impact of each loss on the total loss.
[0183] In the experiment, SDNet was used as the multimodal image fusion model. To ensure that the network can effectively learn multimodal features, the PyTorch framework was used to implement the entire algorithm, and the model was trained using the Adam optimizer. The initial learning rate was set to 0.001, and the batch size for model training and evaluation was set to 64. The entire training process ran for 15 epochs. During the training process, the gradient of the model was updated through the back-propagation mechanism, and the optimizer was used to update the model parameters after each iteration to minimize the fusion loss. All experiments were conducted in an environment accelerated by a graphics processing unit (GPU). , and The hyperparameter values are set to 0.1, 1, and 10.
[0184] The above-mentioned complex dynamic environment perception method based on multimodal fusion can be applied in the field of intelligent security, which includes but is not limited to intelligent monitoring systems, intelligent access control systems, and abnormal behavior detection systems;
[0185] Based on multimodal perception that fuses infrared and visible light images, this innovative technology provides an all-weather, all-round intelligent monitoring network by seamlessly integrating data from high-definition visible light cameras and infrared thermal imaging devices.
[0186] In video surveillance systems, it can not only capture clear and detailed color images during the day, but also accurately identify and track targets through thermal imaging at night or in low-light environments, effectively overcoming the limitations of traditional single-modality systems.
[0187] Its application in intelligent access control systems significantly improves the accuracy of identity authentication and anti-counterfeiting capabilities by combining visible light face recognition and infrared life feature detection.
[0188] In terms of abnormal behavior detection, the combination of infrared and visible light enables the system to accurately identify suspicious activities under various light conditions, such as wandering at night or abnormal heat sources in hidden places, enhancing the deterrence and response speed of the security system. For large public places such as airports, shopping malls or stadiums, this dual-modal fusion technology can achieve all-weather crowd monitoring, suspicious object identification and rapid response to emergencies. Through advanced image processing algorithms, the system continuously learns and adapts from the fused data to improve its understanding of complex scenes. This security solution based on the fusion of infrared and visible light not only improves the efficiency and accuracy of security protection, but also effectively reduces the false alarm rate caused by light changes, allowing security personnel to deal with potential threats more targeted. With the continuous optimization of algorithms and the improvement of hardware performance, the application prospects of infrared and visible light image fusion technology in the field of intelligent security are becoming more and more broad, driving the entire industry to develop in a smarter and more reliable direction, and making important contributions to building an all-weather, efficient security environment.
[0189] To illustrate the effect of this embodiment, a simulation experiment was conducted to fusion process the sample images in the data set. Figure 4As shown in the figure, it can be seen that the image effect after the fusion processing of the example images in the data set in this embodiment shows that the method of this embodiment successfully retains the highlight features of the human target in the infrared image (such as the outline of the person in the figure is clearly visible), and at the same time, the scene structure information in the visible light image (such as the outline of the building, street lights, etc.) is well combined with the thermal radiation information in the infrared image through fusion. The fusion result shows good scene integrity, the structural features such as the outline and windows of the building are clearly distinguishable, and the thermal target in the infrared image is highlighted. While maintaining the advantages of the original images, the fusion method eliminates the limitations of single-mode imaging and provides more comprehensive scene information. Especially in complex dynamic environments, the method of this embodiment effectively solves the difference problem in feature expression of different modal images through a bidirectional adapter module, so that the fused image can not only maintain the advantages of infrared thermal imaging in target detection, but also reflect the characteristics of visible light images in spatial detail description. The experimental results show that the proposed method not only achieves good results in target edge retention, detail texture fidelity and contrast, but also shows strong robustness and adaptability when processing images under different lighting and weather conditions. This excellent fusion performance provides high-quality image input for subsequent intelligent monitoring, target tracking and other tasks, effectively improving environmental perception capabilities in complex scenarios.
[0190] Embodiment 2
[0191] This embodiment discloses an image multimodal fusion system for complex dynamic environments;
[0192] like Figure 5 As shown, an image multimodal fusion system for complex dynamic environments includes:
[0193] An image multimodal fusion system for complex dynamic environments, comprising:
[0194] The image preprocessing module is configured to: acquire a visible light image and an infrared image, and preprocess the visible light image and the infrared image;
[0195] The encoder module is configured to: input the preprocessed visible light image and infrared image into the multi-layer encoder respectively, and obtain a comprehensive feature map of the visible light image and a comprehensive feature map of the infrared image containing multi-scale information;
[0196] The bidirectional adapter module is configured to: input the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image into the bidirectional adapter, respectively extract the multimodal features of the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image, convert the multimodal features into two-dimensional features suitable for linear operations, realize dynamic alignment of the two-dimensional features through dynamic adaptive adjustment, and then restore the two-dimensional features adjusted by dynamic alignment to the original dimension; and semantically align the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image restored to the original dimension;
[0197] The image fusion module is configured to: calculate the image gradient information of the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map, determine the fusion weight based on the image gradient information, and fuse the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map using the fusion weight to obtain a fused image. In the image fusion module, based on the semantically aligned features, the gradient difference between the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map is calculated to generate a dynamic weight factor S1. The weight factor automatically adjusts the fusion ratio of the two modalities according to the gradient features of different regions. When the gradient information of a certain modality is more significant, the modality obtains a higher weight in the fusion process, thereby ensuring that the fusion result can better retain significant features.
[0198] Furthermore, in this embodiment, an image multimodal fusion system for complex dynamic environments also includes a reconstruction module, which acts as a feature decoder to reconstruct visible light images and infrared images based on the fused images. On the one hand, the reconstruction module can guide the network to optimize feature extraction and fusion strategies through the back propagation of reconstruction loss, helping the network to establish associations between different modalities during training and enhance the complementarity of features; in addition, the reconstruction module can also evaluate whether the fused features effectively retain the original image information through the reconstruction quality. More importantly, the reconstruction module significantly improves the robustness of the system by introducing a redundancy mechanism. Even when the quality of a certain modality data is reduced or missing, the system can still use the information of another modality for effective reconstruction, thereby ensuring the stability of perception.
[0199] Embodiment 3
[0200] The purpose of this embodiment is to provide a computer-readable storage medium.
[0201] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for multimodal image fusion in a complex dynamic environment as described in Example 1 of the present disclosure.
[0202] Embodiment 4
[0203] The purpose of this embodiment is to provide an electronic device.
[0204] An electronic device comprises a memory, a processor and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in a method for multimodal fusion of images for complex dynamic environments as described in Example 1 of the present disclosure are implemented.
[0205] The steps involved in the apparatuses of the above embodiments 2, 3 and 4 correspond to the method embodiment 1, and the specific implementation methods can refer to the relevant description part of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0206] Those skilled in the art should understand that the modules or steps of the present invention described above can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0207] Although the above describes the specific implementation mode of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.
Claims
1. A multimodal image fusion method for complex dynamic environments, characterized in that: include: Acquire a visible light image and an infrared image, and preprocess the visible light image and the infrared image; The preprocessed visible light image and infrared image are respectively input into a multi-layer encoder to obtain a comprehensive feature map of the visible light image and a comprehensive feature map of the infrared image containing multi-scale information; the multi-layer encoder includes a plurality of encoders connected in sequence, and except for the first encoder, each of the following encoders is respectively provided with a dense connection block, and all the previous encoders are connected through the dense connection block, and the outputs of all the previous encoders are multiplexed; The comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image are input into the bidirectional adapter, and the multimodal features of the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image are respectively extracted, and the multimodal features are converted into two-dimensional features suitable for linear operations, and dynamic alignment of the two-dimensional features is realized through dynamic adaptive adjustment, and then the two-dimensional features adjusted by dynamic alignment are restored to the original dimension; the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image restored to the original dimension are semantically aligned; The multimodal features are converted into two-dimensional features suitable for linear operations, dynamic alignment of the two-dimensional features is achieved through dynamic adaptive adjustment, and then the two-dimensional features after dynamic alignment adjustment are restored to the original dimension; the process of semantically aligning the comprehensive feature map of the visible light image and the comprehensive feature map of the infrared image restored to the original dimension is as follows: Reshape the multimodal features of the comprehensive feature map of the visible light image and the multimodal features of the comprehensive feature map of the infrared image to flatten them into a two-dimensional matrix: map the flattened two-dimensional matrix to a lower dimension to obtain two-dimensional features suitable for linear operations: use Xavier uniform initialization to initialize the weights of the linear layer and initialize all bias items to zero; A conditional network is introduced to dynamically and adaptively adjust the two-dimensional features. Specifically: the global mean of the two-dimensional features is calculated and global information is extracted: the global information is further processed based on the conditional network to generate a conditional vector for adjusting the features: the conditional vector is multiplied bit by bit with the two-dimensional features to obtain the adaptively adjusted features, ensuring the dynamic alignment of different modal features in the shared space; the adjusted feature representation is further optimized through nonlinear mapping, and the Reshape operation is performed to restore the features to the dimensions before the flattening operation to obtain the deformed low-dimensional features; the adjusted and optimized two-dimensional features are restored to the original four-dimensional form through the Reshape operation; the features are flattened again through the Reshape operation, and then the dimensionality increase operation is performed to restore the low-dimensional features to the same dimension as the input features: finally, the original four-dimensional features are restored through the Reshape operation; Calculate the image gradient information of the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map, and determine the fusion weight based on the image gradient information, wherein the process of determining the fusion weight based on the image gradient information is as follows: generate a dynamic weight factor by calculating the gradient difference between the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map, and the weight factor automatically adjusts the fusion ratio of the two modes according to the gradient characteristics of different regions; specifically, use the Sobel operator to calculate the gradient of the input image and define the gradient amplitude of the input image; perform low-pass filtering on the image to be fused: calculate the gradient difference between the infrared image and the visible light image based on the low-pass filtered image, generate a dynamic weight factor, extract the area with gradient change based on the dynamic weight factor and calculate the fusion loss, and then determine the fusion weight; The semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map are fused using the fusion weight to obtain a fused image.
2. The method for multimodal image fusion in a complex dynamic environment as claimed in claim 1, characterized in that: The preprocessing process includes: cutting the visible light image and the infrared image into image blocks of a specific size; and expanding the number of the image blocks by using a data expansion technology.
3. The method for multimodal image fusion in a complex dynamic environment as claimed in claim 1, characterized in that: The method also includes processing the fused image using a reconstruction module and generating two reconstructed images of different modalities.
4. The method for multimodal image fusion in a complex dynamic environment as claimed in claim 3, characterized in that: The reconstruction module consists of two independent convolution submodules, each of which processes the fusion features through two layers of convolution kernels and generates two reconstructed images of different modalities.
5. An image multimodal fusion system for complex dynamic environments, characterized in that: include: The image preprocessing module is configured to: acquire a visible light image and an infrared image, and preprocess the visible light image and the infrared image; The encoder module is configured to: input the preprocessed visible light image and infrared image into a multi-layer encoder respectively, and obtain a visible light image comprehensive feature map and an infrared image comprehensive feature map containing multi-scale information; the multi-layer encoder includes a plurality of encoders connected in sequence, and except for the first encoder, each of the following encoders is respectively provided with a dense connection block, and all the previous encoders are connected through the dense connection block, and the outputs of all the previous encoders are multiplexed; The bidirectional adapter module is configured to: input the visible light image comprehensive feature map and the infrared image comprehensive feature map into the bidirectional adapter, extract the multimodal features of the visible light image comprehensive feature map and the infrared image comprehensive feature map respectively, convert the multimodal features into two-dimensional features suitable for linear operation, realize dynamic alignment of the two-dimensional features through dynamic adaptive adjustment, and then restore the two-dimensional features adjusted by dynamic alignment to the original dimension; perform semantic alignment on the visible light image comprehensive feature map and the infrared image comprehensive feature map restored to the original dimension; the process of converting the multimodal features into two-dimensional features suitable for linear operation, realizing dynamic alignment of the two-dimensional features through dynamic adaptive adjustment, and then restoring the two-dimensional features adjusted by dynamic alignment to the original dimension; and performing semantic alignment on the visible light image comprehensive feature map and the infrared image comprehensive feature map restored to the original dimension is: Reshape the multimodal features of the comprehensive feature map of the visible light image and the multimodal features of the comprehensive feature map of the infrared image to flatten them into a two-dimensional matrix: map the flattened two-dimensional matrix to a lower dimension to obtain two-dimensional features suitable for linear operations: use Xavier uniform initialization to initialize the weights of the linear layer and initialize all bias items to zero; A conditional network is introduced to dynamically and adaptively adjust the two-dimensional features. Specifically: the global mean of the two-dimensional features is calculated and global information is extracted: the global information is further processed based on the conditional network to generate a conditional vector for adjusting the features: the conditional vector is multiplied bit by bit with the two-dimensional features to obtain the adaptively adjusted features, ensuring the dynamic alignment of different modal features in the shared space; the adjusted feature representation is further optimized through nonlinear mapping, and the Reshape operation is performed to restore the features to the dimensions before the flattening operation to obtain the deformed low-dimensional features; the adjusted and optimized two-dimensional features are restored to the original four-dimensional form through the Reshape operation; the features are flattened again through the Reshape operation, and then the dimensionality increase operation is performed to restore the low-dimensional features to the same dimension as the input features: finally, the original four-dimensional features are restored through the Reshape operation; The image fusion module is configured to: calculate the image gradient information of the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map, and determine the fusion weight based on the image gradient information, wherein the process of determining the fusion weight based on the image gradient information is: by calculating the gradient difference between the semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map, a dynamic weight factor is generated, and the weight factor automatically adjusts the fusion ratio of the two modalities according to the gradient characteristics of different regions; specifically, the Sobel operator is used to calculate the gradient of the input image and the gradient amplitude of the input image is defined; the image to be fused is low-pass filtered: based on the image after low-pass filtering, the gradient difference between the infrared image and the visible light image is calculated, the dynamic weight factor is generated, the area of gradient change is extracted based on the dynamic weight factor and the fusion loss is calculated, and then the fusion weight is determined; The semantically aligned infrared image comprehensive feature map and the visible light image comprehensive feature map are fused using the fusion weight to obtain a fused image.
6. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps of the image multimodal fusion method for complex dynamic environments as described in any one of claims 1 to 4 are implemented.
7. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the image multimodal fusion method for complex dynamic environments as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Infrared and visible light image fusion method with regional attention
CN114782298A
Substation power equipment detection method and system based on infrared and visible light fusion
CN117557775A
Infrared image and visible light image fusion method and device and storage medium
CN118657671A
Data report generation method based on deep learning
CN119150810A